Skip to content

About

Documentation and audit of a self-hosted personal LLM agent system — architecture, permission design, and observed failure modes.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

1 Commit

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 

Repository files navigation

Amber — a self-hosted personal agent system

A small suite of LLM agents I run on my own VPS to handle daily briefings, research, and writing. This repo documents the architecture and what I learned building it — it is not a framework, and it does not contain my configuration.

Running in production on my own hardware since March 2026. This writeup is an audit of my own setup — several of the failure modes below are things I found while writing it.

I built this because I wanted to understand what it actually takes to give an agent access to my email, calendar, and files without it doing something I would regret. Most of what I learned is in What I've learned about constraining agents and Failure modes.


Architecture

flowchart TB
    H["Me (Telegram)"]

    subgraph GW["Agent runtime (self-hosted, single container)"]
        A["Amber<br/>main / chat interface"]
        S["Scout<br/>research"]
        W["Writer<br/>final deliverables"]
        P["Pulse<br/>AI news"]
    end

    subgraph SCHED["Scheduled jobs"]
        C1["morning briefing — daily"]
        C2["AI news — daily"]
    end

    subgraph FS["Workspace (files, git-versioned)"]
        WS["Per-agent workspaces"]
        HO["research-handoffs/"]
    end

    subgraph EXT["External access"]
        G1["Google client: read-only<br/>Gmail / Calendar / Drive"]
        G2["Google client: docs-write<br/>Docs + Sheets only"]
        WEB["Web search + fetch"]
        SYNC["Fitness data sync<br/>(host cron, writes files)"]
    end

    H <--> A
    C1 --> A
    C2 --> P
    A -->|delegates research| S
    S -->|hands off| W
    A --> P

    S --> HO
    HO --> W
    A <--> WS
    S <--> WS
    W <--> WS
    P <--> WS

    A --> G1
    W --> G2
    S --> WEB
    P --> WEB
    SYNC --> WS

    style G1 fill:#2d5016,color:#fff
    style G2 fill:#5c3a00,color:#fff
    style H fill:#1e3a5f,color:#fff
Loading

Everything runs behind a private network overlay. No inbound ports are open to the public internet; the machine answers ping and nothing else.


The agents

Agent Purpose MCP servers Access Human in the loop
Amber Main chat interface. Daily briefing, triage, general assistance. Delegates rather than doing deep work itself. none Google read-only client; web search/fetch; workspace files Every conversation. This is the surface I actually talk to
Scout Deep research. Validates sources, structures comparable findings, decides when to parallelize. Never talks to me directly. none Web search/fetch; workspace + handoff directory Indirect — I see the output through Amber or Writer
Writer Turns consolidated research into finished deliverables: documents, tables, summaries. none docs-write Google client (Docs/Sheets only); workspace I review the deliverable. Cannot delegate to any other agent
Pulse Daily AI news. Explicit editorial priorities; separates fact from speculation and says when something is hype. none Web search/fetch; own workspace I read the daily output

On the MCP column: there are none, and that is worth being explicit about. I expected to need MCP servers and ended up not using them. External data reaches the agents as files: a host-level cron job pulls from an external API, writes a plain-text summary into the workspace, and the agent reads it like any other file. This is less elegant than a live tool call, but it has a property I came to value — the integration cannot be invoked by the agent at all. It is a one-way data drop. The agent cannot call the API, cannot authenticate to it, and cannot make it do anything.


What I've learned about constraining agents

Split credentials, don't just write rules

The control I trust most is not in a prompt. Google Workspace access is split across two separate OAuth clients:

  • a read-only client (Gmail, Calendar, Drive, Contacts) used for everything by default
  • a write client scoped narrowly to creating Docs and Sheets, and nothing else

An agent operating under the default client cannot send email, cannot modify a calendar event, and cannot delete a Drive file — not because it was told not to, but because the credential does not carry the scope. The prompt-level policy sits on top of that as a second layer, with an explicit deny list (send, reply, draft, archive, trash, label, create, update, delete, share, move) and a rule that write access is never inferred from the default client.

The same principle shows up in the unglamorous parts of the system. The nightly backup job pushes to GitHub with a deploy key scoped to a single repository — it can write to the backup repo and nothing else. It cannot create repositories and cannot reach any of my others. I was reminded of this while writing this README: I went looking for a way to publish it straight from the server, and found that the server correctly cannot.

This distinction turned out to be the whole lesson. Instructions constrain a cooperative agent. Credentials constrain any agent.

Approval has to be scoped to the action

Write operations require approval in the same conversation, for that exact action. Not standing approval, not "you can manage my calendar." The narrower rule is easier for me to reason about, because I never have to remember what I authorized last week.

There is one deliberate exception, and it is written down: asking for the morning briefing counts as approval for read-only weather and mail access for that run. Read-only, bounded, stated up front.

Delegation is a permission, and it should be asymmetric

Each agent has an explicit list of which other agents it may invoke. Writer's list is empty — it is a leaf node by design, because it is the agent closest to producing something that leaves the system. Amber can reach the three specialists. The asymmetry is the point: capability to delegate is itself a privilege worth restricting.

Move handoffs off the chat surface

Research moves between agents through files in a dedicated handoff directory, never by pasting into chat. I did this for output quality — long research pasted into Telegram is unreadable — but the security properties turned out to be better too. Handoffs are inspectable after the fact, diffable, and versioned in git. Chat is ephemeral and I would never have reviewed it.

Define what does not deserve your attention

The monitoring agent's spec is mostly a list of things that should not reach me: routine heartbeats, "everything is fine," ordinary weather, non-urgent mail, vague reminders. If nothing qualifies, it emits a single token and stays quiet. Alerting policies tend to be written as "tell me if X"; writing the negative list explicitly is what made it usable, and an agent that cries wolf gets ignored, which is its own safety failure.

Identity as inspectable state

Each agent's identity, operating rules, and memory live as markdown files in a git-versioned workspace rather than being buried in a system prompt. A new agent boots from a BOOTSTRAP.md that it deletes once it has written its own identity file. I can read what an agent believes about itself, diff how it changed, and revert it. This started as a convenience and became my main debugging tool.

What I deliberately left out

No autonomous outbound communication. No agent sends email, posts publicly, or messages anyone but me. No payments, no purchases, no infrastructure changes. The one social-media integration I built was read-only and I have since retired it.


Failure modes observed in production

These are real, and most of them are failures of monitoring, not of agent behavior. That turned out to be the pattern: the agents rarely did something alarming; things quietly stopped working and I did not notice.

1. A silent process leak degraded a service for four months. A third-party agent runtime I run alongside my own leaked one child process per scheduled run and never released them. It hit its process limit in eight days and could no longer fork anything — including its own health check. The container reported unhealthy for four months. Nothing alerted, and I had not looked. The bug was a known upstream issue, fixed in a later release than the one I was running. What I changed: I no longer treat "the container is up" as meaning anything, and I now read health status rather than process status.

2. An auto-healer that plausibly causes the failure it repairs. A watchdog refreshes an OAuth token every five minutes. The token is shared between two services. The scheduled jobs now fail with refresh_token_reused — the error you get when two clients race to rotate the same refresh token. I have not finished confirming this, but a recovery mechanism that creates the fault it is meant to fix is the kind of thing worth writing down before I know the answer. What I changed: nothing yet. Open.

3. A data pipeline died and the dependent agent never noticed. A cron job syncs fitness data every thirty minutes and writes a summary file the agent reads. The sync stopped producing output in April. The cron kept firing on schedule. The agent kept reading the file — it just read four-month-old data, with no signal that it was stale. What I changed: this is the failure that bothers me most, because the agent behaved correctly the entire time. Data freshness has to be part of the data, not assumed.

4. Unbounded logging of fetched content. An agent with web-fetch access logged entire raw HTML pages, tracking scripts included, with no retention limit. Not a breach, but the log became both a disk problem and an unreviewed store of arbitrary third-party content. What I changed: on the list to bound. Not done.

5. A wildcard permission I do not remember granting. Auditing my own config for this repo, I found one agent with allowAgents: ["*"] — permission to invoke any agent — while its sibling is correctly restricted to an empty list. I cannot reconstruct whether that was deliberate. I am leaving it documented rather than quietly fixing it, because "least privilege drifts and you stop being able to explain your own config" is the honest finding, and I only caught it because I sat down to write this.


Public artifacts

Outputs produced through this system:

  • running-plan — a training plan built from synced activity data
  • shop-itinerary — a planning tool built through the same agent workflow

What this is not

This is personal infrastructure, not research. There is no evaluation, no benchmark, no control condition — it is n=1, and "it hasn't broken badly yet" is an anecdote, not evidence. The constraints described here are engineering judgment, not tested guarantees: none of them have been probed adversarially, and I do not claim they would hold against a model actively trying to circumvent them. Parts of the stack are third-party software I don't control. It is also not reproducible from this repo, deliberately — this documents an architecture, it does not ship a configuration.

Open questions

  • The prompt-level policy layer has no technical enforcement behind it. The credential split covers Google; everything else relies on the agent following written rules. I don't know how much that is worth.
  • Alerting is inconsistent. The backup job has a failure trap that messages me when it breaks; nothing else does. Four of the five failure modes above were found by hand, and three had been running broken for months. I knew how to instrument this — I did it once and never extended it to anything else.
  • I don't have a good answer for the wildcard permission in #5, or for how I would have caught it without manually auditing.

License

MIT

About

Documentation and audit of a self-hosted personal LLM agent system — architecture, permission design, and observed failure modes.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors