The Problem With Working Agents

Every agent I built made the same promise: one less thing to remember. After enough of them, the promise inverted. I had scheduled jobs, open issues across several programs, and recurring reviews, and no reliable way to know which of them had quietly stopped moving.

Work rarely fails loudly. An issue sits in progress for weeks. A scheduled job stops running and nothing errors. A task finishes and never gets closed, so the board keeps reporting it as open. Each of these is small. Together they erode trust in the board, and a board nobody trusts is a board nobody reads.

The Scrum Master Model

Human teams solve this with a scrum master. That person runs a short daily standup and asks three questions: what moved, what is stuck, and what is done. The team keeps the work, and the scrum master keeps the picture honest.

Claude and I built the same role for my own work. The overseer runs every weekday morning, reads the state of my open issues and the log of my scheduled agents, and produces a short report. It lists what moved since yesterday, what has gone quiet, what looks finished but was never closed, and which agents ran and which did not. It treats risk, exception, and compliance work as categories of work that need to keep moving, and the report reads the same way for each.

What the Report Looks Like

The report has four headings that repeat every day. Moved lists items that changed since the previous report. Gone quiet lists open items with no recent activity. Finished, still open lists items that appear complete while the record says otherwise. Agent runs lists each scheduled agent and whether it ran. Every heading appears every day, with "None" when there is nothing to list. The fixed layout makes change from day to day easy to see, and it keeps the reading time near a minute.

Each program contributes items to the same four headings. A review that recurs quarterly and a task that recurs daily appear in the same format, and the report treats neither as special.

Report Only

The central design decision was the constraint. The overseer only reports. It never comments on an issue, reassigns one, or changes a status. An agent with write access to stale work can also damage it, and I would have spent my mornings auditing the auditor. A read-only agent has a bounded failure mode: a bad report costs me a minute of reading. That bound lets me run it every day without supervision.

Alternatives We Set Aside

Report-only was one of several designs, and I looked at the others. Write access would let the overseer fix the finished-but-open group directly. It would also put the overseer's mistakes into the record, and every fix would need its own check. Closing items automatically asserts that the work is done, and that assertion belongs to the person who owns the item. Messaging owners would place an agent between me and my team. A dashboard shows current state and waits for a visit, while a scheduled report arrives whether or not I remember to look.

How We Built It

The starting point was one sentence: things are going stale and I am not noticing. Claude and I rewrote that sentence as questions a system could answer from data. How long since this issue changed? Did this job run? Does this finished item still show as open? Each question needed a definition I could state in plain words, and writing those definitions took most of the effort.

We built one report and let its format settle before adding anything. Each later addition came from a moment where I caught myself checking something manually for the second time. A check I had done once did not earn a section. The definitions of quiet and finished came from how I use statuses, and I expect to revisit them as my workflow changes.

Where It Can Go Wrong

The overseer has failure modes of its own. If it cannot read the issue tracker or the run log, the report says so and covers what it could read. A missing report on a weekday is a signal, because the schedule is fixed. A report that grows too long stops getting read, so each section has to earn its place and can be removed later. The report also describes recorded state only. Work that never reaches the tracker does not appear in it, and I supply that context myself.

What I Still Decide

After reading, I decide. A quiet item gets a question to its owner, a new priority, or a closure. A finished item gets closed. An agent that did not run gets investigated. The overseer supplies the picture, and every decision stays with me.

What Changed

I trust the board again. The work of noticing moved from me to the agent, and the decisions I make after reading the report concern priorities. Nothing else about how I work changed; the overseer sits on top of the tools I already use.

The next post covers how we build a new program from a blank page, from a vague need to something running on a cadence.