The Agent Graph Was the Easy Part
I got fast at asking a coding agent to build the thing and waiting. The workflow I actually needed was the part that still exists after I close the chat.
I have a very fast loop now. Ask Cursor, or Claude, or Codex, to build the thing. Wait. Inspect the diff. If it looks plausible, keep going. If it doesn’t, ask again.
I got good at that loop. A little too good.
Last post I opened a different coding agent and sat there, actually stuck, on something that costs me zero seconds in Claude Code. Not a bug. My hands were waiting for a permission prompt that wasn’t coming. The tools are useful because they are good. The problem is I had trained myself on one environment so hard I couldn’t tell “I understand this work” from “I know where the buttons are.”
I do not want to use coding agents less. I want to keep using them a lot. I want to close the chat and still know what I decided, what actually got checked, which model did which job, and how to do the same work next week in a different harness.
A workflow that only exists inside one conversation is not a workflow. It is a session I got attached to.
Boxes are cheap
I drew the obvious graph anyway, because drawing it felt like progress.
plan → build → verify → repair
Cute. Anthropic has a public post full of graphs like that, chaining and routing and an evaluator in a loop, and I nodded along like I had found the missing infrastructure. Then I tried to make it reliable and every interesting question lived outside the boxes. Who owns the state when the chat dies. Whether the reviewer is allowed to see the builder’s little speech about why the tests are fine, actually. How many retries I was pretending were bounded. Which pieces I was about to rebuild even though git already has worktrees and the repo already has a test command.
The graph was the easy part. The easy part is always the diagram.
I almost built a tiny platform about it
This is the embarrassing stretch.
I went deep. Firstmate. no-mistakes. OpenHands. Anthropic’s multi-agent research system. Codex. Aider. Observability. Evals. Routing. Skills. Durable state. At some point I was designing a dashboard. A memory layer. A skill install list justified by how many other people had installed the skills. You can smell it. I was one README away from a product nobody asked for, including me.
The thing that saved it from getting fancier was verification that was not me.
I ran my own plan through an independent council. Different models, different jobs, none of them sitting in the conversation where I had talked myself into it. It came back repair_required. Not “maybe tighten the prose.” The graph contradicted the menu. BEST-OF-N was allowed to pick a winner before the cheap checks had run. Retries were bounded per node in the way a diet is bounded if you only count lunch. I had used install counts as evidence in the same document that told me not to.
I almost did it with a number I liked, too. Anthropic reported a 90.2% gain for a multi-agent research system on an internal eval, and in the same post they said most coding tasks don’t parallelize like research. I had already started quoting the 90.2. The council took that one too. Fair.
A verifier with a different job and a different context will catch things I will not catch by asking the same session to rethink itself. Claude Code already isolates subagents from the parent chat. Cognition has been writing about this in public: long traces rot, and two agents making unstated decisions will hand you a Flappy Bird with a Mario background. I did not invent it. I just finally believed it because it happened to my own plan.
Do not pay a model to notice a type error
The split that actually matters is boring, which is how I know it’s real.
The compiler already knows about lint, types, tests, the build, a leaked secret. A language model should not be the first to discover any of that. Spend the expensive review on whether the design matches what I asked for, and whether the builder smuggled in an assumption the tests cannot see.
Same instinct, one layer up. I do not want a model keeping count of retry 2 versus retry 7. I do not want a Python script deciding whether an architectural assumption is sensible. Scripts for mechanics. Agents for judgment. Firstmate writes that down like a house rule, and I stole it.
Phoenix already exists. So does gitleaks. So do git worktrees. So does Obsidian, if what I need is a place I can study. I kept catching myself about to build the cute version. Stars are not a design argument. I know that. I still almost used them as one.
The chat is a terrible hard drive
The moment that made this stop being theoretical: a long session sitting at about 225K of 256K tokens, and something like 193K of it was just us talking. Not files. Not decisions. Conversation.
I did not discover that long chats get weird. Everyone has felt a session go a little drunk after hour three. The part that got me was watching the project’s brain live in the transcript, then watching the next turn pay rent on all of it.
Chroma has a whole report about models getting worse as you stuff more into the window, even on simple tasks. Focused little prompts beating giant histories. It is a technical report, not a coding-agent bible, and the models in it are already old. Still: long context is not the same thing as useful context. I had been asking how agents could remember more. The better question is how little the next one needs.
I used to call Obsidian “memory,” which is how you can tell I had not thought about it. The vault is for me. I can follow links, argue with myself, keep sources. It is not a soup you pour into a fresh Claude window. The unfinished part is the interesting one: pick the few files this step actually needs, say so out loud, start clean.
Pass decisions forward. Not the conversation.
Smaller on purpose
I do not want an AI manager personality sitting on top of my work. I want a thin decision I can read.
RUN PLAN
shape: NORMAL
planner → strong-reasoning
builder → strong-coding
reviewer → cross-family-strong
gates: lint → typecheck → tests → build
Claude does not “own planning.” GPT does not “own review.” Those are roles I am mapping onto whatever is good this month, and I want to be allowed to change my mind when the runs say so. “Observed across N of my runs” is a finding. “This model is the architecture one” is a group-project identity, and I have had enough of those.
The next real test is parallel work, which is where I will find out if any of this survives a repo. While I was doing the research, multiple subagents hit resource exhaustion writing into the same archive. Roommates. One fridge. Cognition’s polite version is two agents building one game. I got the version where the notes folder filled up.
I am less interested now in the most sophisticated agent graph I can imagine. I want a system I can inspect after the chat is gone.
I do not want to become less dependent on coding agents. I want to become less dependent on using them blindly.
Amisha
Filed under Agents · AI Systems