sg

← Writing

Two Agent Sessions Is What I Can Review

· ai

Coding agents wrote almost all of Fractal's code: 329 commits between March 27 and October 5, in three bursts. Next to that code sits a plans/ directory git never sees, 40 files and about 9,000 lines. Every feature now goes through one of those plans, and I approve each one before it's built. Reading them, and the work that comes back, is the slow part.

Fractal is the infinite-zoom drawing canvas I build on the side, a Go server and a TypeScript client. The split there is the split on all my side projects. Agents research, recommend, and write the code. Most design calls are mine, so every plan waits on me before it's built, and every build comes back to me when it's done. That puts the limit on me, not on the tools: at most two agent sessions per repo, and at most two projects at once.

A feature starts as a plan#

A session usually starts with what I'm thinking of building next, or with something from the backlog. It runs in Claude Code or Codex; I don't have a preference. I tell it to have the architect write the plan, and we go through some grill cycles. In a grill cycle the agent asks me one question at a time about the plan, each with its recommended answer, and the plan is amended from what we settle. Once I'm happy with the plan, I tell it to orchestrate.

The repo's rules hold that shape. plan-before-code says I approve the plan before implementation starts, and "Engineers execute the plan; they don't author it." The orchestrate skill's first rule is "You orchestrate; you don't implement." The session hands each phase to an engineer agent and then a code-reviewer agent. If the fixes haven't converged after three rounds, it commits what's reviewed and verified and escalates the rest to me. Plans that are hard to reverse also get autonomous grill rounds and one review from the other CLI before I approve them.

Two sessions per repo, two projects#

I run at most two agent sessions in a repo. I prefer to have context on what's happening, so almost always only one plan is being built in a repo at a time. When there are two sessions, one is the planning lane, where I'm working on the plan that gets orchestrated next. The other is the build lane, orchestrating the current plan, which I've finalized and approved.

A second build lane would mean two diffs to read against two plans.

A build lane stays cheap to review when its bounds are written down before it starts. The second performance plan records that I authorized "this complete bounded performance run and autonomous implementation." Right below that, it says to promote experiments only on measured wins, and rules out a persistence schema migration, a worker redesign, new dependencies, and a general cache or transport framework.

Across projects the limit is also two. Anything more turns into a mess. It becomes almost impossible to give attention to anything, and the context switching gets to be too much.

What waits for me#

The zoom work shows where a plan stops until I decide. On a separate branch, agents had built a harness to compare two candidate renderers. Then the comparison went on hold. Its plan says "Global chronological paint/erase/repaint ordering is a pending user decision; do not infer it from this confirmed cross-depth requirement." The answer is in the parent plan under "Required behavior (user decision, 2026-09-30)": "Paint order is global and chronological at every depth, for paint over paint and for erasing," and "Later paint covers the hole, including paint authored at a shallower zoom." The eraser post covers what that means on the canvas.

The qualification plan that followed set its gate as "Every target is met or has a recorded user decision." A device check on an iPad was a merge gate until October 1, when I recorded that I had no iPad on hand. The work merged on the browser checks, and the device checklist is still waiting to run.

Rules that keep diffs from growing#

Several of the rules work to keep what comes back to me small enough to read. In scope-discipline, a finding can block only with a trigger the system actually receives, an observable consequence, and "path:line evidence plus a failing test or a precise reproduction." Anything short of that gets recorded, not fixed. Approved plans are frozen during implementation, and "A fix round should leave the diff the same size or smaller."

The orchestrate skill says the same thing about its own loop: "A fix round that grows the diff is a self-expanding loop: stop and cut back to the plan." The autonomous grill skill applies it to plans: "Amendments clarify; they don't grow." If a plan keeps getting longer round over round, the loop stops and reports that.

The failure I'm guarding against#

This doesn't usually happen, because I'm reviewing the plans and the work. It could, if I got overly ambitious: four agents running in parallel, and then merging everything because there was too much to review. Fractal has had at least one unread diff: a May commit whose whole message is "agents", and the rate limiter post covers what was in it.

The same loop elsewhere#

The same habits carry into my day job at Tag-N-Trac. This blog runs the loop too: a writer agent drafts, an editor agent reviews, and "Shipping is the author's call."

What the cap costs#

The cap gives up throughput. I could run more sessions; I don't. The performance work took almost a week, and the eraser plan took two days.

Two is a number about me. The tools would take a third session without complaint, and its plan and diffs would come back to the same reader.