Rendered from docs/process/evaluations/the-build-order-as-a-multi-agent-system.md in the Headwater
corpus. Every document on this half of the site is typed by the taxonomy
the descriptor names: corpus.json.
The build order as a multi-agent system
The build order is the loop that turns a session into an orchestrator, picks issues, dispatches agents, and merges what they build. .claude/commands/next-run.md carried its whole design as prose for eleven runs, and the prose was read by the one agent that could not be reloaded. This evaluation records what one run measured about that shape, the cost model the measurements settle, and the architecture that follows. It also records what was rejected on the way. Six decision records carry the rulings, and this document carries the evidence they cite.
The reader outside this repository is an adopter who runs an agentic loop over a Headwater corpus of their own. Every number here was taken on this repository, and the cost model is stated so that another corpus can take its own.
How the numbers were taken
A session log writes one API response as several lines, one per content block, and every line carries a copy of the same usage record. A count per line inflates turns and cost by about 2.3 times, and it inflates them unevenly. That is because a response with four tool calls is copied more often than one with a single call. Every figure here groups by message.id and takes one usage record per identifier. The orchestrator of the run below reads as 712 turns and $156 per line, and as 304 turns and $67 per message.
The unit that explains the cost is carry. A tool result costs its own size times the number of turns that follow it. Seven thousand tokens read at turn 20 of a 170-turn agent carry 1.05 million token-reads. The same seven thousand read at turn 150 carry 140 thousand. Position is worth as much as size, which is why cheap-looking calls dominate. In the single-agent shape, Bash results were 68% of all carry and file reads 29%, because an iteration made 142 Bash calls against eleven reads.
What the single-agent shape cost, and what the four-agent split bought
Across 80 single-agent iterations, 77% of the cost was cache reads, the context re-read on every turn. Output was 12% and cache writes 11%. The median iteration ran 170 turns, peaked at 286k of context, and re-read 32 million tokens on the way.
Splitting one iteration into adjudication, construction, verification and write-back bought peak context and did not buy dollars. Adjudication at $3 and construction at $17 came to $20 against $21 for the single-agent shape. That is a 4% saving over a ten-hour run that merged seven pull requests. Peak context per agent fell from 286k to 211k, and adjudication alone ran at 116k and $3. That is because settling a premise never needs the test output. Construction grew back to 196 turns against the old 170, because the work did not shrink when it moved.
| Shape | Turns | Peak context | Cache read | Cost |
|---|---|---|---|---|
| One agent, median of 80 | 170 | 286k | 32.1M | $21 |
| Adjudication, median of 7 | 37 | 116k | 2.3M | $3 |
| Construction, median of 8 | 196 | 211k | 28.4M | $17 |
Run 22 measured the shape with verification split out. The parent was 9% of the run, against 27% before, and cache reads were 99% of every token billed. What did not move was growth. The parent climbed from 65k to 687k of context in under three hours with no compaction. Its last 84 turns cost about twice its first 84 for the same work. The built-in tool definitions were 28,987 tokens of that context, and the parent called four of the fourteen tools. Returned reports were its largest input class at 113.7k tokens, with a median of 1,475 each. Three agents notified two or three times, re-injecting the whole report on each.
A builder's late turns cost more than its early turns, and the length of the builder causes that, not the size of its issue. The growth line of tools/run/run-census.sh gives the context at 10% and at 90% of the turns of one transcript. It also gives the cost of the first quarter of the turns against the last quarter, at the rates of each turn's own model. On 2026-09-27 it read all 120 hw-build transcripts of 13 sessions. These include the baseline session 9ab3be93 of 2026-09-20 and eleven sessions after it. The size of an issue is the additions plus the deletions of the builder's own pull request. It leaves out the derived paths .headwater/, site/ and the recorded corpus fixtures. 115 of the 120 builders named a pull request.
| Builders | Count | Median turns | Median growth | Mean growth |
|---|---|---|---|---|
Baseline 9ab3be93, Sonnet 5 |
3 | 582 | 4.2x | 4.0x |
| Other Sonnet 5 builders, before the Opus 5.5 builders | 26 | 190 | 2.8x | 2.8x |
| Opus 5.5 builders, from 2026-09-23 on | 91 | 109 | 1.4x | 1.6x |
| 100 turns or fewer | 43 | 72 | 1.1x | 1.2x |
| 101 to 200 turns | 55 | 140 | 1.9x | 2.0x |
| 201 to 300 turns | 11 | 244 | 2.4x | 2.6x |
| More than 300 turns | 11 | 369 | 3.5x | 3.5x |
Across the 115 builders with a size, the rank correlation of growth with turns is 0.77, and of growth with size it is 0.28. Turns and size correlate at 0.49. With size held constant, turns still correlate with growth at 0.75. With turns held constant, size correlates with growth at -0.16. For the 86 Opus 5.5 builders with a size, the same two figures are 0.58 and -0.01. So a large issue costs more per turn only because it makes a long builder. The issue that asked the question measured #88 at 392 turns while it ran, and that builder stopped at 582.
No bound on the length of a builder is set, because the most that a bound can save is small against what a handoff risks. The code map and hw-explore went into .claude/agents/hw-build.md on 2026-09-23. Since then, 8 of 94 builders in eight sessions went past 200 turns, and 1 went past 300. A fresh builder that takes over at turn 200 saves the cost of the later turns. It pays again for the same number of turns at its own start, and those turns include its orientation. Over the 94 builders, a bound at 200 turns saves at most $22.40 of $459.21, or 4.9%. A bound at 150 turns saves at most $47.33, or 10.3%, and it stops 25 of the 94 builders. The estimate charges nothing for the brief or for work that the new agent does again. The run policy resumes a dead agent and does not replace it, because a fresh agent loses a settled design. HW-PD-0003 asks a dispatch to retire more than it costs. A saving of 5% to 10%, before those two costs, does not meet that test with a margin.
The finding opens again on the last five runs taken together, and one run alone is too small for the test. It opens when more than one builder in five goes past 200 turns, or when the median growth of those builders goes past 2x. On 2026-09-27, the last five runs, from 861a7465 to ee0f0553, had 6 of 59 builders past 200 turns and a median growth of 1.5x. Over those five runs, a bound at 200 turns saves at most $21.28 of $317.28, or 6.7%. One in five is about twice the share of today. At that share, a bound at 200 turns comes near what a bound at 150 turns saves today. That is 10.3% since 2026-09-23, and 11.8% over the five pooled runs. Of the eight single runs since 2026-09-23, e2a9988f passes the growth limit with a median of 2.2x. 6fadc105 sits on the share limit with 2 of 10 builders past 200 turns. One builder moves the test on one run, and that is why the test pools five. The growth line gives both figures for each builder.
The run that measured the orchestrator as the constraint
Session 19108df3 ran for 20 hours and 29 minutes, dispatched 152 agents, and touched 45 pull requests. The human sent eight prompts in that time. The measurements below were taken from its session log after the run.
The fleet ran at a third of its target width. Mean in-flight concurrency was 3.03 against a target of eight. No agent at all was in flight for 20.9% of the active window, which is 3.4 hours across four windows. The largest window was 114 minutes. In each of those windows the parent was merging, building or keeping records by itself.
The parent was the mutex. Of the parent's own serial shell time, 61% contended on four shared resources: the one main checkout, the one engine/target, origin/main, and the post-merge regenerate. cargo build alone was 49% of it, at 67 minutes over 26 calls. The merge itself was cheap, at about eleven seconds over 46 merges. Only 13% of that time was per-job work, and most of that was polling other agents' pull requests.
The batches had a barrier by accident. Subagent runtime had a median of 13.5 minutes, a ninetieth percentile of 46, and a longest of 91. Across the batches of three agents or more, 9.3 hours separated the first completion from the last. No code wrote that barrier. The parent advanced a batch when its slowest member returned.
Merging was not the constraint. Of the 45 pull requests, 44 merged at a median of 40 minutes from opening to merge. The runner queued nothing that showed as idle workers.
Width eight measured worse than width five, and the comparison is confounded. The width was raised from five to eight in the middle of the run. Mean concurrency fell from 4.00 to 3.22, and the share of the window with an idle fleet rose from none to 32.1%. The owner of that run attributes the largest single cause to merging in ready order rather than in artifact-footprint order. So one branch touching every corpus-wide derived artifact was overtaken four times and re-derived four times. The comparison stands as a caution and not as a finding, and the next measurement takes it again under footprint order.
The parent compacted five times. Its single largest gap, 82.8 minutes, ended at a compaction. The doctrine it ran under was 67 KB of prose in its own context. A compaction summarizes prose of that length rather than preserving it.
Of 67 messages the parent sent, two kinds want opposite treatment. Most were authority: a branch sent back with a defect named, or a delta sent for re-verification. A few were coordination: a warning to one worker that another was about to touch the same derived artifact. The first kind is the parent exercising a veto only it may exercise. The second kind is a shared-state problem the parent was relaying by hand.
The finding that set the unit of cost
Pull request #685 measured the same parent's polling. It made 76 gh pr view and gh pr list calls whose combined output was 25 KB. Those calls occupied 77 turns and 21.3 million cache reads, which is 10.3% of the parent's entire 207 million for the run. Those 77 turns and their cache reads carried about 6.3 thousand tokens of payload. That is roughly 3,400 times the payload in carrier cost.
The unit of cost is therefore a parent turn at the parent's full context, and not a token. The same finding rules out the obvious remedy: a fresh agent dispatched to make one poll is worse than the poll. That is because the dispatch is itself a parent turn at the same context, with a prompt on top. A dispatch pays only when it retires more parent turns than it costs. HW-PD-0003 is that rule.
Counting repeated calls does not find this class. The 76 calls covered 43 pull requests at about one and a half calls each, and no shape repeated more than 40 times. An earlier incident, 232 identical calls in twelve minutes, was visible from any angle. This one was visible only by grouping shell calls by purpose and reading the total.
Session 9ab3be93 measured a second class at the same unit of cost, a turn at full context. This time it was paid on the way out of a wait rather than on the way through a poll. The prompt cache holds a copy of a subagent's context for about five minutes. Thirteen turns across three builders read a background-task notification after their own wait had run longer than that. Each of those turns paid to write the whole context back rather than to read it, at 1.25 times the input rate against 0.1 times. The three builders rewrote 3.72M, 1.48M and 0.74M tokens across six, five and two such gaps each, $14.85 of the run's $130.73 total, 11.4%. A wait that came back inside the cache lifetime instead would have cost about $3.84 in cache reads over roughly 32 extra turns. HW-PD-0007 bounds a background wait under that lifetime and re-issues it on return. tools/run/run-census.sh reports the class directly, so the next run is held to a number rather than to this record.
What the following run added
Run 24 ran under the integrator slot and with polling moved out of the parent. Its owner reported four things this evaluation carries forward.
Refusal was the highest-value output. Six of eight slots returned something other than "build as specified". Three rejected part of their own Done-when. One found a pre-existing defect that permanently bricks an output path, reproducible at one run in three with its own change stashed. The owner's reading is that this happens only when adjudication is separate from construction and refusal is licensed in the prompt. That is because an agent handed a prescribed remedy implements it and cannot find the error in it. HW-PD-0002 rules on it.
Widest artifact footprint merges first. Under that rule no branch was overtaken. Its corollary is to never add rebase work to a branch that is still running. That is because a finished branch waiting costs nothing, and a working branch redoing its derives costs a full pass.
A blocking wait outlives the agent that started it. The run ended with 30 orphaned until … sleep 30 loops, the oldest still polling after eleven hours, emitting phantom notifications and pinning deleted worktrees on disk. One of them was still alive nine hours after the run stopped.
A stale instrument reports a correct tree as defective. A binary four hours behind main reported twelve site figures stale. The gate's own printed remedy would have written the wrong values into three pages and staged them. Two sessions read that report and both concluded the pages were stale. Rebuilt at the same commit, the check reported zero stale. The direction of every delta was the tell: a measurement missing something the page knows about is an old binary, not a stale page. Issue #679 holds it as a correctness root.
The cost model
Three quantities move independently, and a change serves at most two of them.
- Fewer parent turns. Served by work that never returns to the parent, by state the parent never reads, and by making every unavoidable turn do everything it can. A dispatch pays only when it retires more turns than it costs.
- Smaller parent context. Served by moving prose out of the command, because every turn re-reads the whole of it.
- Smaller subagent context. Served by trimming what every dispatched agent loads.
CLAUDE.mdat 6,500 tokens loads into each of 150 agents per run, and that is a larger lever than any per-agent model choice.
Decomposing the command into agent definitions serves the second quantity only. It does nothing for the first, and a design that claimed otherwise was corrected by the #685 measurement.
The architecture that follows
HW-PD-0001 states the ownership rule, which decides where any sentence of orchestration prose lives. Under it the command shrinks to a doctrine block of ten numbered lines, a loop, a veto and a dispatch template. Each stage becomes an agent definition with its model in frontmatter. Rules two or more stages obey become skills. Measurement and rationale become this evaluation and the records it cites.
The stages are adjudication, construction, verification and integration, and integration fuses the merge, the regenerate and the write-back. That is because all three are mechanical and all three touch the one shared checkout. A fifth definition writes the queue. The parent takes four turns per issue, which is one more than the fused shape would take, and HW-PD-0002 records why that turn is paid.
HW-PD-0004 moves coordination to create-only claims in a run directory and keeps the veto with the parent. HW-PD-0005 splits the ledger so that its doctrine reads alone and its log reads by the line. HW-PD-0006 keeps /next-run and makes a run resumable from that directory.
What was rejected
Fusing adjudication into construction would retire one parent turn per issue. Run 24's refusals are the evidence against it, and the turn is paid.
A fresh agent per poll fails the dispatch rule. The poll's output is one line, and the dispatch costs a full turn plus a prompt.
Peer-to-peer authority was rejected on a live demonstration. One session declined to edit CLAUDE.md on another session's reasoning, because a peer's message carries no owner authority, and that refusal was correct. In a mesh, every agent adjudicates provenance on every message. On a tree, authority flows down and provenance is never in question.
Renaming the entrypoint was rejected because the name was not what was wrong. The command was the orchestrator's whole brain and its only memory, and a resumable run fixes that where a rename does not.
Re-reading the doctrine every iteration was rejected because a re-read is itself a parent turn and it grows every later turn's context. The doctrine is the header of the ledger instead, and the parent reads it on a turn it already pays.
Per-worktree engine builds were rejected. A fresh worktree has no engine, so the commit gate fails open there, and sccache bakes a dead worktree's path into a test binary. One warm target directory behind the integrator beats several cold ones.
SQLite for the ledger is deferred and not refused. The problem it solves, reading a header without the body, is solved by splitting the file and writing the tabular parts as JSONL. A database earns its place when a cross-run question is asked, and JSONL imports cleanly then.
A headless loop, where a script drives each stage and the session spends turns only on verdicts, is the end state this design points at. This design does not take that state. Of the 45 pull requests in the measured run, 44 merged under the veto, and nothing has measured what merges without one.
The numbers the next run is held against
| Measure | Baseline by hand | Baseline by the tool | Target | New shape, by the tool |
|---|---|---|---|---|
| Parent turns per issue | not measured as such | 19.6, as 880 turns over 45 pull requests | 4 | 24.5, as 98 turns over 4 pull requests |
| Share of the window with no agent in flight | 20.9% | 0%, largest window 0 min | near zero | 0.3%, largest window 0.4 min |
| Mean in-flight concurrency at width 5 | 4.00 | 9.54 over the whole run | at or above 4.00 | 3.91 over the whole run |
| Largest gap after a compaction | 82.8 min | 22.3 min after, 87.6 min before | under the dispatch cost | no compaction ran |
Parent gh calls |
92 | 88 by leading verb | 0 | 2 by leading verb |
cargo build in the parent |
26 calls, 67 min | 8 by leading verb, 15 mentioned | 0 | 0 |
| First-to-last completion spread per batch | 9.3 h total | not taken | not applicable, no batches | 5 min over five adjudications, 1.95 h over five constructions |
tools/run/run-census.sh takes these from a session log and the agent transcripts beside it. The run that follows this design writes its numbers into this table.
The figures in the table were taken by hand, part-way through the run, and the tool was written after them. Over the whole transcript of that run the tool reports 880 turns and 227.3 million cache reads. Mentions of gh pr view cost 66 calls, 66 turns and 18.1 million cache reads, which is 8.0% of the run. Mentions of gh pr list cost 24 calls, 24 turns and 6.7 million, which is 2.9%. The tool counts a call once per turn and a turn once per message, and it reads a verb past a leading cd or set -e. That is because that run wrote nearly every command in that shape. The next run is compared with numbers the same tool takes, and not with the hand count.
The fleet section of the same tool reads the agent transcripts that the harness writes beside the session file. Over the whole run it does not reproduce the hand count. It finds 165 agents at depth one over a span of 20 hours. No window had nothing in flight, and 9.54 agents were in flight on average. The hand count found 20.9% idle and a mean between 3.03 and 4.00. The tool counts an agent as in flight from its first line to its last turn, so every minute inside a blocking wait counts. The hand count was taken over a part of the run, by a method this evaluation does not record. The tool's reading stands, because the next run's reading is taken the same way. The largest gap before a compaction was 87.6 minutes from the parent's last turn. The largest gap after one was 22.3 minutes to its next dispatch. The parent took 5.5 turns per agent over 161 agents of one type, and 19.6 turns per pull request over 45.
What the first run of the new shape measured
Session 8e38c6e8 ran the ledger's RUN 20260907-2208 under the five-agent shape, for two hours and thirty-six minutes, and merged four pull requests. The parent context grew from 55k to 269k tokens over 98 turns, with no compaction. It made two gh calls and no cargo build call, against 88 and 8 for the baseline run. Idle share fell to 0.3%, with a largest gap of 0.4 minutes, against 20.9% before.
Mean concurrency read 3.91 over the whole span, just under the 4.00 target. The reading understates the busy middle of the run. Five construction agents started within eight minutes of each other and finished across a span of 117 minutes, from 28 to 145. That spread thins the fleet after minute 110, once the rest of the run's work had already finished. Over minutes 20 to 100, the tool reads a mean of 4.9 agents in flight and a peak of 6. The whole-span figure is the one this table tracks, and it sits close to the target. The busy-core figure says the design sustains more parallel work than the whole-span number alone would suggest. A staggered batch of long, uneven construction runs narrows the gap between the two figures.
Refusal held on one issue across three rounds. Adjudication refused issue #648 outright, on the ground that two already-merged pull requests had answered its Done-when clause. Verification failed issue #485 twice, on two distinct construction defects. The doctrine's two-fail stop condition held the branch rather than dispatch a third round. A second-opinion agent, dispatched outside the normal construction path, reproduced the second defect for real. It proposed the fix that the third construction round then applied and verification confirmed. A fourth round did not run, for a different reason. The only regression test for that fix does not run in the continuous-integration job. A merge on a green check would repeat the pattern this evaluation names for the corpus tree in general, at line 78. The run left the pull request open for the owner's ruling instead.
One risk surfaced outside the measured numbers. A construction agent for issue #603 proposed a force push, against the standing rule that forbids one. The parent allowed it, after checking by hand that the branch and main were both intact. The rule held on the parent's judgment this round, and not on a check inside the construction agent's own prompt. The construction agent's prompt needs that rule stated, before the next occurrence depends on the same judgment again.
What two later runs added
Four more measurements spend the same unit of cost, a turn at full context. Each one set a rule that the stages now carry without the number.
A parent that checks pays for a turn that learns nothing. The parent of run 20260911-1331 ran a bare true 193 times, and those turns cost 55% of the run. The parent of run 20260920-2058 grew its context from 62K to 721K tokens over 424 turns, with a mean of 409K. Of its 239 turns that did work, 126 were checks on a background agent that found no change. Those turns had a mean context of 462K tokens and read 58M tokens from the cache in total. The parent's cache holds for an hour, so a check keeps nothing warm that the next report would not find warm. Doctrine line 2 of .claude/commands/next-run.md is the rule: with agents in flight, the parent ends its turn.
A builder that sleeps and checks is a poll with a delay. One builder in run 20260920-2058 issued 21 separate turns in the shape sleep 5 && echo ok while it waited on its own CI run. Those turns read 5.9M tokens from the cache. One blocking call to tools/run/wait-for.sh does the same wait in one turn, and the run policy names that form.
A builder that runs the whole suite on each edit pays for the workspace on each turn. In the same run, one builder ran cargo test --workspace eleven times over its own build. Another builder ran cargo directly thirteen times, outside the slots of tools/hw-cargo. .claude/agents/hw-build.md now scopes a test run to one crate while the builder iterates, and it keeps one workspace run before the pull request.
A long report costs the parent on every later turn. In one run, forty of forty-one stage reports were two to four times the 400-token limit. The parent read each one again on every later turn. Each stage now writes its narrative to a file and returns only the fixed block that its definition names.
What the first run with the loop below the parent measured
HW-PD-0022 moved the verify and rework loop for one issue into hw-iterate. It set a target of 3 parent wakes or fewer for each merged issue. Run 20260928-1109 is the first run that measured that number.
The method. A wake is a top-level user turn of the parent transcript that carries a <task-notification>. The count reads one parent session, 2ecbf66e, from 2026-09-28T11:09Z to 2026-09-29T02:35Z, with no restart. A merged issue is an issue that a merge line of the run's log.jsonl names. In that span the run merged 24 pull requests for 23 issues, and one more pull request, #1351, named no issue. HW-PD-0022 counts three kinds of wake: the adjudicate report, the hw-iterate report and the integrator report. The count gives each adjudicate wake and each hw-iterate wake to the issue that its summary names. It gives every integrator wake to the merged issues, because an integrator runs only to merge them. Two scripts outside the tree took this count, and tools/run/run-census.sh does not count wakes of this kind yet. So this count is a hand count in the sense of the section on the numbers the next run is held against.
The number. The merged issues took 17 adjudicate wakes, 29 hw-iterate wakes and 14 integrator wakes. That is 60 wakes, or 2.61 for each of the 23 merged issues. Six of the merged issues had their adjudication in an earlier session. With one more wake for each of those six, the count is 66 wakes, or 2.87 for each merged issue. Both numbers are below the target of 3.
All the wakes. The parent woke 90 times in the span, which is 3.6 for each of the 25 merged pull requests. 30 wakes are outside the count. Of these, 10 adjudicate wakes and 7 hw-iterate wakes were for issues that had not merged yet. The rest were 5 product owner passes, 3 release verifications and 5 single tasks. The owner also wrote 15 turns, which are not stage wakes.
The baseline. HW-PD-0022 gives about 8 wakes for each merged issue in run 20260927-0443, and it does not state how it counted. The same method on session b5554ef1 of that run gives 626 wakes for 52 merged pull requests, which is 12.0 for each. Those pull requests closed 44 issues, so the count is 14.2 for each merged issue. The old shape also woke the parent for each build and each verify. So compare this baseline with the 3.6 of all wakes above, and not with the count of HW-PD-0022.
A live child does not wake its parent. HW-PD-0021 left one question open: whether hw-iterate wakes the parent when it ends its turn while its builder or verifier runs. In this run, 8 issues gave the parent more than one hw-iterate wake. Each repeat wake came after a SendMessage from the parent to that agent, or after a new hw-iterate dispatch for the issue. The messages were rulings on a STOP, a disk that the parent freed, and conflicts after a merge. No repeat wake came while the parent was silent. So the ended turns of hw-iterate cost the parent nothing in this run. The count matches each agent to its issue by the summary of the notification and the time of the send. It does not use the agent id.
What this evaluation cannot show
More than one run measured the orchestrator, and each section above names the run it reads. No two of those runs had the same shape, so no number above is a trend. One run reported the refusals. The width comparison is confounded and is recorded as such. The claim that a compaction preserves a short numbered list and not a long narrative is plausible and untested. The doctrine sizing rests on it. The integrator's net saving of about five parent turns per merge is arithmetic from the measured turn classes. It is not a measurement of the new shape. The saving that a bound on the length of a builder can make is also arithmetic, because no run has tried a planned handoff. Each of these is a thing the next measurement can settle.