Rendered from docs/spec/16-harness-support.md in the Headwater corpus. Every document on this half of the site is typed by the taxonomy the descriptor names: corpus.json.

Harness support

Spec 5 names the four moments at which a corpus serves an agent. The hook contract states what a position passes, what it returns, and what it binds, which is nothing. Both are written against this engine and against no harness. What a particular harness offers at each moment is a fact about that harness, and it moves on the vendor's cadence. This part holds that record in one place, so that a change in a harness moves one table here and no sentence of spec 5.

The reader is a person who binds this engine to a harness, including a harness this repository does not run. The capability vocabulary is the half that holds still. The support table is the half that a vendor moves.

What a harness is

A harness is the program that runs an agent against a checkout. It brokers three things, and every capability below is a claim about one of them. It decides what context the model reads without asking. It hands control to configured scripts at fixed positions in its loop, and takes it back. It keeps the register of what a person, or the model, invokes by name.

A harness is not one program per vendor. "Copilot" names an IDE integration, a terminal program, a hosted coding agent, a code reviewer and a completion engine. Each of the five supports a different subset of the table. A support claim therefore names a surface, and a claim that names only a vendor compresses real variance. The table below carries one column per vendor and states, in the notes, where the surfaces of one vendor disagree.

The eleven capabilities

Eleven capabilities cover the integration this engine asks for. Three carry context to the model, four intercept the loop, three are invoked by name, and one is a protocol rather than a harness feature. Each entry states which moment of spec 5 it serves, and the terms the harness must meet. The terms are the hook contract's vocabulary: what the harness passes, and what it does with what comes back.

Context: standing, scoped, injected

C1 — standing context. A file the harness reads whole into every session, with no action from the model. It carries the conventions and the one directive that names a skill before any description competes. The terms: the file lives in the checkout, and the harness reads it at session start. It costs its size on every request, which is why spec 5 prices it against the subagent.

C2 — scoped context. Instructions attached only when work touches a matching path, which is read-time rule loading. The terms: the checkout declares which file covers which paths, and the harness attaches on match. A harness offers this as a glob a rule file declares, or as a file per directory. A projection follows whichever shape the harness reads. Either way the attached files are generated projections of canonical documents under a declared size budget, which spec 5 requires. None of them is a second authored copy of a rule.

C3 — injected context. A position that runs when a prompt is submitted, whose output joins the model's context before the model acts. It serves intent-time routing. The terms: the harness passes the prompt as text, and it returns standard output to the model verbatim. Silence is a result that costs nothing, because a wrong pointer costs more than a missing one.

Interception: refusal, advisory, gate, read advisory

C4 — write refusal. A position before a tool call that can deny the call with a reason the agent reads. It serves backfill at the write moment: a raw write of a new document is refused, and the refusal names headwater new. The terms: the harness passes the tool name and the one path, and a denial carries prose. The harness gives that prose to the agent rather than to a log.

C5 — write advisory. A position before a tool call that edits a file, which adds context and blocks nothing. It serves impact detection. It comes before the edit, so the agent reads the governing set while the change is still a plan. The same position names a path that the governed scope admits and that no document governs. It then prints the front-matter lines that declare the edge, and it writes nothing. This needs no second binding, because headwater route answers both cases in one call. The terms: the harness passes the tool name and the one path, and it gives the added context to the agent without a decision. The posture is the substance: spec 5 keeps it advisory, because a blocking gate trains the reflex answer that destroys the signal.

C5 also names a derived artifact, which is a file that a producer writes. It has two moments. At the edit moment, the PreToolUse position on Write|Edit runs write.sh. That script adds a part that names the row and the treatment of the record, and the command that rebuilds the file. At the conflict moment, the PostToolUse position on Bash runs derived.sh. A merge, a rebase, a cherry-pick or a revert that stops on a conflict is a shell call. So the session hears of the conflict on that call, and it names each unmerged file that a producer writes. Both positions read headwater derived through one reader in lib.sh. The conflict position is bound on Claude Code alone. For Copilot and Codex, the name of the shell tool on a PostToolUse payload is not measured. The first cost measurement is from 2026-09-27, with the dev-release engine, over six runs at a load average near 28. A shell call with no merge cost 8 to 9 ms. On that call the position reads the payload with the engine and then tests the merge state. It runs headwater derived only inside a stopped merge. A shell call inside a stopped merge cost 20 to 21 ms. One run of headwater derived over this corpus cost 32 ms, and each write.sh call pays that once. That call cost 365 to 440 ms at the same load.

C11 — read advisory. A position before a tool call that reads a file, which adds context and blocks nothing. It serves the read moment with the pointers that C5 names. A session then hears the governing set when it opens a file, and not only when it changes one. It is not C2: C2 attaches a generated rule file on a glob, and C11 is a hook position that calls headwater route. The terms are the terms of C5, and the tool is the read tool of the harness. It has the number 11 so that no earlier number moves.

The cost of C11 is paid on every read. .claude/hooks/fixtures-live.sh times the whole hook with no harness. The first measurement is from 2026-09-24, on this corpus, with the dev-release engine and a warm cache. One read cost 69 to 78 ms, over six runs on each of three paths. On 2026-09-26 one read of .claude/hooks/write.sh cost 232 to 237 ms over six runs, on a host at a load average near 60. The hook before #953, at the same load, cost 197 to 207 ms. The difference is the pointers that the hook now reads from route --json one member at a time.

C6 — turn gate. A position at the end of a turn that can refuse the end and hand the agent a report. It serves the review moment. A binding calls the commit gate itself, so what stops a turn and what stops a commit stay one file. The terms: a blocking verdict that carries the report, and a flag that says the gate already blocked this turn. A gate with no such flag is a loop.

Invocation: commands, skills, agents

C7 — commands. A procedure a person invokes by name. The workflow commands of this repository are three: the build-order iteration, the stacked run, and the board review. The terms: a file per command in the checkout, and the harness offers it under the name the file carries.

C8 — skills. An instruction package the model selects by description, which spec 5 grades as scent and measures rather than trusts. The terms: a directory per skill in the checkout, a description the harness serves to the model, and loading on selection rather than always.

C9 — isolated agents. A scoped instruction set that runs in its own context and reports back, which is what keeps the always-on prompt small. The terms: a file per agent in the checkout, and an isolated context per run.

The protocol surface

C10 — tools. The MCP server is the one surface that is not harness work at all, because the protocol is the point of the protocol. What varies per harness is registration. Either the checkout declares the server, so that every session of that repository gets it, or only the operator's personal configuration can. A checkout that cannot register its own server has the tools only where every operator has done the same work once each.

What every binding owes, whatever the harness

Four terms hold for a binding to any harness, and the first three restate the hook contract at this altitude.

  • A binding calls a verb that ships, and carries no rule of its own. Two entry points to one answer are two answers as soon as one drifts.
  • A binding fails open, even where the harness fails closed. Copilot's documentation makes an erroring pre-tool hook deny the call. On such a harness the binding catches every failure of its own. A missing engine, or an unparseable input, must never deny an edit that the positions below already hold.
  • A binding binds nothing, and the position under it does. The commit gate and the CI job hold every change, no cell of the table below reaches either, and a column prices earliness alone.
  • A fixture suite drives every position over recorded input, including every refusal. .claude/hooks/fixtures.sh is the reference: a hook that has refused nothing is a hook that nobody has seen work.

The support table

The Claude Code column is bound in .claude/. Two suites hold it: .claude/hooks/fixtures.sh for the positions, and .claude/skills/fixtures.sh for the skills. The Codex and Copilot rows for C3 through C6 bind the same way, in .codex/hooks.json and .github/hooks/*.json, and .claude/hooks/fixtures.sh holds their cases too. .claude/hooks/fixtures-live.sh reruns the live confirmation against a real install of each harness, and nothing gates on it, the same posture as headwater probe. The remaining cells of both columns are read from vendor documentation dated 2026-08-25, and no fixture holds them.

A live run against copilot-cli 1.0.83 confirmed every Copilot case in fixtures-live.sh on 2026-09-09. Its own reachability probe and per-case timeout undercounted a working install. A trivial prompt took 44 seconds to answer, well past the 30-second probe. A live install then read as unreachable. Every case skipped in silence. The suite now waits 90 and 180 seconds, not 30 and 100.

Capability Claude Code GitHub Copilot OpenAI Codex
C1 standing context CLAUDE.md .github/copilot-instructions.md, and it reads AGENTS.md AGENTS.md
C2 scoped context a CLAUDE.md in the directory it covers .github/instructions/*.instructions.md, with an applyTo glob an AGENTS.md in the directory it covers
C3 injected context UserPromptSubmit hook userPromptSubmitted hook UserPromptSubmit hook
C4 write refusal PreToolUse, deny with a reason preToolUse, deny with a reason, and an erroring hook denies PreToolUse, exit 2 with the reason on standard error
C5 write advisory PreToolUse, added context and no decision, and PostToolUse adds one line when the edit left a governing edge suspect. PostToolUse on Bash names a derived artifact that a stopped merge left in conflict preToolUse, and whether it gives added context to the agent is not measured PreToolUse, and whether it gives added context to the agent is not measured
C11 read advisory PreToolUse on Read, bound and held by .claude/hooks/fixtures.sh, and whether the agent receives the added context is not measured live not bound: the name of the read tool is not measured, and .claude/hooks/fixtures-live.sh records it not bindable: Codex reads a file through its shell, and a shell call reaches no read matcher
C6 turn gate Stop, exit 2 blocks the turn agentStop, decision: block on the CLI and the cloud agent, and the IDE's Stop cannot block Stop is documented, and whether it blocks is not
C7 commands .claude/commands/*.md .github/prompts/*.prompt.md ~/.codex/prompts/*.md, personal rather than repository configuration
C8 skills .claude/skills/*/SKILL.md .github/skills/, and it reads .claude/skills/ and .agents/skills/ .agents/skills/, selected by name in the composer
C9 isolated agents .claude/agents/*.md --agent <name> reads .claude/agents/<name>.md directly, and .github/agents/*.agent.md mirrors some personas for a picker subagents in the CLI
C10 tool registration .mcp.json in the checkout a workspace file in the IDE, repository settings for the cloud agent [mcp_servers] in config.toml, operator configuration
position registration .claude/settings.json .github/hooks/*.json, and the IDE also reads .claude/settings.json .codex/hooks.json, or ~/.codex/hooks.json, behind a feature flag that ships on

Every Claude Code cell except C2 is bound. The Codex and Copilot rows for C3 through C6 are bound the same way, each held by .claude/hooks/fixtures.sh. C11 is bound for Claude Code alone, and its row states why for the other two. .claude/hooks/fixtures-live.sh is the live confirmation behind that, and it reruns against a real install rather than resting on one session. Every other cell of both columns is a file this repository ships or a claim vendor documentation makes, and no fixture holds either kind.

The shape converged, and it is the shape this repository already ships. All three harnesses read a SKILL.md under a per-skill directory. All three name the same four positions, and all three read a standing file from the checkout. Two of them read this repository's own files: Copilot's IDE discovers hooks in .claude/settings.json, and its skill loader reads .claude/skills/. So part of the Claude Code binding is loaded by a second harness through that harness's own choice, which nothing here has measured. A binding for either other column is registration of the same verbs, not a port of any logic, because the scripts carry none.

The posture inversion is the one hazard the vocabulary has to name. Every position of this repository fails open, and one harness documents a pre-tool position that fails closed on error. The second term of the binding contract above exists for that cell.

A second inversion held on the turn gate, and a live run is what found it. Copilot's agentStop does not fail closed on exit 2. An exit-2 hook there is logged, and the turn ends anyway. The block that works is a decision: block object on standard output at exit 0. .claude/hooks/review.sh branches on the COPILOT_CLI environment variable for that one difference, and Claude Code and Codex keep the exit-2 path the hook contract states.

Copilot resolves an agent from two places, and only one of them is in the table. --agent <name> reads .claude/agents/<name>.md directly, confirmed live by removing this repository's own .github/agents/hw-queue.agent.md and finding dispatch unchanged. The mirror files under .github/agents/ serve a picker, not dispatch. A resumed session drops this identity unless --agent names it again. A bare --resume answers in Copilot's own voice. A repeated --agent restores both the persona and its memory of the session. The same round trip, run twice, confirmed this.

Copilot's own agent can dispatch a persona itself, through a task tool. This table does not yet cover it. A running session can launch a named persona as an in-process sub-task. It reads the report back, synchronously or in the background. The synchronous case is confirmed live. Its model handling is stricter than the CLI's own --agent flag. An unavailable model in a persona's front matter fails the dispatch outright. --agent instead warns, and substitutes a model that exists. No cell of this table rests on the task tool yet, and tools/run/copilot-next-run.md states why.

The IDE differs from the CLI under one vendor name. Copilot's completion surface reads none of this table, and its IDE cannot block a turn where its CLI can. Its code reviewer reads instructions from the base branch rather than the feature branch. The column records the most capable surface, and a person binding one surface reads the vendor's own reference for that surface.

No column moves enforcement. The table prices where a finding arrives, and the hook contract already prices what that is worth. A refusal at write time costs one retry, and the same refusal at review time costs a rewrite. Whether any cell earns its context cost is a probe question, no probe has run, and no cell of this table rests on one.