Rendered from docs/evaluations/README.md in the Headwater
corpus. Every document on this half of the site is typed by the taxonomy
the descriptor names: corpus.json.
Evaluations
The documents on this shelf, in the reading order this corpus derives.
- A probe task section is the prompt, so commentary under that heading is prompt — Two probes leaked commentary into the section a recorder sends as the prompt, and a third declared an answer set with no output contract. (asserted, and no human has accepted it)
- Adjacent work and tooling — Practitioner tools and standards that solve adjacent problems, and what this design specifies that somebody already built.
- The capability-bundle address inventory — The cross-bundle address inventory HW-DR-0044 asks for, and what it shows decision-record actually reaches inside design-spec. (asserted, and no human has accepted it)
- What ships in the box — a first-run walkthrough — The Q3 evidence, which wrote the base package out as real YAML and ran five adopters through their first day against it.
- External evaluation harnesses and benchmarks for the probe campaign — DeepEval, Promptfoo and Inspect AI can drive a probe session but never grade one. The campaign adopts none, because each reads model text or opens a channel the session can write to. (asserted, and no human has accepted it)
- First contact — the Q11, Q12 and Q16 evaluation — The Q11, Q12 and Q16 evidence, which is the license, the migration path, and the public presence that a first reader meets.
- Governs edges: what an anchor reaches, what covers it, when it ages, and where a session meets it — A governs edge reaches one file by equality, 53 of 374 documents declare one, and nothing reports a governed file that changed. One ruling and four designs close the gaps. (asserted, and no human has accepted it)
- The graph, its export, and the tier above it — The Q6, Q13 and Q9 evidence, which is where the corpus graph lives, what it exports, and how a second repository consumes it.
- Harper over this corpus finds no error on a 25-document sample, and version 2.11.0 does not build at the engine's floor of rustc 1.91 — harper-core 2.11.0, fed only the prose an author wrote, found 0 errors in 187 hand-read findings on 25 documents. It adds 496 crates and does not compile at rustc 1.91. (asserted, and no human has accepted it)
- inspect_evals as a worked instance — the first ADR log this project did not write — Ten real architecture decision records from a UK government project, typed against the decision-record entry, and what the run found that their own linters cannot. (asserted, and no human has accepted it)
- The implementation language — the evidence for Q1 — The Q1 evidence, which decides the implementation language on embedding, scope enforcement and sum types rather than on speed.
- The Q1 spike — results — The results of the Q1 risk-retirement spike, which are four items, all passing, and three findings that the argument did not predict.
- The Headwater taxonomy in LinkML — a worked example — The Headwater taxonomy written out in LinkML, and the boundary where the standard stops covering what spec 2 declares.
- n8n as a worked instance — a corpus with no docs root — Forty-six real governing documents from n8n, typed against three entries of the library. This document also records what three runs and one coherence sweep found in a corpus that keeps its prose beside the code. (asserted, and no human has accepted it)
- The taxonomy in OWL and SKOS — a worked example — The taxonomy emitted as OWL and SKOS, the corpus as instance triples, and what a reasoner makes of the result.
- Relation storage — the Q4 evaluation — The Q4 evidence, which makes a relation instance an object in front matter and refuses the annotated prose link as a second edge syntax.
- Choosing the schema format — a cognitive-dimensions walkthrough — The Q2 evidence, a cognitive-dimensions walkthrough over five authoring scenarios, which chose YAML and found five defects in spec 2.
- The Headwater checks in SHACL — a worked example — The Headwater checks written out in SHACL, and the whole-graph line where the constraint language stops.
- Specifying the engine —
docs/spec/is a design-spec series and nothing types the engine as a thing under test, and three traditions each supply one part of what would. - The measurement layer — the Q8 and Q20 evaluation — The Q8 and Q20 evidence, which is what a probe costs, when it runs, and whether the promised instruments can produce their measurements.
- The pointer probe grades a session against the router's own first pick, so a correct session that declines a wrong pointer fails — A session that declined a wrong first pointer answered correctly from two spec parts that the router ranks past 100th. (asserted, and no human has accepted it)
- The serving boundary — what is advertised, what is withheld, what is written back — The Q14, Q17 and Q7 evidence, which is what a corpus advertises, what it withholds, and what a tool may write back.
- The three discovery misses of 2026-09-16 have three different mechanisms, and only one is routing precision — A replay of the router shows three different misses: one true precision failure, one ignored correct offer, and one target the router cannot reach. (asserted, and no human has accepted it)
- Theoretical foundations — Where the research literature confirms, sharpens, or contradicts the design, and what changed in the specification as a result.
- Warrant — what stands behind a document, and who vouched for it — The Q15, Q19 and Q18 evidence, which is what stands behind a document, who vouched for it, and how a disagreement is settled.
- What a check can know — the Q5 and Q21 evaluation — The Q5 and Q21 evidence, which measures what a lexical checker gets wrong on the checker that this repository already runs.
- What the counterfactual campaign of 2026-09-30 measured, by component — The documents under docs/ raised the sufficiency rate by 57.5 points, and the hooks and skills showed no detectable gain, at a cost of $304.46. (asserted, and no human has accepted it)
- Why corpus counts are derived, not stored — Counts over the whole corpus are computed when read, not kept in committed files. Stored counts cause silent merge failures when two branches write the same total but the merged tree holds a different number. (asserted, and no human has accepted it)