A documentation governance system

Documentation rots because nothing holds it accountable.

Headwater types your corpus, checks it as a graph, and accounts for every file it saw. Nothing else can answer β€œis this corpus still true?”

$ curl -fsSL https://headwater.tools/install.sh | sh
Linux x86_64 and macOS on Apple silicon. All install options β†’
---
id: HW-DR-0037
relations:
  governs:
    - site/index.html
    - site/how-it-works/index.html
    - site/compare/index.html
    - site/proof/index.html
    - site/tutorial/index.html
---
This page is not a typed document. It sits outside the corpus root, so no rule reads it. That block is the record which governs it, and the record is checked like any other. A page here with no edge above it is a page the engine would let you delete in silence.
$ headwater check
census
  742 files under the corpus root
      525 typed
        2 untyped
        0 unaccounted for
  404 findings
        16 suppressed, by dated directive
  taxonomy headwater/standard 4.15.0
01

Measured β€” headwater check on this repository, 2026-10-02

Where a vendor site puts logos and percentages, this one puts the run β€” including the ugly number. A full board with a name against each item beats a clean one.

742
files under the corpus root β€” the denominator for the tiles that follow
525
of those files typed, and all 525 checked
0
of those files silently unaccounted for
404
findings reported, every one named and located
16
further findings suppressed, each by a dated directive
9
of the 404 are errors β€” every one is advisory

13944 check instances ran over those 525 documents, from 48 wired rules, of which 7 found anything. The obligation register behind them holds 51 obligations: 48 verified, 2 gap and 1 unverifiable, and each gap names its owner.

Benchmark against comparables unmeasured

A measurement campaign ran on 2026-09-30, and it measured this repository against itself with parts removed, not against the projects it is compared with. So this row stays empty. Principle 11 forbids publishing a number that no run produced.

02

How it works

1 Β·Declare the taxonomy

Shelves, kinds, facets, and relations, in one versioned schema. Your overlay adds what a package cannot know.

shelves:
  decisions: docs/decisions/**
kinds:
  decision:
    purpose: rationale
    facets: [status, summary, last_verified]
    relations: [supersedes, governs, traces_to]

2 Β·Authors declare the graph

Documents are typed nodes. Front-matter references are typed edges, and a required edge owes both ends. No language model issues a verdict.

0002-deliver-at-least-once.md
supersedes ↓
↑ superseded_by
0001-store-attempts-in-postgres.md

3 Β·Every rule constrains that graph

A finding names the rule, the obligation it discharges, the location, and the repair. Mechanical repairs are one command.

0002-deliver-at-least-once.md:9:7 error
relation.reciprocity.missing (OB-REL-1):
`supersedes` requires both ends, so
0001-store-attempts-in-postgres.md
owes `superseded_by`
fix (mechanical): headwater check --fix

4 Β·Everything derived is a projection

Indexes, site navigation, an agent instruction file, a JSON export. A hand-edited projection fails the build.

docs/spec/README.md ← generated
docs/interfaces/README.md ← generated
site navigation ← generated
agent instruction context ← generated
headwater generate --check

The four steps in full, and where the reasoning stops β†’

03

The three commitments

I.

Taxonomy is data.

What shelves exist, what kinds live on them, what metadata they carry, and how they may reference each other β€” all declared in one versioned schema. Customizing the taxonomy never means forking the tooling.

II.

The corpus is a graph.

Documents are typed nodes and front-matter references are typed edges. Every validation rule is a constraint on that graph, and every derived artifact is a projection of it.

III.

AI assistants are readers too, and measured ones.

The corpus feeds agents at intent time, read time, write time, and review time. A probe records whether that context changed what the agent did β€” not whether it was shown.

commitment II, drawn β€” real documents from the tutorial corpus

0002-deliver-at-least-once.md kind: decision Β· ACME-DR-0002 0001-store-attempts-in-postgres.md kind: decision Β· ACME-DR-0001 supersedes ↓ ↑ superseded_by

Every edge above was declared by an author and checked by the engine β€” supersedes owes both ends, and the engine wrote the far half. Edges name targets by identifier, so a rename kills nothing.

04

Against the field

Authors declare the graph here. A language model extracts it there.

LeanCTXShares the Open Knowledge Format, none of the type system. Their pitch is token economics. Ours is whether the documents are true.
TrustGraphThe same pitch, the opposite mechanism. Documents there are feedstock, dissolved into triples. Nothing governs the sources β€” it mines them.
ValeThe closest analog. It checks prose style, line by line. Headwater checks document structure, edge by edge. They compose.
Plain lintersCheck lines, not the graph between documents.
Site generatorsRender whatever you feed them, true or not.

Read the full comparison, including the arguments against us β†’

05

Status, honestly

The engine runs, and one campaign has measured part of what it claims. This repository types its own corpus, checks it on every commit, and scaffolds, explains, routes, generates and exports it. The counterfactual campaign of 2026-09-30 measured it by component. The documents under docs/ have a large measured effect on whether an agent gives a sufficient answer. The hooks and skills showed no detectable gain. The evaluation states each figure, with its interval and its cost, and this page repeats none of them. No arm measured the MCP server or the quality of governance itself, and the efficacy claims stay open until one does. The work still open is on the milestones page.

The command surface is 26 verbs in 7 groups, and all 26 carry a written contract. The export layer names 7 emitter targets, of which 2 are built and 5 parse and report the consumer they wait on. The self-assessment states the rest.

Run it against your own corpus, and read what it can't yet answer.

Run the tutorial β†’