Rendered from docs/process/decisions/0011-crates-publish-in-one-fixed-order-that-a-test-holds-and-a-rerun-continues-where-the-last-run-stopped.md in the Headwater
corpus. Every document on this half of the site is typed by the taxonomy
the descriptor names: corpus.json.
Crates publish in one fixed order that a test holds, and a rerun continues where the last run stopped
Context
Before this record, the reasons lived only in comments of .github/workflows/publish-crates.yml: the header (lines 21 to 65 on 2026-09-27) and the comments on the publish loop (lines 116 to 127 and 147 to 152).
cargo install headwater-cli needs every workspace crate that the binary depends on to be on crates.io. A crate resolves its dependencies through the registry when it publishes. So each crate must publish after every crate that it depends on.
Three behaviors of crates.io shape the loop:
cargo publishrefuses a version that the registry already has.- The crates.io index makes a new crate available to the next
cargo publishafter an interval that crates.io does not state. - crates.io limits how fast a new crate name publishes. A new version of a name that already exists has no such limit. A response with status 429 names the time in GMT at which crates.io accepts the next one.
Decision
One step publishes every crate, in one loop over a fixed list. The list is in the order variable of the workflow, leaves first. A person computed it once from the dependency graph, and no run computes it again. A dependency that a crate adds is then a diff in the list, and not a change of order that nobody sees. engine/crates/cli/tests/publish_order.rs holds the list against the members of the workspace. It holds the set exactly, and it holds the order wherever one member depends on another.
Before it publishes a crate, the loop asks crates.io whether the crate has this version. If it has, the loop skips the crate. So a second run continues where the first run stopped, and does not start again at the first crate. The request sends a User-Agent of its own. The crates.io API refuses the default User-Agent of curl with a 403, which the loop cannot tell apart from a 404.
The loop waits for the index and not for a fixed time. Before it publishes a crate, tools/repo/crates-index-wait.sh reads the crates.io sparse index until it lists each workspace dependency of the crate at this version. After cargo publish returns, the same script waits until the index lists the crate itself. So a warning from cargo that its own wait timed out does not count as a success. If the index does not list a name after 600 seconds, the script names the crate and the dependency, and the run stops. engine/crates/cli/tests/publish_order.rs holds the script against an index on disk.
When cargo publish fails with a 429, the loop reads the time from the error. It waits until that time and 5 seconds more, and tries the same crate again. Any other failure stops the run, because waiting cannot fix it.
Consequences
A maintainer recovers from a partial publish with a second run of the job, and never with a yank. At v0.3.0 the first run stopped at headwater-compat, because the index did not yet list headwater-scaffold after the 30-second wait. A second run published the four crates that were missing. At v0.4.0 the first run stopped again, at headwater-import, because the index did not yet list headwater-scaffold. A second run published the rest. These two occurrences are why the fixed 30-second wait became the index wait (#1197). The index wait has a deadline, so a slow index can still stop a run. The skip makes that stop safe to retry.
The wait for a 429 has no ceiling. HW-OBL-0182 records that gap, and it stays open.
This record states no count of crates. publish_order.rs holds the set, and a count in prose goes stale when a crate joins the workspace.