Chapter 39 — The evidence culture
status: polished · path: Muse Glimmer, pinned Muser tree
Prerequisites: Chapter 38 (the measurement protocol this culture wraps) and passing familiarity with Ch 30 and Ch 32, where exactness policies first appeared as contracts rather than hopes.
39.1 The chapter that makes the other chapters cheap to trust
Chapter 38 gave you the instrument: interleaved ratios, exact-token gates, five-rep means, receipts on an append-only volume. But an instrument only tells you what happened. The questions that come immediately after are the ones this chapter answers: when does a measurement earn the right to be called a claim, who is allowed to say it out loud, and what happens when the number and the sentence disagree?
Everything wrapped around the instrument exists to answer exactly that — the locks, registers, contracts, and tags. Muser calls the whole assembly the evidence culture, and its constitution is one sentence from the working agreements:
“Never weaken a fail-closed check to make a run pass. If a gate rejects your evidence, the evidence is wrong until proven otherwise.”
[AGENTS.md, "Hard rules"]
That sentence inverts the ordinary debugging instinct. The ordinary instinct, when a gate rejects your run, is that the gate is too strict. Here the default is the opposite: the gate is presumed right, your evidence is presumed wrong, and the burden sits exactly there until an audit moves it.
Read it a second way, because the first reading undersells it. The rule does not claim that gates are always correct — gates have bugs like everything else. It claims something narrower and more useful: the cost of being wrong is deliberately loaded onto the person holding the failing run, never onto the check. Loosening a threshold is cheap and quiet; producing an audit that moves the burden is expensive and loud. The culture makes the honest move the cheap one by making the dishonest move impossible to do silently. Every mechanism in the rest of this chapter is that sentence rendered in JSON, and we will walk them in the order a claim itself walks them: refusal, lock, register, release path, wording, live tag, evidence volume, audit.
39.2 Fail-closed, defined and mechanized
Start with the primitive everything else is built from, and start with the question it answers: what should a system do at the moment it cannot prove it is in a good state? There are only two families of answer, and the choice between them decides how much a whole program’s evidence is worth.
Fail-closed means: when a check cannot prove the good state, the system stops, and it stops before the unproven state can be mistaken for a proven one. A fail-open system degrades to permissive under uncertainty; a fail-closed system degrades to refusal. Put the difference in terms of what survives a bad day: a fail-open system’s worst outcome is a number that looks fine and is not, which you may never catch; a fail-closed system’s worst outcome is a stopped run, which you catch immediately by definition. You have already met a dozen instances without the word:
- The producer exits with status 75 on any engine-touched error, and a bare
docker restartis not enough — stale startup receipts, RoPE caches, and sockets must be cleared by the restart ritual[AGENTS.md, "The GX10 lane"](Ch 28). - Serving refuses to load
producer_mode: nativetogether with DFlash, with the error stating the remedy, at[crates/muser-server/src/state.rs:1666-1675]: “native NVFP4 fast-lane speculative decode is unqualified; omit –dflash and use plain NVFP4 decode, or route speculative serving to the kquant lane”. The qualifier wrapper carries the same refusal for its variant (target-plus-dflash) at[scripts/qualify_nvfp4_fast.py:333-336]. - Model identity at startup: configured vs verified SHA-256, and “mismatch
refuses”
[crates/muser-server/src/state.rs:1168-1175]. - The accelerator wrapper’s lease refusals, receipt-immutability refusals, and forbidden-command refusals ([Ch 38 §38.8]).
The design pattern in every case: the operator sees the failure — an exit code, an error string naming the remedy, a latch — and the system never silently proceeds on an unverified state. This is why the culture can be a subject of this book rather than an obstacle to it: each refusal is a documented, testable boundary, not a mysterious crash.
39.3 The release lock — one file that outruns everyone
Suppose every measurement in this book came back gold tomorrow: every lane
exact, every gate green, every receipt retained. What stops that from
becoming a release the same afternoon? The answer is deliberately
unimpressive. At the center of the culture sits a single small file,
release/release-lock.json, short enough to read in a minute and blunt
enough that nobody can misread it. As of the pinned tree its actual state is
(Figure 39.1):
{
"schema": "muser.release-lock.v1",
"state": "containment",
"sealing_enabled": false,
"candidate_creation_enabled": false,
"tagging_enabled": true,
"tagging_policy": {
"class": "non-release-marker",
"allowed_tags": ["v0.1.0-beta.1"],
"operator_go_required": true,
"creates_seal": false,
"creates_candidate": false,
"creates_publication": false
},
"publishing_enabled": false,
"blocked_commits": ["11119bd"],
"unlock_requires": "the exact beta marker requires a separate operator go; sealing, candidate creation, and publication require a new lock amendment"
}
Figure 39.1: release/release-lock.json at the pinned tree, quoted in full
— sealing_enabled: false, candidate_creation_enabled: false,
publishing_enabled: false; the only permitted tag is the non-release beta
marker, and only after an explicit operator go.
What “authoritative” means here is literal: while the lock is in
containment, no seals, tags, or release candidates may be created, no matter
how strong the evidence is [AGENTS.md]. Every ledger entry from the
2026-08 campaigns repeats the reminder in its own preamble — “no entry is a
readiness receipt, seal, tag, candidate, or publication” [ledger, preamble]
— and every chapter of this book inherits the constraint. The numbers you
have read are unsealed engineering evidence, every one of them, and the
ledger stamps the fact onto each restatement as seal_eligible: false
[ledger "Synthetic spec matrix deep-cell restatement"].
That is the distinction the campaign calls notarial versus
non-notarial evidence: notarial evidence is a sealed, independently
reproducible release artifact; non-notarial evidence is everything retained
so far. Say it once more in the negative, because the word invites the wrong
reading. Non-notarial does not mean sloppy, preliminary, or unreproducible —
a non-notarial number can be measured exactly and rerun on our bench all
afternoon. What it lacks is the seal that would let a stranger reproduce it
without us in the room. “Measured” carries a similarly narrow meaning here:
measured on Muser, on this hardware, under a retained receipt — never
“measured once, on any hardware, ever”
[docs/launch-claims.md §Ground rules].
The objection writes itself: the lock is a file in the repo, so what stops
anyone from editing it? Two things, and only the second one is durable. The
lock is tracked rather than advisory — commit 11119bd is listed in
blocked_commits, and the feature contract independently declares that same
commit the non-releasable source baseline
[release/feature-contract-v1.json, "source_baseline"], so a quiet edit to one file contradicts the other and
the contradiction shows up in review. More importantly, the unlock has a
prescribed shape. Deleting or relaxing the lock to make a release happen is
not a move anyone has; the only permitted unlock is “a narrowly scoped,
reviewed change setting sealing_enabled true for this exact
readiness-authorized campaign” [docs/private-release.md §3]. An escape
hatch that has to be argued for in the open is not much of an escape hatch,
which is the point.
39.4 Findings and the feature contract — the campaign’s identity
The lock answers when, and its answer is “not yet.” Two further files
answer what: what the release would consist of, and what still stands
between it and existing. Together with the lock they complete the
constitutional set, and the working agreements warn about them in the same
breath — “changes to them change the campaign identity” [AGENTS.md].
Editing one of these files is not maintenance. It is starting a different
campaign, under a different identity, whose earlier evidence no longer
applies.
release/findings-v1.json is the defect register, and its policy line
only has to close two doors: {"waivers_allowed": false, "release_requires_zero_open": true} [release/findings-v1.json]. Those are
the two escapes a defect register is normally asked for — ship with the
defect and note it (the waiver), and ship with it still open and fix it
next cycle (the deferral). Neither exists here, which means the register’s
open-row count is a gate rather than a status report.
Each finding has id, severity, area, title, status, and resolution; the register
spans 44 rows from REL-001 (the blocked commit) through security
(SEC-001–SEC-005: TLS, CORS, CSRF, WebSocket tickets, CA workflow),
enrollment, replay-ledger durability (REP-001: “generation reservation
ordering and durable fsync protocol incomplete”), scheduling (SCH-001:
the global session mutex replaced by the four-slot pool of
Ch 34), to the performance finding PERF-001.
PERF-001 is the instructive one: it was closed not by a wave of the hand but
by enumerating the retained verdict-grade evidence — the six-depth plain
matrix, the fixed-window spec ratios, the funded-fix 131,008 wall parity,
the disaggregated TTFT/link/determinism/soak gates — and its closure text
still bounds the claim: “Scope remains exactly the measured synthetic and
single-producer lanes” [release/findings-v1.json, PERF-001]. Notice the
shape of that closure. It does not say “fixed.” It is a list of retained runs
plus a fence drawn around what those runs cover — a closure in this register
is itself a piece of evidence, which is why closing PERF-001 took a campaign
rather than a commit.
release/feature-contract-v1.json fixes what the release is: the
hardware contract (one M3 Ultra 96 GB decode host, four slots at 131,072
context, GX10 as prefill/storage node and never a decode destination), the
in-scope list (single model, llama-pinned parity, vision, DFlash, GX10,
dashboard, sessions, migration), the out-of-scope list (LoRA, hot-swap,
infill, hosted-provider APIs, “public-CoreML ANE DFlash routing
(experimental post-release)”), and a release policy that reads like the
ledger’s ethics compressed into six booleans (Figure 39.2):
"release_policy": {
"waivers_allowed": false,
"open_findings_allowed": false,
"qualification_skips_allowed": false,
"owner_tags_or_publishes": true,
"seal_requires_release_readiness_receipt": true,
"post_seal_change_invalidates_campaign": true
}
Figure 39.2: [release/feature-contract-v1.json] — no waivers, no open
findings, no skipped lanes, owner-only publication, and any post-seal change
invalidates the campaign.
That last clause is the sharpest: after a seal exists, any change — source,
artifact, documentation — does not get patched in; it restarts the stage
[docs/private-release.md §3]. A sealed campaign is a photograph, not a
living document.
39.5 The release path — freeze, run, readiness, seal
Grant, for a moment, that the lock does open. What happens then is not “cut a
release.” It is a fixed sequence with a stop at every junction, and the
sequence is worth studying even though it has never run to completion,
because its shape is a list of the things the culture is afraid of. The one
permitted path from “lots of evidence” to “a release” runs like this
(Figure 39.3) [docs/private-release.md]:
flowchart TD
A[Freeze one clean identity:<br/>findings, contracts, provenance,<br/>matrix config, binaries] --> B[Run all 15 mandatory<br/>lanes UNSEALED]
B --> C{Zero open findings?<br/>All lanes exact identity?}
C -- no --> D[STOP: fix and re-freeze]
C -- yes --> E[One readiness receipt]
E --> F[Atomic final campaign:<br/>freshly rerun all 15 lanes<br/>into a hidden directory]
F --> G[One fsync-backed rename<br/>exposes the whole bundle]
G --> H[Candidate built only<br/>from that exact bundle]
H --> I[Two independent verifiers<br/>including a clean-room rebuild]
Figure 39.3: The freeze→run→readiness→seal flow [docs/private-release.md].
Failure at any stage exposes nothing; the seal bundle appears atomically or
not at all.
The fifteen mandatory lanes are enumerated by name — correctness,
sampled, greedy, kvpack, session, vision, baseline, dflash,
remote, serving, onboarding, api-parity, continuous-batching,
migration, security — and “a skipped, unstable, malformed, cross-lane,
wrong-identity, or unsealed=false report is a failure. There are no waivers”
[docs/private-release.md §2]. The final campaign reruns fresh — the
sealed matrix is measured after readiness, not assembled from remembered
numbers — into a hidden sibling directory, fsyncing everything, exposing the
bundle with one rename; “failure exposes nothing” [docs/private-release.md §4].
That last phrase is the atomic-seal idea in one image, and it is worth holding onto. Anyone reading the evidence directory sees either no bundle at all or a complete one; there is no window in which they can catch the campaign mid-sentence and mistake a partial run for a result. A crash halfway through leaves nothing but a hidden directory of garbage — which is the correct outcome of a failed release, and the reason the rename comes last.
Even the candidate verifiers are structural: a second clean-room
verifier “must extract the source archive, perform the offline locked build,
re-hash the resulting binary … and run the loopback smoke request on an
externally offline host” [docs/private-release.md §5], under the
accelerator lease.
None of this has run to completion: the lock is still in containment, and every number in this book is pre-seal. The machinery’s purpose is precisely that this fact is checkable from one file rather than folklore.
39.6 The launch-claims register — copy never outruns the receipt
Evidence decides what is true. Words decide what a reader ends up believing you said. Most technical dishonesty lives in the gap between the two — not in the numbers, which are usually fine, but in the sentence built on top of them, one adjective wider than the measurement supports. So the question this section answers is: where does that gap get closed, and by whom?
docs/launch-claims.md is the answer — the interface between measurements
and words. It is a table — seventeen numbered rows at the pin — where every
row carries its current evidence (with receipt paths), its conditionally approved
wording, and, where wording exists but the owner has not approved it, the
banner OPERATOR REVIEW REQUIRED. The register’s ground rules are the
culture’s most quotable sentences; four verbatim:
“The release lock (
release/release-lock.json) is authoritative: while the feature contract is in containment, no row above goes live regardless of how strong its evidence is. Conditional wordings activate only when the contract leaves containment and the row’s stated reproduction gate passes.”[docs/launch-claims.md §Ground rules]
“A number with no row above does not ship. Add a row (with its evidence citation) before using it.”
[docs/launch-claims.md §Ground rules]
“
[precedent-7B-ferrite]numbers (e.g. 34.9/308 t/s, 21.9–30.1x restore, 24.6 GB/s fabric, 1.42x ANE+GPU concurrency) describe the historical Ferrite research lineage on a 7B model, never this Muser program. They may appear in engineering docs as context but never as a Muser product claim.”[docs/launch-claims.md §Ground rules]
“When evidence and wording conflict, evidence wins and the wording row gets corrected — copy is never allowed to outrun the receipt.”
[docs/launch-claims.md §Ground rules]
Three of those rules govern what a number is allowed to become; the fourth governs what happens when a sentence has already got ahead of its number, and its direction is not negotiable — the wording moves, the evidence does not.
The OPERATOR REVIEW tier is the register’s subtlest device. Rows #2, #6,
#11, #12, #15, #16, and #17 carry evidence-backed proposed wording that
the owner has not approved; the register states the rule outright — such
rows “remain unavailable to launch copy even if its reproduction gate later
passes, until the operator approves it” [docs/launch-claims.md, preamble].
Evidence quality and wording approval are orthogonal axes: a perfect
five-rep matrix still cannot speak until a human owner signs the sentence.
The review package for those rows exists
(docs/launch-claims-review-20260824.md), states each row’s exact proposed
wording, receipts, and risk (“removing ‘synthetic,’ ‘mean,’ or the
tested-depth scope would turn a controlled fixture result into an
unsupported workload-general claim” [docs/launch-claims-review-20260824.md, claim #2]), and closes with the discipline that “this review package
remains the pre-decision record” — it prepares decisions, it does not make
them [docs/launch-claims-review-20260824.md, preamble].
The register also carries the negative space: an “Explicitly post-launch”
list of things that do not exist (node discovery, multi-node scheduling,
revocation, full-depth reuse coverage, remote multimodal, send-during-
prefill) with the instruction that “they simply do not exist yet and must
not be implied” [docs/launch-claims.md §Explicitly post-launch]. A claims
register that only lists what you have is half a register; the other half is
listing what a reader might reasonably assume you have.
39.7 Honesty tags — the metrics schema
The same discipline reaches into the live server payload, and it gets there
by way of a small design question with a load-bearing answer: what should a
dashboard show for a quantity nobody has measured? The tempting answer is
zero. Zero renders cleanly, keeps the layout intact, and is a lie shaped
exactly like data. Muser’s answer instead is that every field in the
telemetry snapshot carries an honesty tag, with the legend enforced in both
prose and code [docs/metrics-schema.md]:
measured— “a live counter, duration, or verified loaded-model fact”;target— “a threshold or modeled goal, never an observed result”;mock— “no backing measurement is available; the dashboard renders the value unavailable.”
The register’s copy legend extends the same idea with five tags for claims —
[measured] / [precedent-7B-ferrite] / [target] / [roadmap] /
[mock] [docs/launch-claims.md, preamble]. The two legends look redundant
until you notice they cover different surfaces: three tags for a live
telemetry payload, five for launch copy, which has to make one distinction
telemetry never faces. That extra distinction is [precedent-7B-ferrite],
and the reason it matters to every chapter of this book is that it keeps
ancestor-lab numbers quarantined where no reader can mistake them for Muser
measurements.
The mock rule has one canonical application. The dashboard’s nodes[]
array — M3/GX10 utilization, memory, power, temperature — “is currently
empty and tagged mock: … collection are not wired to this payload. The
separate node-management API and registry do not manufacture telemetry node
cards” [docs/metrics-schema.md §Cluster and nodes]. And the optimization
card list is empty on purpose: “tricks[] is intentionally empty. No
optimization card appears until its independent correctness and performance
qualification passes for the release identity. Historical Ferrite results
are provenance, not live Muser metrics” [docs/metrics-schema.md §DFlash and optimization claims]. That last sentence is the dashboard-mock-tagging rule
in its general form: fields without measurement render unavailable, and
historical Ferrite results are never inserted as live Muser measurements.
An idle counter may legitimately read a measured zero; a modeled threshold
may never dress up as an observation [docs/metrics-schema.md, preamble].
39.8 Evidence volume discipline — where truth is allowed to live
Two questions sound like one question here, and telling them apart took a measured failure. Where is a receipt allowed to live? And what else is allowed to live beside it? The first has a short answer; the second we got wrong first.
Retained evidence lives on muser-receipt:// and is
append-only [AGENTS.md]. The wrapper’s mechanics make the append-only
property physical: receipts are created through exclusive temp-file +
fsync + rename + directory-fsync, and the publish function’s first act is to
refuse if the target exists — “refusing to replace result receipt”
[scripts/accelerator_safe.py:202-203]; the run journal is opened
O_APPEND and fsynced per record [scripts/accelerator_safe.py:190-197].
That property invites an obvious generalization, and we took it. If the evidence volume is the durable place — exclusive create, fsync, rename, refuse-on-exists — why maintain two storage stories? Put the operational state there as well: replay ledgers, sockets, locks, all on the disk built to never lose a write. We expected the volume’s guarantees to carry over intact.
They did not carry over. The 2026-08-18 durability investigation (fully told
in Ch 31) found that the evidence volume’s
directory-fsync tail produced bimodal ~1 s stalls in the commit path
[AGENTS.md]: a ledger commit that should have been imperceptible would
instead, some of the time, freeze the lane. Nothing was wrong with the
durability. What we had never tested was its latency distribution — a
different property of the very same fsync. That is the lesson worth carrying
out of the episode: a volume tuned so that writes can never be lost is not
thereby a volume on which writes are always quick, and the two properties
have to be measured separately because only one of them was ever on trial.
So the decision went the other way. Operational state — replay ledgers,
sockets, locks — belongs on the internal disk, and because a lesson that
lives only in a document decays, the receiver now probes rather than
trusts: check_ledger_volume measures the reserve-pattern tail
latency and refuses a slow volume before any handoff
[crates/muser-cluster/src/receiver.rs:108-150], with
scripts/gx10/durable_fsync_probe.py as the standalone probe (exit 1 past
--max-tail-ms) [scripts/gx10/durable_fsync_probe.py:19-22]. Evidence and
operations are separated not by convention but by measured failure mode.
39.9 The documentation truth pass — auditing claims against receipts
Documents drift; code moves; receipts stay. Which raises the question this section exists to answer: how do you catch a claim that was true on the day it was written and quietly stopped being true while nobody was watching it? Nothing in the machinery so far catches that one, because nothing so far re-reads old sentences. You have to go looking, deliberately. A documentation truth pass is the genre of audit that re-reads every claim-bearing document against implementation and retained evidence.
Muser’s 2026-08-15 pass checked README
and CLI help against cli.rs, security text against the Axum authorization
policy, architecture against the slot pool and GGUF geometry, dashboard
copy against MetricsSnapshot, performance claims against the retained
representative artifact — result columns recorded surface by surface
[docs/documentation-truth-pass-20260815.md §Sources checked]. Its
performance wording ruling is a model of the form: the one-sample 3.6 %
prefill / 22 % decode figures are “engineering-only” and may appear only
“always with the single-run/non-notarial limitation”; the register
“expressly authorizes no product throughput wording” [docs/documentation- truth-pass-20260815.md §Performance evidence wording].
The same genre runs continuously in the ledger as CORRECTION / RETRACTION / AMENDMENT / SUPERSEDED entries, and in the claims register when evidence moves faster than wording. One worked example, dissected.
The evidence box: a stale claim, handled correctly
The claim. On 2026-08-20 the campaign close-out brief reported, for the external reviewer: “Spec decode vs llama spec: 107.91 vs 81.30 tok/s = 1.327×. PASS,” plus a full spec context matrix “decode means 1.305 / 1.278 / 1.282 / 1.250 / 1.232 at 2k–65k”
[docs/campaign-review-brief- 20260820.md §The campaign]. Every number recomputed exactly from receipts; the red-team review verified “there is no fabrication and no result-shopping”[docs/redteam-review-campaign-brief-20260820.md §Verdict].The staleness. On 2026-08-21 the half-window root cause landed ([Ch 38 §38.7]): every one of those figures was measured while the DFlash draft ran on half its trained sliding window. The measurements were real; the lane they measured was broken. Synthetic speed ~5 % optimistic, natural-text acceptance catastrophically pessimistic.
The handling. Nothing was deleted. The brief now opens with a supersession banner: “SUPERSEDED — 2026-08-21. Every speculative-decode figure below was measured while the DFlash draft was conditioned on half its trained sliding window … All spec claims here are pending re-measurement at the fixed sha. The non-spec content (Phase 2 plain matrix, Phase 4 disaggregated payoff …) is unaffected”
[docs/campaign-review-brief-20260820.md, banner]— the identical banner sits on the red-team brief[docs/redteam-review-campaign-brief- 20260820.md, banner]. The ledger restated the numbers in new entries (1.23692 @2,048 and sisters); the claims register rows #15/#16 now cite only the fixed-window packets; and the old 1.3273/1.3012 figures survive in exactly one role — as the superseded numbers this book tells you not to cite[docs/launch-claims.md #15].What the example teaches. A stale claim is not a scandal; leaving one standing is. The culture’s answer has three moves — the evidence is preserved, the supersession is written on the artifact itself where the next reader cannot miss it, and the replacement claim is scoped tighter than the one it replaces.
The register shows the same move in miniature: claim #6’s wording rule
still instructs “Do not cite the historical 5.83× exact-mirror comparison”
[docs/launch-claims.md #6] — a retired number whose tombstone is kept
inside the very row that replaced it.
39.10 Red-teaming the record
The truth pass audits documents against implementation. That leaves the
uncomfortable question one step further out — who audits the people doing the
auditing, before an outsider ever sees the file? The campaign’s answer was to
red-team itself before asking an external reviewer anything: a review
document produced by “seven independent auditors
(ledger forensics, raw-receipt recompute, Phase-4 packet forensics,
engine-code attribution audit, statistics, framing/honesty,
completeness), plus direct spot-verification of every load-bearing claim.
No measurements were run; nothing in the repo or evidence store was
modified” [docs/redteam-review-campaign-brief-20260820.md, header]. Its
verdict is the culture’s certificate: “The measurements are real. Every
headline number in the ledger recomputes exactly from the retained
receipts; the fail-closed machinery demonstrably worked; discarded runs
were kept and are statistically indistinguishable from counted ones —
there is no fabrication and no result-shopping” [docs/redteam-review- campaign-brief-20260820.md §Verdict].
Note what the red-team pass did not conclude: it did not say the
campaign’s decision was right — its ranked findings argue the decision
was mis-posed, the attribution wrong “four different ways,” and the
statistics inverted [docs/redteam-review-campaign-brief-20260820.md §Findings, ranked]. Honest evidence and a defensible decision are
separate claims; the culture’s machinery certifies the first so the
argument can be about the second. That separation is the whole point:
when the receipts are beyond suspicion, disagreeing well becomes
possible.
39.11 What the culture costs and what it buys
Every discipline sends a bill, and a culture whose costs go unnamed is the kind that gets quietly abandoned the first time it is inconvenient. So here is the ledger, both columns.
- Cost: latency on every claim. An OPERATOR REVIEW row cannot ship even
with perfect evidence; a release cannot seal while the lock says
containment; a finding closure needs enumerated receipts, not a
narrative. Measured consequence: the 2026-08-24 wizard PASS (exact
logits, 9.812/8.887/8.690 Gbps) still “remains the operator-review
draft” for public wording
[docs/launch-claims-review-20260824.md, "Draft new row"]. - Cost: negative space must be maintained. The post-launch list, the
mocktags, the “not measured at every depth” caveats — publishing the sensitivity is part of the claim[docs/launch-claims.md §Explicitly post-launch]. - Buys: auditability in one hop. Any number in this book resolves to a
receipt path; any word resolves to a claims row or is barred; any release
question resolves to one lock file. The red-team review could verify
“no fabrication and no result-shopping” because the evidence chain never
breaks
[docs/redteam-review-campaign-brief-20260820.md §Verdict].
The last entry in the buys column is a story rather than a line item, and it
is the episode we would point at if the whole apparatus had to justify itself
once. Fail-closed buys safety under error. A matrix cell came back with
outputs_match: false at the 65536 warm hit, which for a warm-cache result
is about the worst string the harness can print. The tempting response to a
single red square is to argue with the square rather than with the system:
call it noise, rerun it, move on. Fail-closed forbids exactly that move — the
gate is presumed right and the evidence presumed wrong — so the cell was
investigated instead of defended. The investigation found a mundane cause,
and the correction says so plainly: “the 65,536 warm-hit result was an
infrastructure timeout, not a cache-correctness failure”
[ledger "CORRECTION — the 65536 warm-hit result", 2026-08-21]. The cell was
retracted as an infrastructure timeout rather than explained away, and the
valid cell is the one that then passed its gate.
Notice that the discipline paid off in both directions at once. Had the red square been a real correctness bug, arguing with the gate would have shipped it. Because it was not, the retraction sits on the record where any reader can confirm which cell is being counted — the one that passed its gate on its own merits, not the one that happened to be convenient.
There is one more register the culture keeps, and it is the grimmest one: the list of things measured carefully and then rejected. That is the last chapter of this book.
References
[release/release-lock.json]— quoted in full at Figure 39.1; statecontainment, all release machinery disabled, beta-marker-only tagging.[release/feature-contract-v1.json]— hardware contract, scope lists, the six-boolean release policy (Figure 39.2).[release/findings-v1.json]— zero-waiver policy; the 44-row register; PERF-001’s evidence-enumerated closure.[docs/private-release.md]— the freeze→run→readiness→seal flow (Figure 39.3), the 15 mandatory lanes, atomic bundle semantics, clean-room verification.[docs/launch-claims.md]— the register; preamble (OPERATOR REVIEW semantics, five-tag legend); ground rules (four quoted verbatim in §39.6); rows #2, #6, #15, #16; §Explicitly post-launch.[docs/launch-claims-review-20260824.md]— the pre-decision review package with per-row risk statements.[docs/metrics-schema.md]— honesty-tag legend;nodes[]mock;tricks[]intentionally empty; measured-zero vs modeled-target rule.[docs/documentation-truth-pass-20260815.md]— the audit table; single-run performance wording limits.[docs/campaign-review-brief-20260820.md],[docs/redteam-review-campaign-brief-20260820.md]— the SUPERSEDED banner pair, the seven-auditor method (§39.10), and the no-fabrication verdict (the evidence box).[AGENTS.md]— hard rules (fail-closed sentence quoted §39.1), evidence volume rules, operational-state-on-internal-disk.[scripts/accelerator_safe.py:190-197, 202-203]— append-only journal, immutable receipts.[crates/muser-cluster/src/receiver.rs:108-150]— ledger-volume gate.[scripts/gx10/durable_fsync_probe.py:19-22]— the standalone tail probe and its exit contract.[crates/muser-server/src/state.rs:1666-1675]— the native+DFlash fail-closed serving refusal (quoted).[scripts/qualify_nvfp4_fast.py:333-336]— the qualifier’s matching variant refusal.[ledger …]— preamble; “Synthetic spec matrix deep-cell restatement” (seal_eligible); “CORRECTION — the 65536 warm-hit result”; the J0/J3 entries cited for the notarial/non-notarial distinction.- glossary — terms introduced this chapter: fail-closed, release lock, findings register, feature contract, readiness receipt, atomic seal bundle, launch-claims register, OPERATOR REVIEW, honesty tags, documentation truth pass, notarial evidence, append-only evidence volume.