risc0-ed25519-verified/TRUSTED-BASE.md

267 lines
17 KiB
Markdown
Raw Normal View History

# Trusted base
What you must believe for the theorems in this repository to transfer to the
running Rust code. Everything else is machine-checked.
1. **Lean 4 kernel** (v4.30.0-rc2) and its three foundational axioms
`[propext, Classical.choice, Quot.sound]`. Every certificate is
`#print axioms`-audited against exactly this list.
2. **mathlib** (prebuilt oleans fetched by `lake exe cache get`).
3. **Charon + Aeneas** (pinned `9dd7f23c` / `bf13c42e`): the translation
from Rust MIR to the Lean model is assumed faithful. The generated
`gen/` files are never edited (comments only); proofs are stated ABOUT them.
4. **External-function models** (`gen/*/FunsExternal.lean`): Rust items that
Aeneas cannot translate (constant-time `subtle` primitives, iterator
plumbing, formatting) are axiomatized as opaque symbols. The axiom audit
proves none of these axioms enters the dependency cone of any certificate,
except where a model is explicitly listed below.
5. **The signature-apex boundary (signature layer only)**: FOUR apex-tier
certificates — `CurveFieldProofs.verify_accepts_iff` (byte apex:
accepted iff compress([s]·B [k]·A) = R byte-for-byte),
`verify_accepts_iff_point` (half-lift: R is the canonical encoding of
the recomputed point), `verify_accepts_iff_point_eq` (point equation:
canonically-encoded Q accepted iff Q equals the recomputed point), and
`verify_accepts_iff_decompress` (full lift: R decompresses to a valid
on-curve point that equals the recomputed point) — are each
`#print axioms`-audited by check.sh Phase 3b against EXACTLY the
standard three plus this documented set, and the build fails on any
deviation:
`ed25519.Signature` (wire-format type), the single SHA-512 oracle
`verifying.sha512_hash3` (semantically `Sha512(R ‖ A ‖ msg)`),
`ed25519.Signature.to_bytes`, and `signature.error.Error`/`Error.new`
(opaque error type). The hash is an oracle with no algebraic properties
assumed — the theorems hold for whatever bytes it produces; the SHA-512
implementation itself is NOT verified. Zero curve, scalar, or backend
axioms are in any of the four cones. The constructive decompress theorem underneath the full lift
(`decompress_of_canonical`) carries the standard three axioms ONLY.
P2-a: attack the arithmetic/apex tier boundary itself selftest-tiers.sh tests the property this repository exists to assert and that nothing had tested: the arithmetic tier rests on the three kernel axioms and NOTHING else. Five cases, all green. control both tiers pass on the untouched tree case 1 an apex axiom injected into an ARITHMETIC certificate's proof, statement untouched so only the cone moves -> Phase 3 rejects case 2 apex boundary widened by one name -> Phase 3 rejects case 3 apex boundary narrowed by one name -> Phase 3 rejects restored both tiers pass again Which axiom to inject is derived per fork from this repo's own documented boundary intersected with the victim module's import closure; no name is hard-coded, so the same script ships unchanged in all four forks. Two defects in the test were found and fixed before it was trusted. The lifted driver first omitted `set -euo pipefail`: the phase's Lean work runs in a subshell and the phase ends in a bare `echo ""`, so without -e the subshell's exit 1 was masked and the driver reported GREEN while printing APEX AUDIT FAILED. And a line-count sanity check passed an empty driver because the CERTS array padded it; the guard now looks for the diagnostics it means to provoke. Both are recorded in the script's comments. The test restores what it touches and rebuilds the module it edits, so it leaves the tree exactly as it found it. New executable is pinned in HARNESS.sha256 (Phase 0c fails closed on an unpinned one). Certified by a full sweep: both buttons, all four forks, from purged trees, machine otherwise idle. 8/8 green, 0 errors.
2026-07-31 00:39:34 +00:00
That two-tier separation is enforced, not merely observed. Phase 3
requires every arithmetic certificate's cone to be exactly the three
kernel axioms, and every apex cone to equal the documented set above
exactly. `selftest-tiers.sh` attacks it from both sides: it injects one
of the axioms above into an arithmetic certificate's *proof*, leaving the
statement untouched so that only the cone moves, and it shifts the
documented apex boundary by one name in each direction. All three must be
rejected, and are. Before those cases existed nothing in the harness
distinguished "this tier needs no hash oracle" from "this tier happens
not to use one today".
6. **`Scalar52::sub::black_box` (scalar layer)**: this fork's v4.1.3 code
implements the constant-time conditional via a local `black_box` =
`unsafe { core::ptr::read_volatile(&value) }`. The volatile read is an
optimization fence whose VALUE semantics is the identity; it is modeled as
`id` in `gen/CurveField/FunsExternal.lean` (merged gen). (Upstream v5 uses `subtle`
here; betrusted v4.1.2 uses a pure arithmetic mask — each fork is verified
against its own strategy.)
7. **Compilation of Rust to machine code** (rustc backend) is out of scope,
as is side-channel behaviour (timing, speculation). The proofs are about
functional correctness at the MIR/LLBC level.
8. **What the axiom gate binds, and what it does not.** `check.sh` Phase 2b
reads every compiled `Proofs/*.olean` and fails the build if any
declaration there is an axiom. It asks the kernel rather than parsing
source text, because the source-text check in Phase 1 is evadable four
ways — an indented `axiom`, `@[simp] axiom`, `unsafe axiom`, and `axiom`
with the name on the following line all compile and all miss its pattern.
Membership self-derives from the filesystem, so `Scalar*` and `AxiomCheck`
are covered as well, and the count of compiled modules must equal the
count of shipped sources, so a deleted `.olean` cannot make the scan pass
vacuously. `selftest-axgate.sh` attacks the shipping gate rather than a
copy of it, and was itself negative-tested by removing the gate's error.
P2-a': can a declaration hide from the inventory walker? Phase 2c exists because a source-regex enumerator proved evadable in ltl-accumulator-verified: attributed, private and `instance` declarations and a nested-namespace basename collision all slipped past it. The fix was to stop reading source text and ask the Lean environment, and that fix was ported here. But a fix ported is not a fix tested. selftest-inventory.sh proves the GATE reacts to a difference; it feeds synthetic observations and never runs the walker. Nothing here had ever asked whether the WALKER SEES a declaration written in an evasive shape. selftest-shapes.sh adds all four shapes to an audited module, recompiles it, runs the real Phase 2c, and requires each one to be NAMED in the UNCLASSIFIED list. Asserting that the gate merely failed would not do: one shape surfacing fails the run while the other three ride along unseen. All four forks report all four. Negative-tested by removing the injection — the run then reports the walker blind and fails. The victim module is derived from each repo's own manifest, not named: the forks do not share a corpus (dalek/anza attack Proofs.Basic, risc0/betrusted Proofs.DecompressMain), and a hard-coded name would have silently found nothing on half of them. It must be manifested, must not be an inventory or audit driver, and must be imported by no other manifest module. Two notes for whoever edits this next. When re-deriving a leaf module, the inventory drivers must be excluded from the set of IMPORTERS as well as from the candidates: they import the whole corpus, so leaving them in makes every module look imported, finds no leaf, and the test silently has no victim at all. And a lift of Phase 2c needs SCALAR_SH/SCALAR_MANIFEST alongside PROOFS, or the coverage check dies on an unbound variable. New executable pinned in HARNESS.sha256. Certified by a full sweep: both buttons, all four forks, purged trees, machine otherwise idle. 8/8 green.
2026-07-31 09:56:05 +00:00
`selftest-shapes.sh` asks the companion question about Phase 2c: can a
declaration HIDE from the walker? It adds four shapes to an audited module
`@[simp]`, `private`, an `instance`, and a nested namespace reusing an
audited basename — and requires the walker to report every one of them by
name, not merely to fail. Those four shapes are the ones that defeated a
source-regex enumerator in ltl-accumulator-verified and caused Phase 2c to
be written against the Lean environment instead; until 2026-07-31 the fix
was ported here but never re-attacked. It too was negative-tested, by
removing the injection and confirming the run then reports the walker blind.
**The residue you must still supply yourself:** this binds *declarations*,
not *statements*. Nothing in the button establishes that a certificate's
theorem says what its name — or this document — suggests it says. A
theorem gutted to a tautology with the same axiom cone would pass every
phase. Reading the statements remains a human act.
verification: bind the statements, the specifications, and the model (P1-a) Phases 3/3b establish what each certificate RESTS ON. Neither says what it SAYS, nor what it is ABOUT. A certificate gutted to a tautology of the same axiom cone passes both; so does one whose reference definition has been redefined to BE the extracted code, at which point the theorem reads `loop = loop` and every cone is byte-identical. Phase 3c closes that. Proofs/Audit.lean emits a canonical block holding the policy constants, every certificate's fully-elaborated statement (pp.all, so implicit arguments, instances and universe levels are visible), and the body of every specification constant transitively reachable from those statements. Its SHA-256 is pinned in check.sh and the block itself is committed as AUDIT-MANIFEST.txt, so a mismatch is DIFFED, not merely reported. 31 certificates, 68 specification constants per repository. Two tiers, not one. These forks have an arithmetic tier that must stay oracle-free and an apex tier carrying this fork's hash and wire-format axioms, and the apex boundary genuinely differs per fork (dalek 8 extra names, anza 4, risc0 and betrusted 5). One shared constant would have widened the arithmetic tier to accept hash oracles, which is the most valuable property these repos have. Each auditor is generated from its own repository's policy. Phase 0b pins the extracted model. This was not a precaution: risc0 and betrusted were observed emitting BYTE-IDENTICAL audit-manifest digests (6c821b8e…) while shipping demonstrably different extracted models, their point-doubling routines differing in operation order. A statement names an extracted function; it does not contain that function's body. Binding statements is not binding the subject. Membership derives from the filesystem, so a new model file fails closed. selftest-statements.sh attacks both phases with ten cases, each asserting a specific diagnostic: an edited model body, an unlisted model file, a widened policy, a hand-edited committed block, a certificate dropped from the auditor WITH the digest refreshed to match, and a gutted statement whose cone is unchanged. It lifts the phases out of check.sh at run time, so it attacks the shipping gate rather than a copy. Two bugs found and fixed during that testing, both mine: Phase 3c read `$0` after `cd "$AENEAS_LEAN"`, and $0 is the caller's relative path; and the axgate self-test compared the tree against a pristine checkout rather than against how it found it. A third expectation was wrong rather than the code — widening the apex boundary is caught by the exact-cone requirement before the digest ever runs, which is a stronger rejection, and the test now says so. All sixteen runs green at these commits: four main buttons, four axgate self-tests, four binding self-tests, four scalar buttons. TRUSTED-BASE.md records what this binds and, at equal length, what it does not: a digest binds identity, not meaning; an author can rotate the pins in one commit and is caught by review, not by the script; and pinning the model says nothing about whether Charon and Aeneas translated the Rust faithfully. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-28 22:38:20 +00:00
9. **What the statement binding covers.** `check.sh` Phase 3c compiles
`Proofs/Audit.lean`, which emits a canonical block containing the policy
constants, every certificate's fully-elaborated statement (`pp.all`, so
implicit arguments, instances and universe levels are all visible), and the
fully-elaborated body of every specification constant transitively
reachable from those statements. The SHA-256 of that block is pinned in
`check.sh` and the block itself is committed as `AUDIT-MANIFEST.txt`, so a
mismatch is diffed rather than merely reported. This is what makes a
certificate gutted to a tautology of the same axiom cone fail, and what
makes a reference definition redefined to BE the extracted code fail — two
attacks that move no cone at all. Phase 0b separately pins the bytes of
every extracted-model file under `gen/`, with membership derived from the
filesystem so a new model file fails closed.
**The residue you must still supply yourself.** Three things, stated
plainly because a reader would otherwise assume them:
· *A digest binds identity, not meaning.* The audit proves the statements
are the ones that were reviewed. Whether those statements say something
worth believing about ed25519 is a question only a human reading them
answers. `AUDIT-MANIFEST.txt` is committed precisely so that reading is
possible without re-running anything.
· *An author can rotate the pins.* Editing a statement and refreshing the
digest in the same commit passes every phase. The defence is that both
changes are visible in the diff, reviewed at the pinned commit — not that
the script prevents it. No harness audits its own author.
· *Phase 0b pins the model; it does not verify the translation.* That the
bytes under `gen/` are the reviewed bytes says nothing about whether
Charon and Aeneas translated the Rust faithfully. That assumption is
item 3 above and is unchanged.
verification: pin the harness, the audit drivers and the policy files (P1-c) Every gate this repository has was executed by scripts that nothing pinned. Round-5 review of the companion SLH-DSA repository stubbed the compiler wrapper alone and its button printed ALL GREEN in 3.6 seconds over deliberately destroyed proofs; flipping two guards in the audit driver disabled every check with the digest byte-identical. Depth of checking is worth nothing if the thing doing the checking is unbound — and every gate added this week made that gap more valuable to an attacker, not less. Phase 0c requires every harness file to match HARNESS.sha256. Two design points carry the weight: - WHICH files must be pinned is POLICY and lives in check.sh, never in the map being consulted. If the required set were read from the pin file, deleting an entry would silently un-pin that file. It is instead derived from the filesystem, so a deleted entry is a set mismatch and a build failure. That is the exact defect SLH-DSA round-6 found, closed here by construction. - Membership self-derives from the executable bit: anything this script can shell out to must be pinned, so a NEW script fails closed until someone pins it deliberately. Load-bearing files that are not executable — the audit driver, the committed manifests, the policy tables — cannot be discovered that way and are listed explicitly. lean-guard is inside the set, which finally makes the standing "lean-guard stays hash-pinned" rule a property of the repository rather than a convention. selftest-harness.sh replays five cases, each asserting a specific diagnostic: an edited lean-guard, a new unpinned executable, a deleted pin entry, a missing pin file, and a positive control. It was itself negative-tested — with the hash comparison removed it goes red on exactly that case while cheerfully reporting "10 harness files match their pins". TRUSTED-BASE.md states the limit at equal length to the claim: pinning a harness from inside that harness is circular, and an author who edits a script and refreshes its pin in the same commit passes every phase. What the pin changes is that the edit can no longer be SILENT — it must appear in the diff at the commit being reviewed. A green button says "this is the apparatus that was reviewed", never "this apparatus is trustworthy". Also fixed, found by this sweep: both self-tests compared the working tree against its starting state with `diff <(echo "$VAR") <(command)`, which is asymmetric — for a clean tree the variable is empty and `echo` emits a blank line the command does not. It reported a difference precisely when nothing was wrong, and only surfaced once P1-a was committed and Proofs/ became clean. Both now compare as strings. Verified green: 20 runs across the four ed25519 repositories (four buttons, four harness self-tests, four axiom-gate self-tests, four binding self-tests, four scalar buttons), zero red. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-29 18:13:00 +00:00
10. **The harness is pinned, and what that is worth.** `check.sh` Phase 0c
requires every executable file under `verification/` — plus the audit
driver, the committed manifests and the policy tables, which are not
executable and are therefore listed explicitly in the script — to match
`HARNESS.sha256`. Membership is derived from the executable bit, so a new
script fails the build until someone pins it deliberately, and the required
set is computed from the filesystem rather than read out of the pin file,
so deleting an entry is a failure rather than a silent un-pinning.
`lean-guard` is inside that set: stubbing the memory-capped compiler
wrapper is the cheapest known route to a false green, demonstrated
elsewhere in this estate as ALL GREEN in 3.6 seconds over deliberately
destroyed proofs. `selftest-harness.sh` replays that attack and four
others.
**What it does NOT buy, stated plainly.** Pinning a harness from inside
that harness is circular, and no amount of engineering removes the
circularity. An author who edits `check.sh` — or `lean-guard`, or the
audit driver — and refreshes its pin in the SAME commit passes every
phase. What the pin changes is that the edit can no longer be silent: it
must appear in the diff, at the commit you are reviewing. That is why the
consumer's protection is, and has always been, *review at the pinned
commit* rather than the button's own verdict. A green button says "this is
the apparatus that was reviewed", never "this apparatus is trustworthy".
verification: pin the whole declaration surface (P1-b) Phase 2b asks the kernel whether any AXIOM is declared under Proofs/. Phase 3 pins the cones of the named certificates. Between them sat every other declaration in the corpus — around three thousand of them — and a helper lemma quietly acquiring a hash oracle in its cone moved nothing either phase looked at. Phase 2c closes that. Ported from ltl-accumulator-verified, where a nine-attack self-test proved a source-regex enumerator evadable by attributed, private, indented and `instance` declarations and by a nested-namespace basename collision. Reading the compiled environment sees what the kernel saw; no name shape hides. Every constant contributes module, name, kind and full axiom cone, and the observed set must equal inventory-allowlist.txt exactly in BOTH directions, with a count trailer so a truncated run cannot pass as an empty diff. FOUR THINGS THIS BUILD GOT WRONG, each caught by a check rather than by review: - The number of inventory drivers is a per-repo FACT, not an assumption. dalek and anza cannot import their corpus as one environment (Proofs.Basic and Proofs.ConstSpecs both declare CurveFieldProofs.zero_spec); risc0 and betrusted have no Proofs.Basic at all. Determined by compiling a probe. check.sh now DISCOVERS its drivers from the filesystem instead of naming two, and the generator refuses to split out a module the repo lacks. - The split let one real declaration hide behind another's entry. Keyed on name alone, the two zero_specs produced byte-identical records, so 3022 declarations were covered by 3021 allowlist entries. Caught by the count trailer. Every record now carries its originating module. - The gate's success line said "single sanctioned axiom", inherited from the accumulator's policy. This corpus permits NONE. A success message describing a different rule is how an assertion stops meaning anything. - selftest-axgate.sh lifted Phase 2b with a range ending at "Phase 3", so inserting Phase 2c between them made it swallow the new phase and die on variables only check.sh defines — surfacing as the BASELINE case failing, a self-test blaming a gate for its own extraction bug. Both self-tests now stop at the next phase marker whatever it is called, and refuse to run if they capture more than one phase. The guard is the fix; the range was the symptom. WHAT THIS IS NOT, recorded in TRUSTED-BASE.md at the same length as the claim: - No independent cone walker. The accumulator cross-checks collectAxioms against a hand-written walker. Ported here it was wrong in BOTH directions on mathlib's inductive shapes: EdPoint gave [] against the kernel's three axioms, and once extended, ProjPoint gave three against the kernel's none. Two implementations disagreeing both ways are a second wrong answer, not a check. These cones rest on collectAxioms alone. - Thirteen Proofs/Scalar* modules are inventoried by nothing — the second-button seam, still open. Phase 2c names every uncovered module on every run so the omission is visible rather than inferred. selftest-inventory.sh exercises the shipping gate with six cases, each asserting a specific diagnostic, including the one that matters: a cone widened by one oracle while name, module and kind stay put. Negative-tested by disabling the gate's diff, which turns two cases red including one for the wrong reason, correctly reported as such. Verified green: 20 runs across the four repositories (four buttons, four harness, four inventory, four axgate, four binding self-tests), zero red. The four check-scalar.sh greens from the preceding sweep stand: that script neither reads the pin file nor changed. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-29 23:20:19 +00:00
11. **The whole declaration surface is pinned, not just the certificates.**
`check.sh` Phase 2c reads the compiled environment and records, for EVERY
constant originating in an audited module — roughly three thousand of them,
compiler-generated auxiliaries included — its originating module, fully
qualified name, declaration kind and complete axiom cone. The observed set
must equal `inventory-allowlist.txt` exactly, in BOTH directions: a
declaration present but not allowlisted (`UNCLASSIFIED`) and an allowlist
entry with no declaration (`STALE`) are both build failures, and a count
trailer disagreeing with the lines received is a third.
The gap this closes: Phase 2b asks the kernel only whether an AXIOM is
declared, and Phase 3 pins the cones of the named certificates. Between
them sat every helper lemma in the corpus. One of those quietly acquiring
a hash oracle in its cone moved nothing either phase looked at.
**Two limits, stated because a reader would otherwise assume neither.**
· *No independent cone walker here.* The companion accumulator runs a
hand-written closure walker alongside the kernel's `collectAxioms` and
requires the two to agree on every constant, so each checks the other.
Ported to this corpus on 2026-07-29 that walker was wrong in BOTH
directions on mathlib's inductive shapes — it reported no axioms for
`CurveFieldProofs.EdPoint` where the kernel reported three, and after
being extended it reported three for `CurveFieldProofs.ProjPoint` where
the kernel reported none. Two implementations disagreeing both ways are
not a cross-check; they are a second wrong answer. The cone figures here
therefore rest on `collectAxioms` alone. The accumulator keeps its
cross-check, its corpus being mathlib-free.
· *The scalar layer is outside this phase.* Thirteen `Proofs/Scalar*`
modules belong to `check-scalar.sh` and are inventoried by nothing. That
is the two-button seam, still open. Phase 2c prints every uncovered
module by name on every run, so the omission is visible rather than
inferred.
verification: close the two-button seam and level up the scalar button (P0-b) THE SEAM. This repository is checked by two scripts, and until now neither asserted anything about the other's scope. check.sh's dead-file gate simply SKIPPED anything named Scalar*, so a new Proofs/ScalarX.lean was gated by nothing at all: absent from one manifest by exemption, from the other by omission, compiled by neither, inventoried by neither. Each button now reads the other's manifest and requires every shipped proof source to belong to EXACTLY ONE of them — neither orphaned nor double-claimed, both directions, plus a phantom check on entries naming files that do not exist. Negative-tested four ways, including the exact hole this item names. THE SCALAR BUTTON. Closing the seam exposed it as the estate's weakest link, having been left behind by every hardening round while the main button gained five phases. 45 lines to 227: - source-integrity check over its sources; - harness-pin verification, so running THIS button alone is protected and not only running it after check.sh; - a kernel-side axiom-declaration gate over the compiled artifacts, replacing a source-text grep that is evadable four ways on v4.30.0-rc2; - a declaration inventory of ~1880 constants against its own allowlist, diffed both directions with a count trailer. These 13 modules were the only part of the proof corpus with no inventory: check.sh Phase 2c named them as uncovered on every run, and now names the button that covers them instead; - per-certificate exact-cone assertions replacing `-eq 13` over matching output lines. A count cannot say WHICH certificate is clean and passes just as happily if one cone is reported twice. Every fork-specific fact was read from the existing script rather than assumed: risc0 and betrusted audit sub_loop1_one_spec where dalek and anza audit cond_add_l_one_spec, untouched. THREE BUGS, ONE ROOT CAUSE, all found by the gates rather than by review. Each reasoned about how a thing is SPELLED instead of what it BELONGS TO, and the corpus punished each: Proofs/ScalarPackSpec.lean is named like the scalar layer and owned by the main button. - the scalar dead-file gate globbed Scalar* and demanded ScalarPackSpec be scalar-owned. REMOVED rather than special-cased: the seam check tests membership in exactly one manifest, which is strictly stronger than any prefix; - the scalar axiom gate scanned Scalar*.olean, reporting "14 modules" for a 13-module manifest. On a tree where check.sh had not run that artifact is absent and the button would have failed for a false reason. It now scans the manifest by membership and fails closed on a missing artifact; - Phase 2c's driver discovery globbed Inventory*.lean and claimed the other button's driver, then correctly complained its own manifest lacked those modules. This is the family the campaign began with: a source-text axiom grep reasoning about spelling. Recorded in TRUSTED-BASE.md because it generalises. Also fixed: the first negative test of the scalar gate's absence check passed for the wrong reason — the button recompiles before the gate runs, so removing an artifact merely caused it to be rebuilt. Retested against the lifted phase, where absence is a persistent condition. Verified green: 24 runs across the four repositories — four main buttons, four scalar buttons, and sixteen self-tests — zero red. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-30 10:30:29 +00:00
12. **Both buttons, and the seam between them.** This repository is checked by
two scripts: `check.sh` covers the field, curve and signature layers,
`check-scalar.sh` the scalar layer. Until 2026-07-30 neither asserted
anything about the other's scope and the main one simply SKIPPED anything
named `Scalar*`, so a new `Proofs/ScalarX.lean` was gated by nothing at all
— absent from one manifest by exemption, from the other by omission,
compiled by neither, inventoried by neither. Each button now reads the
other's manifest and requires every shipped proof source to belong to
EXACTLY ONE of them, both directions: neither orphaned nor double-claimed.
The scalar button was also brought to the main one's standard, having been
left behind by every hardening round: it now checks source integrity,
verifies the harness pins (so running it alone is protected too), asks the
KERNEL about axiom declarations instead of grepping source text, inventories
its ~1,880 declarations against its own allowlist, and asserts each of its
13 certificates by name rather than counting how many lines of output
matched. A count cannot say WHICH certificate is clean, and passes just as
happily if one cone is reported twice.
**What this cost, recorded because the lesson generalises.** Three separate
gates in this work reasoned about how a thing is SPELLED rather than what it
BELONGS TO, and the corpus punished each one: `Proofs/ScalarPackSpec.lean`
is named like the scalar layer and owned by the main button. A dead-file
gate globbing `Scalar*` demanded it be scalar-owned; an axiom gate scanning
`Scalar*.olean` swept in an artifact this button does not compile, which on
a fresh tree is absent and would have failed the run for a false reason; and
the inventory driver discovery globbing `Inventory*.lean` claimed the other
button's driver. All three now test membership in a manifest. This is the
same family as the source-text axiom grep that began this campaign:
reasoning about names instead of about the thing itself.
verification: --audit-only mode, and the guard that keeps it from becoming evidence (T1) Gate work dominates this estate's wall-clock: on 2026-07-29, 3.9 hours of a session went to Lean re-elaborating proofs nobody had edited while the audit phases themselves took about fifteen seconds. --audit-only runs every gate against the artifacts a previous full run left behind: ~60s against ~1280s. IT IS SAFE ONLY BECAUSE IT REFUSES. - It requires every shipped .lean to be BYTE-IDENTICAL to a basis recorded by a previous full run. Not mtimes: `touch` defeats those, and a stale-artifact check that fails open is worse than no shortcut at all, because a green button would then describe a corpus that is no longer on disk. - The basis is gitignored build state, so a fresh clone cannot inherit permission to skip compiling. - The closing banner differs and says in words that the run is not evidence. selftest-auditonly.sh exercises seven cases: no basis, an edited comment character, a deleted source, a new source, a missing artifact, a truncated basis, and — asserted as a PASS — every source's mtime touched with bytes unchanged, which pins the bytes-not-mtimes decision rather than leaving it implicit. Negative-tested: with the basis comparison disabled a changed source is wrongly accepted, exit 0 and zero refusals, so the guard is load-bearing. A PHASE TERMINATOR, because this broke twice. Every self-test lifts a phase from check.sh by scanning to the next phase marker. The last phase had no marker after it, so a lift ran to end-of-file and swallowed whatever was appended later — first Phase 2c into the axgate lift, then T1's tail into the binding lift, where it referenced $AUDIT_ONLY and died under `set -u`. Both surfaced as the BASELINE case failing: a self-test blaming a gate for its own extraction bug. The phases now end at an explicit sentinel and both lifters stop there, so nothing appended below can silently become part of the last phase from a lifter's point of view. TRUSTED-BASE.md records what an audit-only transcript does and does not establish, and — because it cost a confusing red run today — that lean-guard's memory clamp presents as `FAIL: Proofs/<module>` while being a resource condition, not a broken proof. Verified green: 8 full button runs (four check.sh, four check-scalar.sh) and 20 self-tests across the four repositories, zero red. One earlier run failed on the memory clamp because the author ran a test suite concurrently; re-run on a quiet machine, green. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-30 17:16:20 +00:00
13. **`--audit-only`, and why a green transcript from it is not evidence.**
`check.sh --audit-only` runs every gate but skips recompilation, against the
`.olean` files a previous full run left behind: about 60 seconds against
about 1280. It exists because gate work dominates this estate's wall-clock,
and it is safe only because it refuses.
It refuses unless every shipped `.lean` is BYTE-IDENTICAL to a basis
recorded by a previous full run — not mtimes, which `touch` defeats, and a
stale-artifact check that fails open would be worse than no shortcut at all:
a green button would then describe a corpus that is no longer on disk. The
basis is gitignored build state, so a fresh clone cannot inherit permission
to skip compiling, and the closing banner says in words that the run is not
evidence.
**The kernel re-elaborates nothing in such a run.** What it establishes is
that the gates accept artifacts produced earlier — useful while developing a
gate, worthless as a record. `formal-verification-control/tools/record-run.py`
enforces that: it refuses to archive any transcript bearing the audit-only
markers, and also any transcript without a terminal success banner, any
containing `error:`, and any repository whose tree is dirty at record time.
A banner is a request; that tool is the gate.
Also worth knowing when reading a red run: `lean-guard` clamps Lean's memory
budget to what the machine can spare, and under load that clamp can be too
small to elaborate a large module. It surfaces as `FAIL: Proofs/<module>`,
which reads exactly like a broken proof and is not one — it is a resource
condition, and the cap is what protects this machine from the global OOM
that killed a session on 2026-07-02. Check the transcript for a `clamping`
line before concluding anything about the mathematics.
verification: build hygiene, and the hidden dependency it exposed (P0-a) Phase 0a purges every .olean before compiling, bans stray Lean files at the verification root (LEAN_PATH contains $PWD, so they join the build unaudited), and requires gen/ to be exactly the model manifest plus its pinned templates. The templates are KEPT, unlike SLH-DSA which deletes them: extract.sh directs the operator to diff the hand-written external models against them, so they are the reference for that comparison and P2-c will enforce it. The purge is skipped under --audit-only, which exists to audit the artifacts a previous full run produced. Those two features would otherwise destroy each other, and it is a further reason an audit-only transcript is not evidence: it has not had this hygiene applied. WHAT THE PURGE EXPOSED, and it is the point of the whole item: This button had never compiled the corpus from nothing. The signature apex rests on scalar arithmetic — PointLiftSpec -> ScalarPackSpec -> ScalarFromBytesSpec, and SigApexSpec -> ScalarDenote — and TWELVE of the scalar layer's thirteen modules are transitive prerequisites of this manifest. They were never compiled here. The button worked because check-scalar.sh had run at some earlier point and left its .olean files behind. .olean is gitignored, so no git status could ever have shown that the verdict rested on untracked artifacts produced by a different script. Nothing about the proofs was wrong. The evidence was resting on something invisible, for the entire life of these repositories, and it surfaced the moment something finally cleaned up before verifying. Those twelve are now compiled here as PREREQ — BORROWED, NOT OWNED. check-scalar.sh still audits them; Phase 1b asserts every borrowed name belongs to the other manifest and to neither twice, so the list cannot become a second ownership claim. Two consequences fixed along the way, both the spelling-versus-membership error that ScalarPackSpec has now taught four times: - Phase 2b globbed Proofs/*.olean and would have demanded artifacts this button never builds. It now scans its manifest by membership and fails closed on a missing one. - The three inventory drivers were exempted from the dead-file gate and compiled in a later phase; after a purge they were absent when Phase 2b ran. They are now in the manifest like everything else, and three exemptions are gone. The sweep runner now reports RESOURCE rather than RED when it sees a memory_exception: lean-guard's clamp is not a broken proof, and it has misled the operator once and the author once. Verified green: 8 full runs from completely purged trees — four check.sh, four check-scalar.sh — zero red, zero resource. Every artifact rebuilt from committed source. These are the first runs in this repository's history whose verdict provably depends on nothing but the bytes in git. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-30 20:54:05 +00:00
14. **The verdict depends on committed bytes, not on build state — and what
proving that revealed.** `check.sh` Phase 0a purges every `.olean` under
`verification/` before compiling, forbids stray Lean files at the
verification root (they join the build through `LEAN_PATH`, which contains
`$PWD`), and requires `gen/` to be exactly the model manifest plus its
pinned Aeneas templates. The templates are KEPT here, unlike the companion
SLH-DSA repository which deletes them: `extract.sh` directs the operator to
diff the hand-written external models against them, so they are the
reference for that comparison. The purge does not run under `--audit-only`,
which exists to audit the artifacts a previous full run produced; that is a
further reason an audit-only transcript is not evidence.
**What the purge exposed, on 2026-07-30.** This button had never in its life
compiled the corpus from nothing. The signature apex rests on scalar
arithmetic — `PointLiftSpec``ScalarPackSpec``ScalarFromBytesSpec`, and
`SigApexSpec``ScalarDenote` — and TWELVE of the scalar layer's thirteen
modules are transitive prerequisites of this manifest. They were never
compiled here. The button worked because `check-scalar.sh` had run at some
earlier point and left its `.olean` files behind, and `.olean` is gitignored,
so no `git status` could ever have shown a reader that the verdict rested on
untracked artifacts produced by a different script. Nothing about the proofs
was wrong; the *evidence* was resting on something invisible.
Those twelve are now compiled here as `PREREQ`**borrowed, not owned**.
`check-scalar.sh` still audits them: their cones, their declaration
inventory, their axiom gate. Phase 1b asserts that every borrowed name
belongs to the other manifest and to neither twice, so the list cannot
quietly become a second claim of ownership.
The general lesson, which is why the purge is worth its minutes: a
verification that never cleans up cannot distinguish "these proofs check"
from "these proofs check given whatever happens to be lying around".