Gate work dominates this estate's wall-clock: on 2026-07-29, 3.9 hours of a
session went to Lean re-elaborating proofs nobody had edited while the audit
phases themselves took about fifteen seconds. --audit-only runs every gate
against the artifacts a previous full run left behind: ~60s against ~1280s.
IT IS SAFE ONLY BECAUSE IT REFUSES.
- It requires every shipped .lean to be BYTE-IDENTICAL to a basis recorded by
a previous full run. Not mtimes: `touch` defeats those, and a stale-artifact
check that fails open is worse than no shortcut at all, because a green
button would then describe a corpus that is no longer on disk.
- The basis is gitignored build state, so a fresh clone cannot inherit
permission to skip compiling.
- The closing banner differs and says in words that the run is not evidence.
selftest-auditonly.sh exercises seven cases: no basis, an edited comment
character, a deleted source, a new source, a missing artifact, a truncated
basis, and — asserted as a PASS — every source's mtime touched with bytes
unchanged, which pins the bytes-not-mtimes decision rather than leaving it
implicit. Negative-tested: with the basis comparison disabled a changed source
is wrongly accepted, exit 0 and zero refusals, so the guard is load-bearing.
A PHASE TERMINATOR, because this broke twice. Every self-test lifts a phase
from check.sh by scanning to the next phase marker. The last phase had no
marker after it, so a lift ran to end-of-file and swallowed whatever was
appended later — first Phase 2c into the axgate lift, then T1's tail into the
binding lift, where it referenced $AUDIT_ONLY and died under `set -u`. Both
surfaced as the BASELINE case failing: a self-test blaming a gate for its own
extraction bug. The phases now end at an explicit sentinel and both lifters
stop there, so nothing appended below can silently become part of the last
phase from a lifter's point of view.
TRUSTED-BASE.md records what an audit-only transcript does and does not
establish, and — because it cost a confusing red run today — that lean-guard's
memory clamp presents as `FAIL: Proofs/<module>` while being a resource
condition, not a broken proof.
Verified green: 8 full button runs (four check.sh, four check-scalar.sh) and 20
self-tests across the four repositories, zero red. One earlier run failed on
the memory clamp because the author ran a test suite concurrently; re-run on a
quiet machine, green.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
13 KiB
Trusted base
What you must believe for the theorems in this repository to transfer to the running Rust code. Everything else is machine-checked.
-
Lean 4 kernel (v4.30.0-rc2) and its three foundational axioms
[propext, Classical.choice, Quot.sound]. Every certificate is#print axioms-audited against exactly this list. -
mathlib (prebuilt oleans fetched by
lake exe cache get). -
Charon + Aeneas (pinned
9dd7f23c/bf13c42e): the translation from Rust MIR to the Lean model is assumed faithful. The generatedgen/files are never edited (comments only); proofs are stated ABOUT them. -
External-function models (
gen/*/FunsExternal.lean): Rust items that Aeneas cannot translate (constant-timesubtleprimitives, iterator plumbing, formatting) are axiomatized as opaque symbols. The axiom audit proves none of these axioms enters the dependency cone of any certificate, except where a model is explicitly listed below. -
The signature-apex boundary (signature layer only): FOUR apex-tier certificates —
CurveFieldProofs.verify_accepts_iff(byte apex: accepted iff compress([s]·B − [k]·A) = R byte-for-byte),verify_accepts_iff_point(half-lift: R is the canonical encoding of the recomputed point),verify_accepts_iff_point_eq(point equation: canonically-encoded Q accepted iff Q equals the recomputed point), andverify_accepts_iff_decompress(full lift: R decompresses to a valid on-curve point that equals the recomputed point) — are each#print axioms-audited by check.sh Phase 3b against EXACTLY the standard three plus this documented set, and the build fails on any deviation:ed25519.Signature(wire-format type), the three SHA-512 wrapper oraclesverifying.sha512_new/update/finalize_bytes(+ the opaquesha2.Sha512state type),ed25519.Signature.to_bytes, andsignature.error.Error/Error.new(opaque error type). The hash is an oracle with no algebraic properties assumed — the theorems hold for whatever bytes it produces; the SHA-512 implementation itself is NOT verified. Zero curve, scalar, or backend axioms are in any of the four cones. The constructive decompress theorem underneath the full lift (decompress_of_canonical) carries the standard three axioms ONLY. -
Compilation of Rust to machine code (rustc backend) is out of scope, as is side-channel behaviour (timing, speculation). The proofs are about functional correctness at the MIR/LLBC level.
-
What the axiom gate binds, and what it does not.
check.shPhase 2b reads every compiledProofs/*.oleanand fails the build if any declaration there is an axiom. It asks the kernel rather than parsing source text, because the source-text check in Phase 1 is evadable four ways — an indentedaxiom,@[simp] axiom,unsafe axiom, andaxiomwith the name on the following line all compile and all miss its pattern. Membership self-derives from the filesystem, soScalar*andAxiomCheckare covered as well, and the count of compiled modules must equal the count of shipped sources, so a deleted.oleancannot make the scan pass vacuously.selftest-axgate.shattacks the shipping gate rather than a copy of it, and was itself negative-tested by removing the gate's error. The residue you must still supply yourself: this binds declarations, not statements. Nothing in the button establishes that a certificate's theorem says what its name — or this document — suggests it says. A theorem gutted to a tautology with the same axiom cone would pass every phase. Reading the statements remains a human act. -
What the statement binding covers.
check.shPhase 3c compilesProofs/Audit.lean, which emits a canonical block containing the policy constants, every certificate's fully-elaborated statement (pp.all, so implicit arguments, instances and universe levels are all visible), and the fully-elaborated body of every specification constant transitively reachable from those statements. The SHA-256 of that block is pinned incheck.shand the block itself is committed asAUDIT-MANIFEST.txt, so a mismatch is diffed rather than merely reported. This is what makes a certificate gutted to a tautology of the same axiom cone fail, and what makes a reference definition redefined to BE the extracted code fail — two attacks that move no cone at all. Phase 0b separately pins the bytes of every extracted-model file undergen/, with membership derived from the filesystem so a new model file fails closed.The residue you must still supply yourself. Three things, stated plainly because a reader would otherwise assume them:
· A digest binds identity, not meaning. The audit proves the statements are the ones that were reviewed. Whether those statements say something worth believing about ed25519 is a question only a human reading them answers.
AUDIT-MANIFEST.txtis committed precisely so that reading is possible without re-running anything.· An author can rotate the pins. Editing a statement and refreshing the digest in the same commit passes every phase. The defence is that both changes are visible in the diff, reviewed at the pinned commit — not that the script prevents it. No harness audits its own author.
· Phase 0b pins the model; it does not verify the translation. That the bytes under
gen/are the reviewed bytes says nothing about whether Charon and Aeneas translated the Rust faithfully. That assumption is item 3 above and is unchanged. -
The harness is pinned, and what that is worth.
check.shPhase 0c requires every executable file underverification/— plus the audit driver, the committed manifests and the policy tables, which are not executable and are therefore listed explicitly in the script — to matchHARNESS.sha256. Membership is derived from the executable bit, so a new script fails the build until someone pins it deliberately, and the required set is computed from the filesystem rather than read out of the pin file, so deleting an entry is a failure rather than a silent un-pinning.lean-guardis inside that set: stubbing the memory-capped compiler wrapper is the cheapest known route to a false green, demonstrated elsewhere in this estate as ALL GREEN in 3.6 seconds over deliberately destroyed proofs.selftest-harness.shreplays that attack and four others.What it does NOT buy, stated plainly. Pinning a harness from inside that harness is circular, and no amount of engineering removes the circularity. An author who edits
check.sh— orlean-guard, or the audit driver — and refreshes its pin in the SAME commit passes every phase. What the pin changes is that the edit can no longer be silent: it must appear in the diff, at the commit you are reviewing. That is why the consumer's protection is, and has always been, review at the pinned commit rather than the button's own verdict. A green button says "this is the apparatus that was reviewed", never "this apparatus is trustworthy". -
The whole declaration surface is pinned, not just the certificates.
check.shPhase 2c reads the compiled environment and records, for EVERY constant originating in an audited module — roughly three thousand of them, compiler-generated auxiliaries included — its originating module, fully qualified name, declaration kind and complete axiom cone. The observed set must equalinventory-allowlist.txtexactly, in BOTH directions: a declaration present but not allowlisted (UNCLASSIFIED) and an allowlist entry with no declaration (STALE) are both build failures, and a count trailer disagreeing with the lines received is a third.
The gap this closes: Phase 2b asks the kernel only whether an AXIOM is declared, and Phase 3 pins the cones of the named certificates. Between them sat every helper lemma in the corpus. One of those quietly acquiring a hash oracle in its cone moved nothing either phase looked at.
Two limits, stated because a reader would otherwise assume neither.
· No independent cone walker here. The companion accumulator runs a
hand-written closure walker alongside the kernel's collectAxioms and
requires the two to agree on every constant, so each checks the other.
Ported to this corpus on 2026-07-29 that walker was wrong in BOTH
directions on mathlib's inductive shapes — it reported no axioms for
CurveFieldProofs.EdPoint where the kernel reported three, and after
being extended it reported three for CurveFieldProofs.ProjPoint where
the kernel reported none. Two implementations disagreeing both ways are
not a cross-check; they are a second wrong answer. The cone figures here
therefore rest on collectAxioms alone. The accumulator keeps its
cross-check, its corpus being mathlib-free.
· The scalar layer is outside this phase. Thirteen Proofs/Scalar*
modules belong to check-scalar.sh and are inventoried by nothing. That
is the two-button seam, still open. Phase 2c prints every uncovered
module by name on every run, so the omission is visible rather than
inferred.
- Both buttons, and the seam between them. This repository is checked by
two scripts:
check.shcovers the field, curve and signature layers,check-scalar.shthe scalar layer. Until 2026-07-30 neither asserted anything about the other's scope and the main one simply SKIPPED anything namedScalar*, so a newProofs/ScalarX.leanwas gated by nothing at all — absent from one manifest by exemption, from the other by omission, compiled by neither, inventoried by neither. Each button now reads the other's manifest and requires every shipped proof source to belong to EXACTLY ONE of them, both directions: neither orphaned nor double-claimed.
The scalar button was also brought to the main one's standard, having been left behind by every hardening round: it now checks source integrity, verifies the harness pins (so running it alone is protected too), asks the KERNEL about axiom declarations instead of grepping source text, inventories its ~1,880 declarations against its own allowlist, and asserts each of its 13 certificates by name rather than counting how many lines of output matched. A count cannot say WHICH certificate is clean, and passes just as happily if one cone is reported twice.
What this cost, recorded because the lesson generalises. Three separate
gates in this work reasoned about how a thing is SPELLED rather than what it
BELONGS TO, and the corpus punished each one: Proofs/ScalarPackSpec.lean
is named like the scalar layer and owned by the main button. A dead-file
gate globbing Scalar* demanded it be scalar-owned; an axiom gate scanning
Scalar*.olean swept in an artifact this button does not compile, which on
a fresh tree is absent and would have failed the run for a false reason; and
the inventory driver discovery globbing Inventory*.lean claimed the other
button's driver. All three now test membership in a manifest. This is the
same family as the source-text axiom grep that began this campaign:
reasoning about names instead of about the thing itself.
--audit-only, and why a green transcript from it is not evidence.check.sh --audit-onlyruns every gate but skips recompilation, against the.oleanfiles a previous full run left behind: about 60 seconds against about 1280. It exists because gate work dominates this estate's wall-clock, and it is safe only because it refuses.
It refuses unless every shipped .lean is BYTE-IDENTICAL to a basis
recorded by a previous full run — not mtimes, which touch defeats, and a
stale-artifact check that fails open would be worse than no shortcut at all:
a green button would then describe a corpus that is no longer on disk. The
basis is gitignored build state, so a fresh clone cannot inherit permission
to skip compiling, and the closing banner says in words that the run is not
evidence.
The kernel re-elaborates nothing in such a run. What it establishes is
that the gates accept artifacts produced earlier — useful while developing a
gate, worthless as a record. formal-verification-control/tools/record-run.py
enforces that: it refuses to archive any transcript bearing the audit-only
markers, and also any transcript without a terminal success banner, any
containing error:, and any repository whose tree is dirty at record time.
A banner is a request; that tool is the gate.
Also worth knowing when reading a red run: lean-guard clamps Lean's memory
budget to what the machine can spare, and under load that clamp can be too
small to elaborate a large module. It surfaces as FAIL: Proofs/<module>,
which reads exactly like a broken proof and is not one — it is a resource
condition, and the cap is what protects this machine from the global OOM
that killed a session on 2026-07-02. Check the transcript for a clamping
line before concluding anything about the mathematics.