mirror of
https://github.com/saymrwulf/anza-ed25519-verified.git
synced 2026-09-04 20:24:06 +00:00
4 commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
| a8998dd94b |
verification: pin the whole declaration surface (P1-b)
Phase 2b asks the kernel whether any AXIOM is declared under Proofs/. Phase 3
pins the cones of the named certificates. Between them sat every other
declaration in the corpus — around three thousand of them — and a helper lemma
quietly acquiring a hash oracle in its cone moved nothing either phase looked
at.
Phase 2c closes that. Ported from ltl-accumulator-verified, where a nine-attack
self-test proved a source-regex enumerator evadable by attributed, private,
indented and `instance` declarations and by a nested-namespace basename
collision. Reading the compiled environment sees what the kernel saw; no name
shape hides. Every constant contributes module, name, kind and full axiom cone,
and the observed set must equal inventory-allowlist.txt exactly in BOTH
directions, with a count trailer so a truncated run cannot pass as an empty
diff.
FOUR THINGS THIS BUILD GOT WRONG, each caught by a check rather than by review:
- The number of inventory drivers is a per-repo FACT, not an assumption.
dalek and anza cannot import their corpus as one environment (Proofs.Basic
and Proofs.ConstSpecs both declare CurveFieldProofs.zero_spec); risc0 and
betrusted have no Proofs.Basic at all. Determined by compiling a probe.
check.sh now DISCOVERS its drivers from the filesystem instead of naming
two, and the generator refuses to split out a module the repo lacks.
- The split let one real declaration hide behind another's entry. Keyed on
name alone, the two zero_specs produced byte-identical records, so 3022
declarations were covered by 3021 allowlist entries. Caught by the count
trailer. Every record now carries its originating module.
- The gate's success line said "single sanctioned axiom", inherited from the
accumulator's policy. This corpus permits NONE. A success message
describing a different rule is how an assertion stops meaning anything.
- selftest-axgate.sh lifted Phase 2b with a range ending at "Phase 3", so
inserting Phase 2c between them made it swallow the new phase and die on
variables only check.sh defines — surfacing as the BASELINE case failing,
a self-test blaming a gate for its own extraction bug. Both self-tests now
stop at the next phase marker whatever it is called, and refuse to run if
they capture more than one phase. The guard is the fix; the range was the
symptom.
WHAT THIS IS NOT, recorded in TRUSTED-BASE.md at the same length as the claim:
- No independent cone walker. The accumulator cross-checks collectAxioms
against a hand-written walker. Ported here it was wrong in BOTH directions
on mathlib's inductive shapes: EdPoint gave [] against the kernel's three
axioms, and once extended, ProjPoint gave three against the kernel's none.
Two implementations disagreeing both ways are a second wrong answer, not a
check. These cones rest on collectAxioms alone.
- Thirteen Proofs/Scalar* modules are inventoried by nothing — the
second-button seam, still open. Phase 2c names every uncovered module on
every run so the omission is visible rather than inferred.
selftest-inventory.sh exercises the shipping gate with six cases, each
asserting a specific diagnostic, including the one that matters: a cone
widened by one oracle while name, module and kind stay put. Negative-tested by
disabling the gate's diff, which turns two cases red including one for the
wrong reason, correctly reported as such.
Verified green: 20 runs across the four repositories (four buttons, four
harness, four inventory, four axgate, four binding self-tests), zero red. The
four check-scalar.sh greens from the preceding sweep stand: that script neither
reads the pin file nor changed.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
|
|||
| b3f4ee0d33 |
verification: pin the harness, the audit drivers and the policy files (P1-c)
Every gate this repository has was executed by scripts that nothing pinned.
Round-5 review of the companion SLH-DSA repository stubbed the compiler
wrapper alone and its button printed ALL GREEN in 3.6 seconds over
deliberately destroyed proofs; flipping two guards in the audit driver
disabled every check with the digest byte-identical. Depth of checking is
worth nothing if the thing doing the checking is unbound — and every gate
added this week made that gap more valuable to an attacker, not less.
Phase 0c requires every harness file to match HARNESS.sha256. Two design
points carry the weight:
- WHICH files must be pinned is POLICY and lives in check.sh, never in the
map being consulted. If the required set were read from the pin file,
deleting an entry would silently un-pin that file. It is instead derived
from the filesystem, so a deleted entry is a set mismatch and a build
failure. That is the exact defect SLH-DSA round-6 found, closed here by
construction.
- Membership self-derives from the executable bit: anything this script can
shell out to must be pinned, so a NEW script fails closed until someone
pins it deliberately. Load-bearing files that are not executable — the
audit driver, the committed manifests, the policy tables — cannot be
discovered that way and are listed explicitly.
lean-guard is inside the set, which finally makes the standing "lean-guard
stays hash-pinned" rule a property of the repository rather than a convention.
selftest-harness.sh replays five cases, each asserting a specific diagnostic:
an edited lean-guard, a new unpinned executable, a deleted pin entry, a
missing pin file, and a positive control. It was itself negative-tested — with
the hash comparison removed it goes red on exactly that case while cheerfully
reporting "10 harness files match their pins".
TRUSTED-BASE.md states the limit at equal length to the claim: pinning a
harness from inside that harness is circular, and an author who edits a script
and refreshes its pin in the same commit passes every phase. What the pin
changes is that the edit can no longer be SILENT — it must appear in the diff
at the commit being reviewed. A green button says "this is the apparatus that
was reviewed", never "this apparatus is trustworthy".
Also fixed, found by this sweep: both self-tests compared the working tree
against its starting state with `diff <(echo "$VAR") <(command)`, which is
asymmetric — for a clean tree the variable is empty and `echo` emits a blank
line the command does not. It reported a difference precisely when nothing was
wrong, and only surfaced once P1-a was committed and Proofs/ became clean.
Both now compare as strings.
Verified green: 20 runs across the four ed25519 repositories (four buttons,
four harness self-tests, four axiom-gate self-tests, four binding self-tests,
four scalar buttons), zero red.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
|
|||
| 21f80d29eb |
verification: bind the statements, the specifications, and the model (P1-a)
Phases 3/3b establish what each certificate RESTS ON. Neither says what it SAYS, nor what it is ABOUT. A certificate gutted to a tautology of the same axiom cone passes both; so does one whose reference definition has been redefined to BE the extracted code, at which point the theorem reads `loop = loop` and every cone is byte-identical. Phase 3c closes that. Proofs/Audit.lean emits a canonical block holding the policy constants, every certificate's fully-elaborated statement (pp.all, so implicit arguments, instances and universe levels are visible), and the body of every specification constant transitively reachable from those statements. Its SHA-256 is pinned in check.sh and the block itself is committed as AUDIT-MANIFEST.txt, so a mismatch is DIFFED, not merely reported. 31 certificates, 68 specification constants per repository. Two tiers, not one. These forks have an arithmetic tier that must stay oracle-free and an apex tier carrying this fork's hash and wire-format axioms, and the apex boundary genuinely differs per fork (dalek 8 extra names, anza 4, risc0 and betrusted 5). One shared constant would have widened the arithmetic tier to accept hash oracles, which is the most valuable property these repos have. Each auditor is generated from its own repository's policy. Phase 0b pins the extracted model. This was not a precaution: risc0 and betrusted were observed emitting BYTE-IDENTICAL audit-manifest digests (6c821b8e…) while shipping demonstrably different extracted models, their point-doubling routines differing in operation order. A statement names an extracted function; it does not contain that function's body. Binding statements is not binding the subject. Membership derives from the filesystem, so a new model file fails closed. selftest-statements.sh attacks both phases with ten cases, each asserting a specific diagnostic: an edited model body, an unlisted model file, a widened policy, a hand-edited committed block, a certificate dropped from the auditor WITH the digest refreshed to match, and a gutted statement whose cone is unchanged. It lifts the phases out of check.sh at run time, so it attacks the shipping gate rather than a copy. Two bugs found and fixed during that testing, both mine: Phase 3c read `$0` after `cd "$AENEAS_LEAN"`, and $0 is the caller's relative path; and the axgate self-test compared the tree against a pristine checkout rather than against how it found it. A third expectation was wrong rather than the code — widening the apex boundary is caught by the exact-cone requirement before the digest ever runs, which is a stronger rejection, and the test now says so. All sixteen runs green at these commits: four main buttons, four axgate self-tests, four binding self-tests, four scalar buttons. TRUSTED-BASE.md records what this binds and, at equal length, what it does not: a digest binds identity, not meaning; an author can rotate the pins in one commit and is caught by review, not by the script; and pinning the model says nothing about whether Charon and Aeneas translated the Rust faithfully. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
|||
| c898ac4284 |
verification: kernel-side axiom-declaration gate (Phase 2b) + self-test
Phase 1's anti-smuggling check reads source text. Measured today on Lean
v4.30.0-rc2, four distinct declarations compile cleanly and slip past its
anchored pattern:
` axiom cheat : ...` one leading space
`@[simp] axiom cheat : ...` line starts with the attribute
`unsafe axiom cheat : ...` `unsafe` absent from the modifier list
`axiom` <newline> ` cheat` no space follows the keyword
Any of them yields a repository that proves False while the button prints
ALL GREEN. Only the tab variant is blocked, and by Lean, not by us.
Hardening the pattern would fix the exhibited syntax rather than the class,
which is the mistake this estate has made before. Phase 2b stops parsing text
and asks the kernel instead: it reads every compiled Proofs/*.olean with
readModuleData and rejects any declaration that is an axiom.
Design notes:
- reads compiled artifacts rather than importing the modules, because
Proofs.Basic and Proofs.ConstSpecs deliberately reuse `zero_spec` and a
whole-corpus import is impossible by construction;
- membership is self-deriving from the filesystem, so Scalar* and
AxiomCheck are covered too — both are skipped by the CERTS audit and by
the dead-file gate;
- fails closed on absence: a missing .olean would make the scan vacuous, so
the count of compiled modules must equal the count of shipped sources;
- removes its temp source AND artifact on both paths, since a bare `rm`
after the call never runs under `set -e` when the gate goes red — exactly
how this repo accumulated 101 orphan .olean files;
- ~3 s for the whole corpus, against ~53 s for one module-importing run.
Phase 1's grep stays as a fast first line of defence. Phase 2b is the gate
that is load-bearing.
selftest-axgate.sh attacks the shipping gate, lifted out of check.sh at run
time rather than copied. It asserts the specific diagnostic, so a rejection
for an unrelated reason fails too, and it was itself negative-tested: with
the gate's throwError removed, the self-test goes red on exactly that case.
No proof, statement, specification or certificate is touched. No attested
commit is altered — the log binds specific commit hashes, all of which remain
ancestors of HEAD.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
|