ltl-accumulator-verified/KNOWN-GAPS.md
mrwulf 2da0a79981 Review round 4: F1* absorbed (lied-size boundary), acceptCons_sound, kit reproducibility
Round-3 verdicts: GPT-5.6 conditionally approves (blockers closed, one
portability finding); the Claude reviewer's Socratic addendum produced
F1*, the strongest finding of the series — deployed verify_consistency
and mechanized ConsRec are NOT extensionally equal. Reproduced exactly
(witness verify_consistency(1,3,R2,R3,P(2→3))=True vs ConsRec reject;
3,405 divergences n<60; strictly one-sided; power-of-two seeding
mechanism confirmed in source).

- KNOWN-GAPS gap 14: witness, mechanism, one-sidedness, and the
  pinned-pair side condition under which Theorem 3 transfers to the
  deployed verifier (pacta's pin-store flow supplies it by
  construction). No pacta code change; deployed behavior matches
  upstream RFC 9162 implementations.
- fidelity: lied-size family — 73,573 boundary cases, 3,867 expected
  divergences PINNED, one-sided direction asserted per case. Banner
  rescoped: agreement over pinned families, not extensional equality.
- Theorem3.lean: acceptCons_sound (F2) — soundness over the named
  acceptCons predicate, n₀=0 discharged from the non-prefix premise,
  size bound derived from acceptance via new consRec_some_le. Cones
  read from #print axioms; CONES/AxiomCheck/allowlist updated
  (218 → 222 constants, diff = the two theorems + two generated
  auxiliaries).
- F3/GPT§7: verification/lean-toolchain pin + run_bare.sh (reviewer's
  standalone runner, plain public lean — verified green: 61 cones, 222
  constants, gate green) + AENEAS_ENV override in check.sh and
  selftest_audit.sh.
- F4: awk field-equality replaces regex-with-dots in Phase 3b.
- F5: git-tracked .pyc removed (worse than reported — it was in the
  repo, not just the kit); __pycache__ gitignored; round-4 kit ships a
  corpus MANIFEST.sha256 + pinned commit (also GPT's governance
  condition).

check.sh exit 0, ATTESTATION GREEN; selftest exit 0, 9/9 + control.
Live LTL untouched (12 leaves, bcd15f9d…); attestation still gated on
ePrint decision + author review + explicit operator order.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-12 15:07:57 +02:00

106 lines
6.8 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# Known gaps and scope boundaries (honest ledger)
Deliberate, documented, and none silent. Reviewers should verify this
list is COMPLETE, not merely that the items are acceptable.
1. **SHA-256 is opaque** — the single boundary axiom (`LTLAcc.sha256`),
by design identical to the paper's posture: soundness theorems
construct collisions, never assume collision resistance.
2. **No consistency-completeness theorem** (honest ConsRec acceptance).
Matches the paper (its Theorem 1 is inclusion-only); honest
consistency behavior is covered by the fidelity harness's honest
cases (within the 230,016-case consistency agreement with the
deployed verifier).
3. **Lemma 2 is mechanized as specializations, not as one general
theorem.** The paper's Lemma 2 is a single statement quantified over
an abstract hash-fold `F` and a connected subtree `S`. The corpus has
no hash-fold datatype/predicate; it proves the needed instances
directly — whole-tree (`extractMTH_correct`), ConsRec
(`consRecBinding`), inclusion (`extractIncl_correct`), and the width
fact (`hnode_preimage_inj`). These suffice for Theorems 23. Two
consequences: (a) the abstract lemma itself is not a mechanized
object; (b) the path-instance receipt-uniqueness for `Root` (removed
with the vacuous `root_binding`) is not restored — optional, unused.
Any paper claim that "Lemma 2 is mechanized" must read "its
specializations sufficient for Theorems 23 are mechanized."
4. **Signature layer abstract** — Ed25519 EUF-CMA, the poison/evidence
retention state, and transferability of fork evidence (paper Prop
1(2)) are not modeled; `fork_distinct` is the Merkle-layer share only.
5. **Transliteration bridge**`fidelity/lean_defs.py` mirrors the Lean
definitions by quoted-source inspection (the Lean defs are
noncomputable over the opaque hash, so the bridge cannot be #eval'd
closed). Same inspection bridge the paper's own harness uses.
6. **Proposition 2 (verdict integrity) out of scope** — per paper §10's
mechanization list (iv). It is a property of the consumer tooling's
construction, enforced and regression-tested in the pacta repo.
7. **Multi-step pin monotonicity** — mechanized per-step
(`pinAccept_monotone`); the paper's multi-step chain is its
reflexive-transitive iterate, not separately mechanized.
8. **Process history** (candor): three cone pins were guessed (not read)
during S5.3S6 and the audit's failure went unnoticed until S7
because green was claimed from tailed output rather than the exit
code. No theorem was affected (kernel-checked throughout); pins were
corrected, the audit surface defined, and the standing rule is now:
exit code + ALL GREEN, cones read from #print axioms only.
9. **Asymptotic cost not mechanized.** Paper Theorems 2 and 3 assert the
extractors run in `O(n)` / `O(n₁)` hash evaluations. The mechanization
proves functional correctness of the named extractors only — no cost
semantics, recurrence, or computability-after-hash-instantiation. (The
extractors are `noncomputable` over the opaque `sha256`.)
10. **Pin-store initialization from the empty pin not modeled**, and
`pin_prefix_correct` assumes `0 < n`. Trust-on-first-use / the size-0
initial state is a separate operation; the theorems cover transitions
from a positive-size pin. (Related to gap 7's per-step scoping.)
11. **acceptIncl now named (was review F1).** The consumer's inclusion
acceptance `m < n ∧ Root … = some r` is now the Lean object
`acceptIncl`, with `acceptIncl_complete`/`acceptIncl_sound` routing
completeness/soundness through it, and the fidelity harness exercises
the out-of-range families (`m ≥ n`). `Root` alone still accepts
out-of-range `m`; that is by design (it is the reconstruction, not the
accept predicate).
12. **Audit-gate lineage** (candor; was round-2 GPT H1 / Claude NEW-1,
both round-1 "fail-closed" claims were overclaims). The round-2
coverage gate enumerated declarations with a source regex and was
evadable (attributes, indentation, private/protected, `instance`,
nested-namespace basename collisions). Round 3 replaced it with an
environment-derived inventory (`Proofs/Inventory.lean` +
`inventory-allowlist.txt`, fully-qualified names, no filtering) and
`selftest_audit.sh`, which runs the published evasion table plus a
namespace collision, an axiom smuggle, a stale-entry case, and two
unmanifested-module cases against the exact production gate.
Residual honesty: the inventory sees what the compiled environment
contains; it cannot see source that is never compiled (which the
dead-file checks cover) or defeat a hostile Lean toolchain.
13. **Review-kit fidelity target was not self-contained in round 2**
(GPT H2: missing load-time imports made `run_fidelity.py` unrunnable
from the kit). Round 3 ships the complete stdlib-only import closure
of `pacta.transparency`, content-addressed against pacta commit
`3d81d53`, plus the clean-extraction transcript with exit code.
14. **Deployed `verify_consistency` accepts strictly more than the
mechanized `ConsRec` on lied-size inputs** (round-3 Claude addendum
F1*, reproduced by the operator against deployed pacta). Witness:
for the honest proof P between sizes 2→3, `verify_consistency(1, 3,
R2, R3, P)` returns True — a semantically false claim ("R2 is the
root of a size-1 prefix") — while `ConsRec` rejects; 3,405 such
divergences exist for n < 60, ALL one-sided (the mechanized model
never accepts anything the deployed verifier rejects; inclusion
shows zero divergences under identical abuse). Mechanism: when the
claimed old size is a power of two, the deployed RFC 9162 iterative
algorithm seeds the walk with the old root and uses the sizes only
as bit-navigation state, so several size claims navigate one proof
identically. Consequences: (a) fidelity between the two consistency
verifiers is agreement over the pinned case families, NOT
extensional equality the harness's lied-size family pins the
boundary (73,573 cases, 3,867 expected divergences, direction
asserted one-sided); (b) Theorem 3 / `acceptCons_sound` cover the
MECHANIZED accept set, and their soundness transfers to the deployed
verifier only under the side condition that the consumer's
`(n₀, r₀)` is an authentic pinned pair and `n₁` is the authentic
size of the tree behind `r₁` which pacta's pin-store flow supplies
by construction (`n₀` comes from the consumer's own pin, never from
the peer; `(n₁, r₁)` arrive together in one signed head). No
exploitability against that flow is claimed or ruled out here;
assessing it requires the signature/STH layer (gap 4). No pacta code
change is made (deployed behavior matches upstream RFC 9162
implementations; the consumer flow enforces the side condition).