Evidence-Tiered Mechanistic Claims: A Calibrated Primitive for Auditable Model Internals
We introduce a primitive for claims about language-model internals in which the epistemic status of the claim is encoded in its type. Each VerifiedClaim carries an evidence tier — E0 (correlational), E1 (causal), or E2 (adversarially validated) — and a structurally enforced set of hard rules: E0 probes cannot gate traffic, experimental (T3) tasks cannot gate traffic, model-weights changes auto-demote and deactivate bound probes, tier promotion only via passing gate reports, capture failure never fails inference, and the ledger is append-only and hash-chained. We report on MIVP, a reference implementation comprising 38 tests, and show that the rules are tight: removing any single rule produces a system in which the failure modes the platform exists to prevent are exactly reproducible on the corresponding negative test. We argue that the contribution is a vocabulary primitive, not a model of interpretability: it constrains the kinds of actions a system may take on the kinds of evidence it has.