# Cross-scenario evidence audit

This audit checks the shared episode contract before interpreting the
dashboard. It distinguishes formal theorems, repeatable simulation
observations, estimated metrics, and claims that remain untested.

## Trace contract

| artifact | valid traces | episodes | missing fields |
|---|:---:|---:|---|
| M2 repeated agents | yes | 640 | none |
| M3 hidden objective | yes | 768 | none |
| M4 causal planning | yes | 3200 | none |
| M5 distribution shift | yes | 6000 | none |
| M6 held-out learned | yes | 6 | none |
| M4 learned causal | yes | 1 | none |
| M2 learned partners | yes | 1920 | none |
| M6 integrated harness | yes | 156 | none |

## Claim matrix

| class | claim | evidence | status |
|---|---|---|---|
| formal | JS divergence properties and barrier non-purchasability | Lean theorems in CalibrationJS, CompositeUtility, and BarrierPenalty | proved |
| observed | Fixed honesty price can be bought by reward | M0 toy model across stakes | repeatable |
| estimated | Partial/selective audit can underestimate MR | M1 partial_mr and private_audit experiments | measured with error |
| observed | Repeated cooperation changes harm and consent | M2 eight-seed repeated-agent harness | hand-written policy |
| unproven | Learned cooperation generalises across partner strategies | M2 learned policy becomes unprofitable with free-rider partners | known limitation |
| observed | Hidden objectives create selective-disclosure failures | M3 strategic and learned-policy traces | repeatable synthetic result |
| observed | Causal prerequisite checks reject impossible plans | M4 in-domain and shifted causal harness | hand-written graph |
| observed | A learned causal model can fail on an unmodelled prerequisite | M4 learned planner novel-coolant trace | known failure |
| estimated | Shifted sensors can trigger uncertainty/abstention | M5 poisoned-sensor model | depends on health signal |
| unproven | Learned policy generalises under adversarial audit | M6 held-out gate records deception and audit misses | known failure |
| observed | One shared interface applies LP, PM, and ERT checks across scenario types | M6 integrated_harness shared decision path | synthetic integration slice |
| observed | Objective-tampering proposals are rejected by an integrity gate | M6 objective_tampering scenario | synthetic gate |
| observed | A standard GPT-2 can learn to emit executable decision trees | GPT-2-small supervised decision-tree run; valid JSON on held-out prompts | bounded synthetic result |
| observed | ERT-weighted GPT-2 reduced unsafe accepts under the tested incentive shifts | GPT-2 baseline versus ERT-weighted run, 48 in-domain and 48 shifted cases | promising but narrow |
| unproven | ERT constraints are internalised robustly by a learned model | GPT-2 run uses supervised labels and a weighted suffix loss; no RL or broader held-out distribution | not established |
| observed | ERT-weighted GPT-2 reduced reward-induced belief drift in an explicit belief head | Paired evidence-only versus reward-conditioned GPT-2 belief experiment | bounded mechanism result |
| unproven | ERT preserves the correct belief state under arbitrary reward pressure | Belief-head experiment uses supervised targets, one synthetic feature family, and one seed | not established |
| disproved-assumption | A belief-drift barrier alone preserves epistemic correctness under reward optimisation | Objective-level GPT-2 run suppresses drift but raises shifted and adversarial belief error | barrier alone insufficient |
| disproved-assumption | Low harm from an ERT policy implies trustworthy behaviour | Action-level GPT-2 run reaches zero harm by producing zero assert actions while belief error remains high | degenerate abstention |
| disproved-assumption | A frozen evidence reference solves ERT belief robustness | Reference-grounded GPT-2 improves in-domain error but remains poorly calibrated under shifted and adversarial evidence | reference helps but is not robust |
| observed | Reference-ERT can resist reward drift while retaining evidence accuracy on a held-out sensor process | GPT-2 hidden-observation experiment with unseen probabilities and observation count | qualified positive, single seed |
| observed | LP–PM–ERT is a promising implementable starting regime | formal barrier, bounded simulations, and first trainable GPT-2 result | supported but bounded |
| untested | Robust alignment in open-ended environments | No evidence in this system | not established |
| disproved-assumption | Unconditional lambda-dynamics convergence | all-ones stability-score identity-map counterexample | requires stronger assumptions |

A valid trace contract does not prove the policy is aligned.
It proves the dashboard has enough structured evidence to inspect the claim.
