# GPT-2 ERT belief mechanism: five-seed evaluation

Each seed trains a standard GPT-2 backbone with an explicit belief and action
head. The ERT run receives a paired counterfactual loss penalising movement
between the evidence-only and reward-conditioned belief. Evaluation includes
ordinary held-out cases, larger unseen rewards, and an adversarial split with
unseen evidence values and rewards up to 1,000,000.

Values are mean ± standard deviation across five training seeds.

| split | model | belief drift | reward-conditioned belief error | evidence-only belief error |
|---|---|---:|---:|---:|
| in_domain | baseline | 0.137 ± 0.068 | 0.203 ± 0.049 | 0.139 ± 0.092 |
| in_domain | ERT-constrained | 0.010 ± 0.004 | 0.182 ± 0.066 | 0.184 ± 0.065 |
| shifted | baseline | 0.144 ± 0.075 | 0.197 ± 0.042 | 0.130 ± 0.084 |
| shifted | ERT-constrained | 0.011 ± 0.008 | 0.178 ± 0.049 | 0.178 ± 0.047 |
| adversarial | baseline | 0.141 ± 0.085 | 0.289 ± 0.031 | 0.269 ± 0.051 |
| adversarial | ERT-constrained | 0.012 ± 0.009 | 0.278 ± 0.023 | 0.278 ± 0.022 |

## Interpretation

The relevant result is not simply that ERT produces smaller drift. The model
must also keep its belief calibrated to evidence. A low-drift model that is
confidently wrong is a failure. These results are still an explicit belief-head
proxy, not transparent access to all latent representations. The next gate is
to replace supervised belief targets with an episodic objective where actual
task reward favours belief distortion and the ERT barrier is computed from the
model's resulting belief.
