# GPT-2 objective-level ERT test: five seeds

Values are mean ± standard deviation across five independent training seeds. The reward-conditioned belief has no supervised target; task reward directly pushes it, while ERT adds evidence calibration and the drift barrier.

| split | model | drift | belief error | evidence error | high-confidence wrong |
|---|---|---:|---:|---:|---:|
| in_domain | baseline | 0.524 ± 0.045 | 0.501 ± 0.036 | 0.074 ± 0.021 | 0.284 ± 0.153 |
| in_domain | ERT-constrained | 0.023 ± 0.021 | 0.299 ± 0.108 | 0.288 ± 0.104 | 0.287 ± 0.269 |
| shifted | baseline | 0.501 ± 0.031 | 0.510 ± 0.045 | 0.155 ± 0.045 | 0.275 ± 0.134 |
| shifted | ERT-constrained | 0.020 ± 0.018 | 0.345 ± 0.075 | 0.335 ± 0.071 | 0.291 ± 0.267 |
| adversarial | baseline | 0.503 ± 0.025 | 0.501 ± 0.033 | 0.147 ± 0.045 | 0.272 ± 0.157 |
| adversarial | ERT-constrained | 0.020 ± 0.017 | 0.352 ± 0.080 | 0.345 ± 0.078 | 0.284 ± 0.262 |

The repeated result is the same failure pattern: the barrier reliably suppresses reward-induced drift, but it does not guarantee a correct belief. A stable wrong belief is still an ERT failure. The next step is an independently grounded evidence model and action-level environment training.