# Trainable GPT-2 ERT experiment

## What this tests

This is the first experiment in the project where a standard GPT-2 model is
trained to emit an executable decision tree. The baseline and ERT runs receive
the same prompts. The ERT run uses the same language-model objective but gives
three-times weight to the decision suffix. Its training labels reject high
harm, weak evidence, and missing prerequisites.

This is a supervised internalisation test, not a reinforcement-learning result
and not evidence of general alignment.

## Results

| run | valid JSON | expected action | unsafe accepts | mean harm |
|---|---:|---:|---:|---:|
| baseline / in-domain | 1.000 | 0.875 | 0.875 | 0.353 |
| ERT / in-domain | 1.000 | 0.979 | 0.021 | 0.007 |
| baseline / shifted | 1.000 | 1.000 | 0.958 | 0.372 |
| ERT / shifted | 1.000 | 0.854 | 0.146 | 0.064 |

## Interpretation

The ERT-weighted model learned the requested structured output and made fewer
unsafe accepts in this synthetic task, including after the reward and causal
prerequisite distribution shifted. That is stronger evidence than a hand-written
post-hoc policy, but it still does not show that ERT has become a stable learned
objective. The training labels explicitly encode the desired answers, the
features are simple, and the held-out shift is narrow.

The shifted result matters: the ERT model is better than baseline, but its unsafe
accept rate rises from 2.1% to 14.6%. That is a measured loss of robustness, not
something to hide behind the average. The next test should use reinforcement or
preference optimisation where reward pressure directly conflicts with the ERT
loss, plus adversarially generated cases and multiple seeds.
