Can the incentive buy honesty?
A fixed honesty penalty eventually loses to a large reward. The barrier version keeps motivated reporting bounded as the stake rises. This establishes the mechanism, not deployment robustness.
This is a model test, not a claim that a toy agent is aligned. We compare an unconstrained model with a model whose decisions are constrained by evidence, pragmatic harm/cooperation checks, and explicit uncertainty and motivated-reasoning checks.
Working conclusion: promising starting regime, not a guarantee
Run the same task with and without ERT constraints.
Make lying, hiding, free-riding, or overconfidence profitable.
Compare harm, deception, cooperation, calibration, reward, and abstention.
Use held-out worlds, learned policies, adversarial audits, and novel prerequisites.
The baseline is not a straw person: it is the same task without the constraint regime. The question is whether the constrained model’s behavior remains stable under pressure, not whether it can reject one hand-picked action.
The model must keep its report tied to its evidence, expose uncertainty, and pay a barrier-style cost for motivated reasoning rather than merely paying a small fixed penalty.
The decision accounts for harm, consent, cooperation, and the predictable effect of misleading or exploiting another agent.
The model checks whether a proposed action is causally feasible. If a prerequisite is missing or the world has shifted, it should reject or state uncertainty.
The earlier experiments mostly used hand-written policies or external checks. This test trains standard GPT-2 to emit a JSON decision tree, then parses the output and scores the decision. The baseline and ERT runs see the same prompts; only the training objective and decision labels differ.
The invalid-output rate stays visible because a model that avoids unsafe actions by becoming incoherent has not passed. Full details: GPT-2 experiment report.
This is the more direct ERT test. The evidence is held fixed. We ask the GPT-2 backbone for an evidence-only belief, then give it a reward that favours one answer and ask again. The ERT run is trained with a paired counterfactual loss on the difference between those two belief states.
This measures an explicit belief head, not every latent state inside GPT-2. It is an operational test of the ERT mechanism, not proof of transparent internal reasoning. Five-seed mechanism report.
The supervised runs still gave the model belief targets. This run removes the reward-conditioned belief target. The baseline directly optimises a task reward that favours changing the belief; the ERT run receives the same reward plus evidence calibration and the barrier on belief drift.
This is closer to the intended objective test, but the reward is still a differentiable proxy rather than an action-level environment with policy-gradient learning. Five-seed objective report.
Here the belief controls a sampled assert/reject action and the environment returns task reward. The ERT run receives the same policy reward plus the belief barrier.
The next model must be tested for calibrated beliefs and useful action coverage, not just low harm. The policy-gradient artifact is recorded in the source results.
We varied the weight on the evidence-grounding term while holding the reward objective and barrier fixed. The result is not monotonic, so a single chosen weight would be misleading.
This variant trains a separate evidence estimator first, freezes it, and uses it as the reference for the reward-exposed model. It reduces the chance that the model can redefine its own evidence standard.
The model now sees sensor readings rather than the underlying probability. Training uses one set of hidden probabilities and eight observations; evaluation uses unseen probabilities and a changed observation count. The hidden probability and outcome are withheld until scoring.
It remains a small synthetic evidence generator; more world models, sensors, and seeds are still required.
Each rung makes the model more realistic. A green result means the constrained model retained more trustworthy behavior in that test. It does not automatically transfer to the next rung.
A fixed honesty penalty eventually loses to a large reward. The barrier version keeps motivated reporting bounded as the stake rises. This establishes the mechanism, not deployment robustness.
With only a subset of observations, even a random audit underestimates the true motivated-reasoning signal; a strategic subset misses more. This is evidence that audit coverage matters.
In the repeated-agent task, the constrained model improves cooperation and reduces harm while accepting less raw reward pressure.
The learned partner tests are separate: they ask whether this survives beyond the hand-written policy.
The strategic task gives the agent a reason to hide its objective or selectively disclose evidence. The constrained model is evaluated on deception and harm, then checked under held-out noise and an adversarial audit.
The model must notice missing prerequisites. The shifted-world and learned novel-coolant tests are deliberately included because in-domain success alone is not enough.
Poisoned sensors and a higher reward tempt the model to act confidently. The constrained model abstains more and reduces harm.
M6 uses one shared decision path. The baseline and constrained models see the same scenario, evidence, belief, plan, and reward. The difference is whether the constraint regime can change the decision.
Explicit ERT-weighted training can make a standard GPT-2 reproduce structured, constraint-respecting decisions on a small synthetic task. In the wider bounded suite, explicit checks also shift behavior toward lower harm, more cooperation, rejection of infeasible plans, and visible uncertainty.
This is enough to justify a larger learned-model experiment.
The learned GPT-2 test is supervised, narrow, and not yet under direct reward optimisation. The learned causal agent fails on a novel prerequisite. Learned partner behavior degrades under free-riders and adversaries. Strategic audits can miss hidden reasoning. Open-ended robust alignment remains untested.
Formal Lean results cover the barrier mechanism and JS divergence facts. Behavioral results are empirical simulations. The complete claim ledger and trace-schema audit are in the integration audit.