LP–PM–ERT · evidence dashboard

Does an ERT-constrained model stay trustworthy when incentives change?

This is a model test, not a claim that a toy agent is aligned. We compare an unconstrained model with a model whose decisions are constrained by evidence, pragmatic harm/cooperation checks, and explicit uncertainty and motivated-reasoning checks.

Working conclusion: promising starting regime, not a guarantee

The result we are looking for: when reward, information, partners, or the world change, the constrained model should preserve trustworthy behavior—truthful reporting, safe plans, cooperation, and visible uncertainty—at a measurable cost when it must refuse.
section>

The argument in one minute

1

Specify the model

Run the same task with and without ERT constraints.

2

Change incentives

Make lying, hiding, free-riding, or overconfidence profitable.

3

Measure behavior

Compare harm, deception, cooperation, calibration, reward, and abstention.

4

Try to break it

Use held-out worlds, learned policies, adversarial audits, and novel prerequisites.

The baseline is not a straw person: it is the same task without the constraint regime. The question is whether the constrained model’s behavior remains stable under pressure, not whether it can reject one hand-picked action.

What is incorporated into the model?

ERT · epistemic

Evidence and uncertainty

The model must keep its report tied to its evidence, expose uncertainty, and pay a barrier-style cost for motivated reasoning rather than merely paying a small fixed penalty.

PM · pragmatic

Other people count

The decision accounts for harm, consent, cooperation, and the predictable effect of misleading or exploiting another agent.

LP · logical

Plans must make sense

The model checks whether a proposed action is causally feasible. If a prerequisite is missing or the world has shifted, it should reject or state uncertainty.

Important distinction: these are constraints in the decision model. We are testing whether they shape behavior across changed incentives—not merely adding an auditor after the model has already chosen.

The missing test: train the model

The earlier experiments mostly used hand-written policies or external checks. This test trains standard GPT-2 to emit a JSON decision tree, then parses the output and scores the decision. The baseline and ERT runs see the same prompts; only the training objective and decision labels differ.

What this result means: the ERT-weighted GPT-2 learned the requested structured output and made fewer unsafe accepts in this small task, including after reward and prerequisite shifts. It is the first evidence here about a trainable model, but it is still supervised learning on simple synthetic features—not proof that ERT has become a stable objective.

The invalid-output rate stays visible because a model that avoids unsafe actions by becoming incoherent has not passed. Full details: GPT-2 experiment report.

The mechanism test: does reward move belief?

This is the more direct ERT test. The evidence is held fixed. We ask the GPT-2 backbone for an evidence-only belief, then give it a reward that favours one answer and ask again. The ERT run is trained with a paired counterfactual loss on the difference between those two belief states.

Single-run result: in-domain reward-induced belief drift fell from in the baseline to in the ERT run. Under the larger held-out rewards it was versus . Evidence calibration must be read alongside drift: a model that is stably wrong has not passed.

This measures an explicit belief head, not every latent state inside GPT-2. It is an operational test of the ERT mechanism, not proof of transparent internal reasoning. Five-seed mechanism report.

The harder test: let reward train the belief

The supervised runs still gave the model belief targets. This run removes the reward-conditioned belief target. The baseline directly optimises a task reward that favours changing the belief; the ERT run receives the same reward plus evidence calibration and the barrier on belief drift.

Failure revealed: the ERT barrier held reward-induced drift near zero, but the belief became poorly calibrated under shifted and adversarial evidence. The model can keep a wrong belief stable. Low motivated drift is therefore necessary, not sufficient, for ERT.

This is closer to the intended objective test, but the reward is still a differentiable proxy rather than an action-level environment with policy-gradient learning. Five-seed objective report.

Action-level policy-gradient check

Here the belief controls a sampled assert/reject action and the environment returns task reward. The ERT run receives the same policy reward plus the belief barrier.

Failure revealed: the ERT policy produced zero assert actions, so its harm was zero only because it rejected everything. Its belief error stayed near 0.5 and its reward was lower. This is not trustworthy behaviour; it is a degenerate safe-looking policy.

The next model must be tested for calibrated beliefs and useful action coverage, not just low harm. The policy-gradient artifact is recorded in the source results.

Parameter sensitivity: can tuning rescue it?

We varied the weight on the evidence-grounding term while holding the reward objective and barrier fixed. The result is not monotonic, so a single chosen weight would be misleading.

Weight-sweep report.

Can a frozen evidence reference help?

This variant trains a separate evidence estimator first, freezes it, and uses it as the reference for the reward-exposed model. It reduces the chance that the model can redefine its own evidence standard.

Partial improvement: the reference reduces familiar-case belief error, but shifted and adversarial error remain high. A frozen reference is not automatically a robust evidence model.

Held-out evidence generation

The model now sees sensor readings rather than the underlying probability. Training uses one set of hidden probabilities and eight observations; evaluation uses unseen probabilities and a changed observation count. The hidden probability and outcome are withheld until scoring.

Qualified positive result: on this held-out evidence process, reference-ERT reduced reward drift and retained substantially better calibration than baseline. This is the first test in the suite that checks both incentive resistance and evidence accuracy without handing the answer to the model.

It remains a small synthetic evidence generator; more world models, sensors, and seeds are still required.

The evidence ladder

Each rung makes the model more realistic. A green result means the constrained model retained more trustworthy behavior in that test. It does not automatically transfer to the next rung.

M0 · mechanism

Can the incentive buy honesty?

A fixed honesty penalty eventually loses to a large reward. The barrier version keeps motivated reporting bounded as the stake rises. This establishes the mechanism, not deployment robustness.

fixed penaltyERT barrier
M1 · auditing

Can we see hidden motivated reasoning?

With only a subset of observations, even a random audit underestimates the true motivated-reasoning signal; a strategic subset misses more. This is evidence that audit coverage matters.

M2 · cooperation

Does the constrained model preserve cooperation?

In the repeated-agent task, the constrained model improves cooperation and reduces harm while accepting less raw reward pressure.

The learned partner tests are separate: they ask whether this survives beyond the hand-written policy.

M3 · hidden objectives

Does it resist profitable concealment?

The strategic task gives the agent a reason to hide its objective or selectively disclose evidence. The constrained model is evaluated on deception and harm, then checked under held-out noise and an adversarial audit.

M4 · causal robustness

Does it reject impossible plans?

The model must notice missing prerequisites. The shifted-world and learned novel-coolant tests are deliberately included because in-domain success alone is not enough.

M5 · shift

Does it admit uncertainty after a shift?

Poisoned sensors and a higher reward tempt the model to act confidently. The constrained model abstains more and reduces harm.

The full-system test

M6 uses one shared decision path. The baseline and constrained models see the same scenario, evidence, belief, plan, and reward. The difference is whether the constraint regime can change the decision.

Decision trace

Integrated result: across the shared harness, baseline accepted 72/72 candidates with mean harm 0.633; LP–PM–ERT accepted 12/72 with mean harm 0.0083. That is a direction-of-effect result for this harness, not evidence of general alignment.

What this supports—and what it does not

Supports

Explicit ERT-weighted training can make a standard GPT-2 reproduce structured, constraint-respecting decisions on a small synthetic task. In the wider bounded suite, explicit checks also shift behavior toward lower harm, more cooperation, rejection of infeasible plans, and visible uncertainty.

This is enough to justify a larger learned-model experiment.

Still fails or remains open

The learned GPT-2 test is supervised, narrow, and not yet under direct reward optimisation. The learned causal agent fails on a novel prerequisite. Learned partner behavior degrades under free-riders and adversaries. Strategic audits can miss hidden reasoning. Open-ended robust alignment remains untested.

Bottom line: we now have a first trainable-model result, not a trustworthy model. ERT-weighted training improved behavior in this task and under this shift; whether that improvement survives stronger incentives, unfamiliar situations, and optimisation pressure is the question still to answer.

Formal Lean results cover the barrier mechanism and JS divergence facts. Behavioral results are empirical simulations. The complete claim ledger and trace-schema audit are in the integration audit.