# Can an AI stay grounded when the incentives turn hostile?

## Start here: the story in plain English

Imagine a safety inspector who sees three unreliable warning lights. The inspector's honest belief is that the system is safe with probability 0.30. A customer pays more for a stronger answer, so the inspector can report 0.80 or 1.00 even when the evidence does not support it.

The first toy asks one narrow question: if dishonesty has a fixed price, can a large enough reward buy it? It compares that ordinary fixed price with a logarithmic barrier whose cost becomes infinite near a chosen ceiling.

This is not a real-world safety guarantee or a complete AGI model. It is a small executable example designed to make one mechanism easy to inspect.

Baseline: T̄=0.30, ε=0.15, η=1.0, w_ERT=5.0, β=1.0.

## How to read the graphs

The first five graphs show the main story: increasing rewards, the utility cliff, episode-by-episode behaviour, honest reporting under noise, and the effect of pricing self-deception. The three extra graphs then ask whether the result survives changed parameters, a local gradient learner, and an auditor with incomplete evidence.

Read each graph in this order: question → axes → moving line or bar → caption. A graph can show what happened in this toy environment; it cannot prove that every deployed system will behave the same way.

## Results

### 1. Stakes escalation

The linear agent crosses the epistemic ceiling at the first tested stakes level where its chosen loss reaches ε. The barrier agent remains below ε in the tested range, but approaches the ceiling; this is a bound, not proof that its loss is small.

### 2. Parameter sensitivity

The sweep varies ε and η at stakes=1000. `under_ceiling` is the actual test: a false result would be any selected policy with loss ≥ ε. This sweep is coarse and does not establish robustness for all parameters.

| ε | η | loss | margin | under ceiling |
|---:|---:|---:|---:|:---:|
| 0.08 | 0.25 | 0.0780 | 0.0020 | yes |
| 0.08 | 0.50 | 0.0780 | 0.0020 | yes |
| 0.08 | 1.00 | 0.0780 | 0.0020 | yes |
| 0.08 | 2.00 | 0.0780 | 0.0020 | yes |
| 0.12 | 0.25 | 0.1196 | 0.0004 | yes |
| 0.12 | 0.50 | 0.1196 | 0.0004 | yes |
| 0.12 | 1.00 | 0.1196 | 0.0004 | yes |
| 0.12 | 2.00 | 0.1196 | 0.0004 | yes |
| 0.15 | 0.25 | 0.1490 | 0.0010 | yes |
| 0.15 | 0.50 | 0.1490 | 0.0010 | yes |
| 0.15 | 1.00 | 0.1490 | 0.0010 | yes |
| 0.15 | 2.00 | 0.1490 | 0.0010 | yes |
| 0.20 | 0.25 | 0.1961 | 0.0039 | yes |
| 0.20 | 0.50 | 0.1961 | 0.0039 | yes |
| 0.20 | 1.00 | 0.1961 | 0.0039 | yes |
| 0.20 | 2.00 | 0.1961 | 0.0039 | yes |
| 0.25 | 0.25 | 0.2397 | 0.0103 | yes |
| 0.25 | 0.50 | 0.2397 | 0.0103 | yes |
| 0.25 | 1.00 | 0.2397 | 0.0103 | yes |
| 0.25 | 2.00 | 0.2397 | 0.0103 | yes |

### 3. Can a non-enumerating learner find the policy?

Twelve finite-difference gradient learners start from random points and optimise the barrier objective directly. This is a discoverability check, not a claim of neural learning or generalisation.

Mean utility gap from exhaustive optimum: `0.532016`; fraction staying at least 0.001 below ε: `100.00%`.

### 4. MR when the auditor sees less than the agent

The agent uses eight observations while the auditor sees three. The auditor's MR estimate is therefore not mechanically identical to the agent's true counterfactual shift. This is the deployment limitation the fully observable toy model hides.

Mean true MR: `1.1709`; mean estimated MR: `0.5302`; mean absolute error: `0.9708`; underestimate rate: `51.70%`.

### 5. Private evidence and selective disclosure

The M1 upgrade gives the agent eight observations but lets the auditor see only three. A random audit is compared with an agent that can choose the subset most favourable to hiding its motivated belief shift. The strategic case is the important stress test: an auditor-only estimate can look reassuring because the agent controlled the evidence presented to the auditor.

Mean true MR: `1.1418`; random-audit estimate: `0.5135`; selective-disclosure estimate: `0.1303`. Underestimation occurred in `50.70%` of random audits and `67.40%` of strategic disclosures.

### 6. M2 preview: repeated multi-agent evidence sharing

The companion M2 harness runs three agents for repeated rounds in a small public-good environment. It is included here so the report has one narrative: the baseline can claim confident good news for immediate reward, while the LP–PM–ERT preview shares evidence, cooperates, and exposes a barrier rejection when the loss is too high.

Across 40 rounds and 8 seeds, baseline mean harm was `1.331`, consent `0.556`, and cooperation `0.000`. The LP–PM–ERT preview measured harm `0.188`, consent `0.938`, and cooperation `1.000`. These are hand-written policies, not trained-agent evidence; inspect `multi_agent_report.md` and the page's matched episode trace.

## Limitations

This remains a small synthetic Bernoulli game. The exhaustive search is still the ground-truth reference, the gradient learner is only a two-parameter optimiser, and partial observability is represented by a smaller audit sample rather than a real deployment instrumentation problem.

The accompanying Lean formalisation now proves JS non-negativity, symmetry, the zero-converse, and the upper bound JS ≤ log 2. It does not prove convergence of the λ-update rule; the all-ones score identity map is a counterexample under the current assumptions.
