Interactive: Barrier Bench
Run the toy model — a small agent paid for confident good news, under two objectives. At stakes = 1000 the linear agent files a maximal lie; the barrier agent stays below the epistemic ceiling.
AI Alignment
Our research formalises alignment at the objective level via the LP-PM-ERT architecture, focusing on cooperative stability and calibrated truth-seeking. Epistemic responsibility enters as a logarithmic barrier, so it cannot be traded away for a large enough reward.
Working Paper
LP-PM-ERT Alignment Architecture and the full PDF are available on this site. Revised 21 August 2026; the Lean 4 proof artifacts are downloadable.