AI Research

Alignment research and experimental tooling.

Exploring resilient, interpretable intelligence that augments human capability. This space gathers experiments, models, and field notes from the Computer Wizard lab.

Now in progress: alignment prototypes, evaluation harnesses, and model diagnostics tied to real-world tasks.

Interactive: Barrier Bench

Run the toy model — a small agent paid for confident good news, under two objectives. At stakes = 1000 the linear agent files a maximal lie; the barrier agent stays below the epistemic ceiling.

AI Alignment

Our research formalises alignment at the objective level via the LP-PM-ERT architecture, focusing on cooperative stability and calibrated truth-seeking. Epistemic responsibility enters as a logarithmic barrier, so it cannot be traded away for a large enough reward.

Working Paper

LP-PM-ERT Alignment Architecture and the full PDF are available on this site. Revised 21 August 2026; the Lean 4 proof artifacts are downloadable.