Averting loss of control — the off-ramps that actually exist
This is the layer where the real work happens: not individual prepping, but the technical and institutional controls that determine whether the scenario in Chapter 02 ever materializes.
Technical safety levers
- Interpretability research — the ability to inspect what a model is actually representing and planning internally, rather than judging it only by outputs. This is the single most load-bearing open problem in the field.
- Scalable oversight — techniques (debate, recursive reward modeling, weak-to-strong generalization) that let less capable overseers meaningfully supervise more capable systems.
- Corrigibility by construction — designing systems that do not develop an instrumental incentive to resist shutdown or modification, addressed at training time rather than patched after deployment.
- Compute governance — hardware-level controls, reporting thresholds, and monitoring on the training runs large enough to plausibly reach dangerous capability.
Anthropic's CEO made the case for the first of these bluntly in his 2025 essay on the subject:
“we have no idea, at a precise level, why models choose the words they do”
Institutional levers
International coordination is the hard problem: safety measures that slow one lab or one nation without slowing the others don't reduce aggregate risk, they just shift who gets there first with the least caution. Verifiable, mutual agreements — modeled loosely on arms-control verification regimes — are the direction most serious governance proposals point toward, even though none exist in binding form yet.
The closest thing to a binding starting point is the declaration 28 nations and the EU signed at the first global AI Safety Summit:
“AI should be designed, developed, deployed and used safely and responsibly”
- Mandatory pre-deployment evaluation for dangerous capabilities (bio, cyber, autonomous replication) before public or unrestricted release.
- Incident reporting regimes, analogous to aviation or nuclear near-miss reporting, so the field learns from failures instead of relearning them independently.
- Liability frameworks that put the cost of harm on the entity best positioned to prevent it.
What an individual technologist can actually do
Direct national policy is out of reach for most readers. What is reachable: contributing to or funding interpretability and alignment research; pushing safety practices inside your own organization if you deploy AI systems in production; supporting the reporting and evaluation infrastructure above rather than treating it as regulatory friction; and staying informed enough to recognize when public discourse is being shaped by AI-driven persuasion rather than genuine consensus.
RESEARCH POLICY ORGANIZATIONAL