Scalable Oversight: How Humans Supervise Systems Smarter Than Them
This page is a structured working draft — real analysis, not yet expanded with the full expert sourcing given to the flagship pages. Safe to build on; treat specifics as provisional until sourced.
The problem this solves
Standard oversight assumes the supervisor understands the task at least as well as the system being supervised. That assumption breaks down for a system operating beyond human-expert level in a domain — how do you verify work you can’t independently check?
Leading approaches
- Debate — two AI systems argue opposing sides of a question in front of a human or weaker AI judge, on the theory that it’s easier to spot a flaw in an opposing argument than to independently verify a complex claim from scratch.
- Recursive reward modeling — using AI assistance to help humans evaluate AI outputs on tasks too complex to judge unaided, building up supervisory capability in layers.
- Weak-to-strong generalization — studying whether a weaker model’s supervision signal can still reliably steer a stronger model’s behavior in the right direction, even though the weaker model can’t fully verify the stronger one’s reasoning.
Where the research stands
These techniques are active, published research directions with encouraging but early results in controlled settings — none has been demonstrated at the scale or stakes a genuinely superhuman system would require. This is one of the areas where the gap between current research maturity and the capability trajectory (see Timeline Forecasts) is most concerning to safety researchers.
Why this can’t be solved by regulation alone
Scalable oversight is fundamentally a technical problem — no law can make an unsolved verification problem solved. Governance (see International Treaties and Summits) can buy time and set incentives for labs to invest in this research; it can’t substitute for the research succeeding.