Compute and Scaling Laws, Explained
This page is a structured working draft — real analysis, not yet expanded with the full expert sourcing given to the flagship pages. Safe to build on; treat specifics as provisional until sourced.
The empirical pattern
Since the early 2020s, AI capability has tracked a strikingly predictable curve: models trained with more compute, more data, and more parameters perform measurably better on a wide range of tasks, in a relationship researchers call a scaling law. This predictability, more than any single breakthrough, is what convinced many skeptics that continued investment would keep paying off.
Why this makes compute a governance lever
If capability scales predictably with compute, then compute becomes a rare, physically trackable proxy for capability — you can meter and monitor the large data centers and chip supplies capable of training frontier models far more easily than you can monitor algorithms or code. This is the entire logic behind export controls on advanced AI chips and reporting thresholds tied to training-run size (see Compute Governance).
Where scaling might bend
Three things could break or bend the curve: physical limits on chip manufacturing and energy supply; a shortage of high-quality training data as the internet’s easily available text is exhausted; and diminishing returns if the next capability jump requires qualitatively different architectures rather than more of the same recipe. Researchers are actively divided on how close any of these limits actually are.
The practical takeaway
Scaling laws are the reason lab leaders’ timelines compressed rather than lengthened after 2023 — for several years, simply spending more on compute kept working better than almost anyone expected.