The Measurement Gap
Authors/Creators
Description
This essay argues that the AI capability question posed by "Situational Awareness" (Aschenbrenner, 2024) has, on its own terms, mostly been answered, and has turned into a measurement question that no current benchmark can settle.
In the first week of September 2026, OpenAI's GPT-6 Astra scored 62.7% and 99.9% on the ARC-AGI-3 benchmark with the same underlying weights, depending only on which evaluation harness was used. This essay treats that result as the central data point.
Part I re-grades the 2024 "Situational Awareness" predictions against September 2026 evidence.
Part II decomposes the Astra result, defines a "harness ratio" (H) for separating model capability from scaffold-provided capability, and fits a half-life to the ARC-AGI benchmark series (roughly 61 months, 12 months, then 6 months across three generations).
Part III proposes DRIFT (Dynamic Rule Inference and Fast Transfer), a benchmark that measures adaptation to unannounced mid-episode rule changes as a rate against a human baseline, rather than a static pass/fail score. A minimum-viable prototype design is included.
Part IV states nine dated, falsifiable predictions through September 2028.
Written with Claude Fable 5.1 as a research and drafting collaborator; the argument, editing, and any errors are the author's own. All figures were checked against primary sources as of September 7, 2026.
Companion site with rendered mathematics: https://measurementgap.com
Files
measurement_gap (1).md
Files
(193.1 kB)
| Name | Size | Download all |
|---|---|---|
|
md5:8041fc466f426bba2f066937b9e6642f
|
62.8 kB | Preview Download |
|
md5:01f5856de6b69797cf26cf99aa764e3d
|
130.4 kB | Preview Download |
Additional details
Additional titles
- Subtitle (English)
- Situational Awareness, two years on, and what to build after ARC