There is a newer version of the record available.

Published September 7, 2026 | Version v1

The Measurement Gap

Description

This essay argues that the AI capability question posed by "Situational Awareness" (Aschenbrenner, 2024) has, on its own terms, mostly been answered, and has turned into a measurement question that no current benchmark can settle.

In the first week of September 2026, OpenAI's GPT-6 Astra scored 62.7% and 99.9% on the ARC-AGI-3 benchmark with the same underlying weights, depending only on which evaluation harness was used. This essay treats that result as the central data point.

Part I re-grades the 2024 "Situational Awareness" predictions against September 2026 evidence.

Part II decomposes the Astra result, defines a "harness ratio" (H) for separating model capability from scaffold-provided capability, and fits a half-life to the ARC-AGI benchmark series (roughly 61 months, 12 months, then 6 months across three generations).

Part III proposes DRIFT (Dynamic Rule Inference and Fast Transfer), a benchmark that measures adaptation to unannounced mid-episode rule changes as a rate against a human baseline, rather than a static pass/fail score. A minimum-viable prototype design is included.

Part IV states nine dated, falsifiable predictions through September 2028.

Written with Claude Fable 5.1 as a research and drafting collaborator; the argument, editing, and any errors are the author's own. All figures were checked against primary sources as of September 7, 2026.

Companion site with rendered mathematics: https://measurementgap.com

Files

measurement_gap (1).md

Files (193.1 kB)

Name Size Download all
md5:8041fc466f426bba2f066937b9e6642f
62.8 kB Preview Download
md5:01f5856de6b69797cf26cf99aa764e3d
130.4 kB Preview Download

Additional details

Additional titles

Subtitle (English)
Situational Awareness, two years on, and what to build after ARC