Published October 8, 2026 | Version 1.0

Three Hazards Under One Horizon: Constant-Hazard, Degrading-Step and Log-Logistic Accounts of Agent Failure Describe Different Levels of the Same Data, the 50 Percent Trend Barely Depends on Which Is Right, and Every High-Reliability Forecast Does

Authors/Creators

  • 1. SONYTECH

Description

(c) 2026 Pranay Mahendrakar. Licensed under CC BY 4.0.

Forecasts of autonomous AI capability increasingly rest on one number: the length of task, measured in human working time, that an agent completes half the time. That number has doubled roughly every seven months, and three accounts of why agents fail on longer tasks now sit behind its extrapolation. One fits a logistic curve in the logarithm of task length; one proposes a constant failure rate per minute of human task time, giving each agent a half-life; one reports that per-step accuracy falls as a run proceeds, because models condition on their own earlier errors. They are usually read as rival models of one process. This paper argues that they describe three different levels of the same data - the hazard inside a run, the hazard per unit of a task's length, and the success curve across a heterogeneous task suite - and that the aggregation linking those levels is the classic problem of unobserved heterogeneity, under which a falling pooled hazard is compatible with constant or rising hazards inside every task. Published evidence at each level points in different directions: rising, flat and cliff-shaped within-run failure all appear, depending on model, harness and task. Two consequences follow, one reassuring and one not. The 50 percent trend is nearly indifferent to which account is right, since every account is calibrated to it. Every forecast at higher reliability is not. Holding the 50 percent horizon fixed, curves that all match the published 80 percent ratios place the 99 percent horizon anywhere from about one hundredth to about one eight-hundredth of it, and the within-task hazard a deployment would face, which those ratios do not identify, could put it as high as one eighth. A reading procedure for horizon claims and six measurements that would settle the open part are offered. The analysis is built from published measurements, each credited to the source that reported it.

Files

three-hazards-one-horizon.pdf

Files (426.3 kB)

Name Size Download all
md5:d4ebadbf9515add032c4e1281f2b458e
426.3 kB Preview Download