stadion: A Benchmark for Operational Decisions with Classical Baselines and Exact Optima
Description
Most benchmarks used to evaluate autonomous agents share a structural limitation: nobody knows what the right answer was. A score of 61% on a browsing benchmark or an agent-versus-agent leaderboard says that one agent beat other agents on a particular task instance; it does not say whether the winning behaviour was close to what any reasonable procedure would recommend. We present stadion, an open-source benchmark for operational decisions in which the ground truth exists and is computable. Six tasks — inventory ordering, dynamic pricing of a perishable, admission to a finite-buffer queue, energy arbitrage against a daily price cycle, two-echelon supply-chain replenishment, and joint price-and-restock — are drawn from operations-research problems whose small state spaces admit exact solutions by backward induction. Every agent decision is scored against two references it cannot argue with: the classical operations-research method for the problem and the exact optimum. The result is a normalised score with a paired bootstrap confidence interval, on which indistinguishable from the classical method is a first-class outcome rather than a rounding error. We describe the scoring protocol, the six tasks, the harness self-verification procedure, and the headroom (0.4%–26.6%) between the classical rule and the optimum on each task — a distribution that argues against the common benchmark practice of selecting tasks where the classical answer is bad. The library is Python and treats reinforcement-learning policies and language-model agents through the same interface.
Files
Drobyshev_2026_stadion.pdf
Files
(226.3 kB)
| Name | Size | Download all |
|---|---|---|
|
md5:43875d302dab1df11f7a2d198f26e601
|
226.3 kB | Preview Download |
Additional details
Software
- Repository URL
- https://github.com/DrobyshevDev/stadion
- Programming language
- Python
- Development Status
- Active