Published September 24, 2026 | Version 1.0

Curriculum Is Three Claims, Not One: Ordering, Selection and Decomposition in RLVR, and the Random-Order Control Almost Nobody Runs

Authors/Creators

  • 1. SONYTECH

Description

(c) 2026 Pranay Mahendrakar. Licensed under CC BY 4.0.

Curriculum learning entered large-scale reasoning training through reinforcement learning with verifiable rewards (RLVR): difficulty-ordered or difficulty-filtered problem schedules are now routine in systems built on GRPO and DAPO, and dozens of 2024-2026 papers report a gain from some version of "curriculum." Five years earlier, a controlled study spanning thousands of orderings on standard image and language benchmarks found that curricula beat random ordering only under a restricted training budget or noisy labels, and that even those gains were attributable to a dynamically expanding training set rather than to the ordering itself (Wu et al., 2021). This paper reads the RLVR curriculum literature against that finding and against the literature's own most careful recent attempt to re-run it. Of the RLVR papers surveyed here that report a curriculum or difficulty-selection gain, only one holds total training steps and total unique problems fixed while randomizing the schedule -- the control Wu et al. specify -- and that paper, tested across multiple model families on synthetic reasoning benchmarks, reports no robust advantage of difficulty-based sequencing over random sampling in either accuracy or response length (Mordig et al., 2026). The remaining papers, spanning math, writing, multi-domain and preference-data settings, compare against an unfiltered or uniformly-sampled baseline that changes what the model trains on, not merely the order it trains on it in, and a controlled theoretical treatment of the RLVR setting attributes the provable benefits of adaptive problem choice specifically to changing the training distribution, not to sequencing a fixed one (Rajaraman et al., 2026a). What the field calls "curriculum" in RLVR names at least three distinct mechanisms -- static ordering, adaptive selection, and structural decomposition -- with three different evidentiary records, and the one sharing its name and its instrumentation with a mechanism that failed a matched-control test twice, five years apart, is the one still invoked as the field's working premise.

Files

curriculum-three-claims-rlvr.pdf

Files (6.1 MB)

Name Size Download all
md5:2c0d3c91a2e69f1f789304eb5c1a05e9
6.1 MB Preview Download