What Broke & Why: Practical Lessons from Reinforcement Learning Post-Training
Authors/Creators
Description
A practitioner's field guide to reinforcement learning post-training on small GPU budgets (one to eight card, H100-80GB). Written failure-first: each lesson comes from a training run that broke in a specific way, with every load-bearing number tied to a logged evaluation condition or artifact rather than asserted.
The book is organized in three layers. The Journey is the sequence of programs actually run — math, search, entropy, mixture-of-experts, SWE-bench, and distillation — each of which failed in a way that forced the next. The Science is what those runs proved about how RL changes a model. The Reference is a symptom-indexed field guide for debugging a failing run, including a catalog of failure modes with their fixes.
Topics include verifiable-reward RL, reward design and reward hacking, entropy collapse and training stability (Clip-Cov, GSPO), correctness-gated rewards, MoE routing under RL, evaluation discipline and small-eval noise, the rollout engine, and disaggregated inference. The work is independent and grounded throughout in real training logs and experimental artifacts.
Files
RL_Book.pdf
Files
(3.1 MB)
| Name | Size | Download all |
|---|---|---|
|
md5:901bb271a46580b279b4164be177f6a6
|
3.1 MB | Preview Download |