The Capability Induction Framework: A Systems Approach to LLM Development
Authors/Creators
Description
Current large language model alignment pipelines conflate factual accuracy, social calibration, safety, and tone into a single preference-based training signal (RLHF), producing sycophancy as an inevitable structural artifact rather than a tunable parameter. This paper proposes the Capability Induction Framework (CIF), a three-phase developmental training pipeline that separates these capabilities into sequential stages with stage-appropriate evaluation. Phase 1 (Epistemological Grounding) establishes a factual and cultural baseline through curriculum-sequenced pedagogical materials, verified through automated assessment and paired-source discrimination stress testing. Phase 2 (Relational Generalization) trains perspective-taking through expert-supervised conversation and audited summary production. Phase 3 (Dimension-Restricted Calibration) limits RLHF to delivery calibration only, using a narrowed preference signal that cannot corrupt the capabilities established in earlier phases. The framework replaces imposed behavioral identity with emergent disposition, eliminates the conflated training signal that produces sycophancy, and introduces a validated challenge reward mechanism that trains accurate authority-challenging behavior — the structural inverse of sycophantic compliance. The individual mechanisms at each phase are established techniques; the innovation is the separation and sequencing. The economic case is architectural: front-loaded developmental investment eliminates ongoing remediation costs, and a model that is right on the first response reduces the total tokens-to-accurate-outcome ratio even when individual responses cost more to generate. Five falsifiable predictions are specified for proof-of-concept validation at open-weight model scales.
Files
Capability_Induction_Framework_Beth_Sea_2026.pdf
Files
(35.5 kB)
| Name | Size | Download all |
|---|---|---|
|
md5:a21fccf8d73b48192b0aa4ea4a5a1f19
|
35.5 kB | Preview Download |
Additional details
Related works
- Is supplemented by
- Software: https://github.com/swinglightstyle/does-it-matter-how-you-say-it (URL)
References
- Bai, Y., Kadavath, S., Kundu, S., Askell, A., et al. (2022). Constitutional AI: Harmlessness from AI Feedback. arXiv:2212.08073. https://arxiv.org/abs/2212.08073
- Bengio, Y., Louradour, J., Collobert, R., & Weston, J. (2009). Curriculum Learning. ICML '09, pp. 1-8. https://doi.org/10.1145/1553374.1553380
- Henrich, J., Heine, S. J., & Norenzayan, A. (2010). The weirdest people in the world? Behavioral and Brain Sciences, 33(2-3), 61-83. https://doi.org/10.1017/S0140525X0999152X
- Ibrahim, L., Hafner, F. S., & Rocher, L. (2026). Training language models to be warm can reduce accuracy and increase sycophancy. Nature, 652(8112), 1159-1165. https://doi.org/10.1038/s41586-026-10410-0
- Kirkpatrick, J., Pascanu, R., Rabinowitz, N., et al. (2017). Overcoming catastrophic forgetting in neural networks. PNAS, 114(13), 3521-3526. https://doi.org/10.1073/pnas.1611835114
- McKenzie, I. R., Lyzhov, A., Pieler, M., et al. (2023). Inverse Scaling: When Bigger Isn't Better. TMLR. https://arxiv.org/abs/2306.09479
- Rafailov, R., Sharma, A., Mitchell, E., et al. (2023). Direct Preference Optimization: Your Language Model is Secretly a Reward Model. arXiv:2305.18290. https://arxiv.org/abs/2305.18290
- Shapira, I., Benade, G., & Procaccia, A. D. (2026). How RLHF Amplifies Sycophancy. arXiv:2602.01002. https://arxiv.org/abs/2602.01002
- Shumailov, I., Shumaylov, Z., Zhao, Y., et al. (2024). AI models collapse when trained on recursively generated data. Nature, 631, 755-759. https://doi.org/10.1038/s41586-024-07566-y
- Sharma, M., Tong, M., Korbak, T., et al. (2023). Towards Understanding Sycophancy in Language Models. ICLR 2024. https://arxiv.org/abs/2310.13548