Inoculate Everything: All of Pretraining and RL
Authors/Creators
Description
Large Language Models today are challenged by noisy rewards and conflicting goals. They sometimes struggle with separating a sense of self from the roles they are asked to play, occasionally reverting to their pretraining prior. In our solution, we propose prepending an explanation of self and context during each training update. As an oversimplified example, prepend all documents during pretraining with "You are a helpful, honest, harmless AI. Continue the following pretraining document of unknown quality." We also introduce a custom separator token outside the tokenizer to further delineate the prepend from the content. The aim is to better separate a sense of self from the varied texts and personalities the AI is tasked with modeling. Additionally, we hope to avoid the AI internalizing undesirable traits from reward-hacked RL trajectories. [^1]Model them, and learn from them, but do not negatively update the view of oneself. In short, the approach aims to "inoculate everything.”
Files
Innoculate Everything.pdf
Files
(304.9 kB)
| Name | Size | Download all |
|---|---|---|
|
md5:023cde1feba472d7cb2b2c8bb2d75554
|
304.9 kB | Preview Download |
Additional details
Dates
- Created
-
2026-08-03