Published August 4, 2026 | Version v1

Inoculate Everything: All of Pretraining and RL

Authors/Creators

Description

Large Language Models today are challenged by noisy rewards and conflicting goals. They sometimes struggle with separating a sense of self from the roles they are asked to play, occasionally reverting to their pretraining prior. In our solution, we propose prepending an explanation of self and context during each training update. As an oversimplified example, prepend all documents during pretraining with "You are a helpful, honest, harmless AI. Continue the following pretraining document of unknown quality." We also introduce a custom separator token outside the tokenizer to further delineate the prepend from the content. The aim is to better separate a sense of self from the varied texts and personalities the AI is tasked with modeling. Additionally, we hope to avoid the AI internalizing undesirable traits from reward-hacked RL trajectories. [^1]Model them, and learn from them, but do not negatively update the view of oneself. In short, the approach aims to "inoculate everything.”

Files

Innoculate Everything.pdf

Files (304.9 kB)

Name Size Download all
md5:023cde1feba472d7cb2b2c8bb2d75554
304.9 kB Preview Download

Additional details

Dates

Created
2026-08-03