Why the Third Axis Is Weakness
Authors/Creators
Description
Scored one sample at a time, training cannot tell a model that only ever produces one hotdog from one that has learned the whole distribution of hotdogs. Explorative Modeling instead draws $K$ candidates and trains on whichever is closest to the training example, calling exploration a third pretraining axis. But what exactly does it scale? I connect exploration to weakness, which previous work proved optimal for generalisation. Weakness is measured on the trained model as it samples at inference, counting the further commitments it could still make. I derive an exact identity for best-of-$K$ risk in any measurable output space. The identity allows the candidate law to depend on the target itself. The bridge from risk to weakness therefore requires training to draw candidates exactly as the model samples at inference. Suppose every valid output is equally likely and candidates score only by exact match. For every $K>1$, among models that permit only valid outputs and spread probability evenly across them, the chance of a match rises strictly with the number of permitted outputs, and so with weakness. Raising $K$ steepens that relation. The unique optimum spreads mass evenly across every valid output, the weakest policy that never emits an invalid one. When valid outputs differ in frequency, each optimum above $K=1$ is unique and approaches that same limit as $K$ grows. Now let every unseen context independently draw a uniformly random nonempty set of required outputs. The probability of meeting every requirement is exactly proportional to weakness over the unseen contexts. What this means is that one should select and audit generators by measured weakness rather than mode counts. Training data define a problem, parameters are the form of a solution, and weakness measures its future compatibility.
Files
EAW-5.pdf
Files
(344.3 kB)
| Name | Size | Download all |
|---|---|---|
|
md5:18fc0c081dfe5abd2d708f8251536f50
|
344.3 kB | Preview Download |