Published August 22, 2026 | Version 1.0

Three Guarantees Under One Word: What Importance-Based Selective Forgetting in Language Models Can and Cannot Promise

Authors/Creators

  • 1. SONYTECH

Description

(c) 2026 Pranay Mahendrakar. Licensed under CC BY 4.0.

Unlearning methods are asked to deliver a guarantee, and the word is used for three different ones: that a model's outputs no longer reveal the target under some stated class of queries, that the target is absent from the weights, and that no adversary within a stated budget can restore it. This paper separates the three, reads the published record as establishing that they come apart in practice, and asks which one each of the field's motivating use cases actually requires. Read through that separation, a literature that appears to disagree about whether unlearning works turns out to be reporting different guarantees under the same heading: benchmark forget-quality numbers are the first guarantee measured under the weakest access model, and the recovery results that appear to refute them are the third guarantee measured at budgets the benchmarks never applied. The paper then audits the specific premise that licenses importance-based and weight-attribution methods - that the target is concentrated in identifiable parameters, that those parameters can be found, and that changing them removes rather than reroutes - and finds published counter-evidence against each link, with the third link the weakest and the least addressed. On the demand side, only one of the three motivating use cases is well served by the guarantee the field is optimising, and for the right-to-erasure case an unprovability result suggests that the auditable deliverable is a documented procedure rather than a property of the weights. This paper reports no experiments and no measurements of its own. It states what the published record establishes, states flatly what it does not, proposes a three-part disclosure that would let a reader tell the guarantees apart, and names five studies that would settle the open part - including one matched-compute ablation that has been run for image classifiers and, as far as the author can determine, never for language models.

The literature search, drafting and citation verification for this paper were carried out with AI assistance under the author's direction. Every citation was machine-verified against the arXiv API and Crossref before inclusion, and every quantitative claim was read back against the cited source's own text. The author is responsible for the final text and for all claims made in it.

Files

unlearning-three-guarantees.pdf

Files (495.1 kB)

Name Size Download all
md5:a96c2f57016bd27aeff3ee641021b28a
495.1 kB Preview Download