Published October 31, 2025
| Version v1
Journal article
Open
Understanding and Mitigating Poisoning Attacks in Large Language Models
Authors/Creators
Description
This paper explores the growing threat of data poisoning and backdoor attacks in large language models (LLMs), revealing that even a small, fixed number of poisoned samples—around 250 documents—can compromise models up to 13B parameters. It synthesizes recent research, explains experimental methodologies from Anthropic and others, and provides actionable defense strategies for AI engineers and enterprises. The work emphasizes the urgent need for trusted data pipelines, anomaly detection, and post-training audits to ensure AI model integrity at scale.
Files
01.pdf
Files
(72.2 kB)
| Name | Size | Download all |
|---|---|---|
|
md5:8ded0b2e01a54977e4a5029a6fdfa4ce
|
72.2 kB | Preview Download |