Published October 31, 2025 | Version v1

Understanding and Mitigating Poisoning Attacks in Large Language Models

  • 1. ROR icon IU International University of Applied Sciences

Description

This paper explores the growing threat of data poisoning and backdoor attacks in large language models (LLMs), revealing that even a small, fixed number of poisoned samples—around 250 documents—can compromise models up to 13B parameters. It synthesizes recent research, explains experimental methodologies from Anthropic and others, and provides actionable defense strategies for AI engineers and enterprises. The work emphasizes the urgent need for trusted data pipelines, anomaly detection, and post-training audits to ensure AI model integrity at scale.

Files

01.pdf

Files (72.2 kB)

Name Size Download all
md5:8ded0b2e01a54977e4a5029a6fdfa4ce
72.2 kB Preview Download