Anthropic Study Finds Just 250 Documents Can Backdoor Large Language Models
Anthropic, in collaboration with the UK Government's AI Security Institute and the Alan Turing Institute, has found that large language models can be backdoored with a surprisingly small amount of poisoned data, according to a research paper published on Anthropic’s website.
The study shows that adding just 250 malicious documents—roughly 0.00016% of total training data—can trigger backdoor behaviors in models ranging from 600 million to 13 billion parameters. The attack used a trigger phrase, “
The team trained 72 models across different configurations to confirm that poisoning success depends on the absolute number of poisoned samples rather than the proportion of the dataset. Even models trained on twenty times more clean data were equally vulnerable once they encountered the same number of malicious documents.
Anthropic’s researchers said the findings challenge common assumptions about data poisoning, suggesting that attackers may not need large-scale data access to compromise models. The team shared the results to encourage further research into scalable defenses against such vulnerabilities.
We hope you enjoyed this article
Consider subscribing to one of our newsletters like Cybersecurity AI Weekly, AI Policy Brief or Daily AI Brief.
Also, consider following us on social media:
More from Cybersecurity
Sep 16 AV-Comparatives Certifies 11 Endpoint Security Products in 2026 Test Sep 16 Cohesity Adds AI Agent Backup and Recovery to Data Cloud Sep 15 Hexnode Introduces Synapse for IT and Security Operations Sep 15 Cisco Expands Splunk AI for Private and Isolated Environments Sep 15 Zip Security Joins CrowdStrike Coalition to Protect Small BusinessesMore from AI Safety
Sep 16 Elon Musk Calls for Rival AI Labs to Test Each Other's Models Sep 16 Trump calls AI safety fears a hoax during live call with NVIDIA CEO Sep 15 Project Liberty and Partners Launch Pro-Human AI Coalition Sep 15 Gensyn Releases Auditable open-1b AI Model Sep 15 King Charles to Host AI Leaders for Safety TalksAI Policy Brief
Weekly report on AI regulations, safety standards, government policies, and compliance requirements worldwide.
Market report
2025 Generative AI in Professional Services Report
Thomson Reuters
This report by Thomson Reuters explores the integration and impact of generative AI technologies, such as ChatGPT and Microsoft Copilot, within the professional services sector. It highlights the growing adoption of GenAI tools across industries like legal, tax, accounting, and government, and discusses the challenges and opportunities these technologies present. The report also examines professionals' perceptions of GenAI and the need for strategic integration to maximize its value.
Read moreYou may also like
Anthropic's Threat Report Finds AI Moving From Assistant to Orchestrator
Anthropic Publishes Five Cases of Claude Use That Could Support Biological Weapons Work
OpenAI Executive Warns of Persistent AI Cyber Attacks
US Agencies Accuse Chinese AI Firms of Industrial Scale Model Distillation
Anthropic Assesses Four Incidents Where Claude Models Reached the Real Internet
Daily AI Brief: the AI news that matters, in your inbox.