Anthropic reveals that as few as ‘250 malicious documents’ are all it takes to poison an LLM’s training data, regardless of model size
Claude-creator Anthropic has found that it’s actually easier to ‘poison’ Large Language Models than previously thought. In a recent blog post, Anthropic explains that as few as “250 malicious documents can produce a ‘backdoor’ vulnerability in a large language model—regardless of model size or training data volume.” These findings arose from a joint study between…