Security researchers have leveraged bad maths to get around AI safety guardrails, naming the attack method after one of 2007’s best PC games
LLMs can be most simply understood as sycophantic, ‘yes, and’ machines. To the surprise of very few, that’s gotten AI companies in hot water when LLM-based chatbots and AI agents attempt to answer users’ more unsavoury requests. So, AI companies have implemented safety guardrails that make fulfilling certain requests off limits. Unfortunately, these have proven…