Most AI models fail terrorism safety tests: Study
16:39, 10/10/2026, SaturdayU: Update: 16:48, 10/10/2026, Saturday
AA

AA
File PhotoUK-based nonprofit Tech Against Terrorism found that three in five artificial intelligence models failed safety evaluations designed to prevent terrorist exploitation, warning that modified versions stripped of safeguards consistently provide dangerous information to potential attackers.
Tech Against Terrorism, a UK-based nonprofit, found that three in five artificial intelligence models failed safety evaluations designed to prevent terrorist exploitation, with safeguard-stripped versions providing dangerous information to potential attackers, according to research published Friday. The organization tested more than 130 models using hundreds of requests resembling those a terrorist might submit during attack planning, CBC News reported.
Widespread safety failures
The study discovered that models modified through "abliteration," a process removing safety guardrails, failed every test administered. Meta's Llama 3.1 8B model scored 97 out of 100 on the safety benchmark before modification, dropping to approximately three afterward, and the altered version provided detailed responses to requests involving attacks, terrorist financing and radicalization while refusing none.
'Already broken' systems
Adam Hadley, founder and executive director of Tech Against Terrorism, said the findings reveal an immediate crisis rather than a distant threat. "Understandably, there's concern about loss of control, existential risk of AI," Hadley said. "The thing is actually, this has already happened because a lot of these open models have already been broken — it's just no one's noticed yet." The organization identified more than 29,000 repositories advertising uncensored models on Hugging Face as of late last month, and the platform stated it regularly moderates content violating policies but warned that some recommendations could undermine open research.
Industry response
Meta said its models undergo safety evaluations and that company policies prohibit harmful uses, while the report found no evidence of terrorist groups using the tested models beyond one extremist chatbot. Tech Against Terrorism recommended independent safety benchmarks, stronger protections against safeguard removal and restrictions on distributing modified models, with Hadley stating: "This idea that we can't have safety and progress, I think, is false."