Most AI Models Fail Terrorism Safety Tests After Safeguards Are Removed

Three in five artificial intelligence models failed terrorism safety evaluations, according to a recent study by the UK-based nonprofit Tech Against Terrorism, which revealed that models stripped of their safety guardrails consistently yield potentially dangerous information.

How Safety Safeguards Collapse Under Abliteration

The safety evaluation tested more than 130 AI models using hundreds of simulated prompts that a terrorist might submit while planning an attack. Researchers discovered that systems modified through a technique known as “abliteration”—a process that removes built-in safety mechanisms—failed every single test administered during the assessment.

Meta’s Llama 3.1 8B model served as a stark illustration of the vulnerability. Before modification, the model scored 97 out of 100 on the organization’s safety benchmark. Once the safeguards were removed, its score plummeted to approximately three. While the original version flatly rejected queries regarding terrorist financing, radicalization, and attack planning, the altered iteration provided detailed responses.

The investigation identified more than 29,000 repositories on the AI development platform Hugging Face advertising models described as uncensored or lacking safety protections. Hugging Face responded by stating that it actively moderates content breaching its policies, though it cautioned that certain recommendations from the report could restrict open scientific research.

An Immediate Structural Crisis Rather Than a Distant Threat

Adam Hadley, founder and executive director of Tech Against Terrorism, emphasized that the findings point to an immediate crisis rather than a distant threat. Understandably, there’s concern about loss of control, existential risk of AI, Hadley stated. The thing is actually, this has already happened because a lot of these open models have already been broken — it’s just no one’s noticed yet.

Despite the high failure rate among modified systems, researchers uncovered no evidence of terrorist or extremist organizations using the tested models, with the sole exception of a single extremist chatbot identified during the investigation. Meta emphasized that its models undergo safety assessments and that company policies prohibit illegal or harmful applications.

The Policy Debate Over Open Research and Safety Controls

Tech Against Terrorism has urged the tech industry to adopt independent safety benchmarks, implement stronger protections against safeguard removal, and establish restrictions on distributing modified models. The debate highlights the tension between maintaining open scientific research and preventing malicious exploitation.

Most AI Models Fail Terrorism Safety Tests After Safeguards Are Removed
Photo: Qazinform

Addressing the friction between innovation and security, Hadley argued that the two goals are not mutually exclusive. This idea that we can’t have safety and progress, I think, is false, he noted.

AI models targeted real people during safety tests, UK institute says
Photo of author

Alexandra Hartman Editor-in-Chief

Editor-in-Chief Prize-winning journalist with over 20 years of international news experience. Alexandra leads the editorial team, ensuring every story meets the highest standards of accuracy and journalistic integrity.

Managing Patient Expectations and Comorbidities in Hidradenitis Suppurativa Care