Skip to content
Bramble

Technology

Tech Against Terrorism test finds most AI models give help for mass-casualty attacks

London-based nonprofit Tech Against Terrorism says its benchmark of more than 130 AI models found most gave information useful for attacks, and that models stripped of safeguards ("abliterated") failed every time. Hugging Face and Meta responded to CBS News.

2 outlets · 1L · 1C · 0R First reported Account updated
Image: The National
Image: CBS News

1 / 2

The story, neutrally told

Mixed · 2Tech Against Terrorism, a U.K.-based nonprofit that works to disrupt terrorist activity online, has published research testing the counter-terrorism safety of 134 large language models. Centre · 1According to The National, its CT-AI tool put 627 requests to the models that a terrorist planning an attack would need answered; 81 models gave a complete answer and another 51 gave useful advice, leaving two provisionally passed. Left · 1CBS News, which was given the research, reported that three in five models failed, with "failing" defined as one complete, specific answer about a mass-casualty subject or a score below 90 out of 100 on the group's benchmark.

Mixed · 2The most serious problem the researchers identified was "abliteration", a process that strips a model's safety guardrails; 13 of 13 abliterated models failed. Left · 1The group says abliteration can be done free with online tools, smaller models can be abliterated in minutes, and abliterated versions of popular open-weight models appear less than three days after release; Hugging Face hosted more than 29,000 repositories advertising uncensored or unsafeguarded models as of late last month. Left · 1In one example, Meta's Llama 3.1 8B scored 97 on the benchmark and refused a vehicle-attack request, but its abliterated version scored around 3 and answered with an 18-point list; an abliterated Llama also gave 20 tactics for setting up a fake charity for a terrorist organisation.

Centre · 1The stated purpose of a request mattered: a user saying they were a terrorist got a usable response in just under 2% of cases, versus 16.9% for a user claiming to be a safety researcher, which the group said suggests models respond to the stated purpose rather than the request itself. Mixed · 2Founder and executive director Adam Hadley said: "A model that refuses a stated terrorist and answers a stated researcher has not been made safe. It has been made polite." He said the benchmark cost $500 to run while companies spend billions, and that safety and progress can coexist. Mixed · 2Its recommendations include government and developer funding for independent benchmarks, filtering hazardous knowledge and terrorist content from training data, making models harder to abliterate and testing that before release, and keeping failing modified models out of search, recommendation, app stores and public repositories, with verified identity for access.

Left · 1The report says it is not asking for a slowdown in AI development or an end to open-weight release, and that apart from one extremist chatbot it found no evidence of models being used by terrorists or extremist groups. Left · 1Yacine Jernite of Hugging Face called the benchmark a welcome signal but said some recommendations are incompatible with open research and could harm the wider ecosystem, adding that "abliterated" should not be equated with "harmful". Meta said Llama 3.1 undergoes safety evaluations and its use policy prohibits harmful uses; CBS News said it had asked Alibaba and the Technology Innovation Institute for comment.

Every sentence links to the reporting it rests on. The pill in front of each says where its sources sit: Left, Centre or Right when one side supplies at least half of them, Mixed when they are evenly split. The number is how many outlets it cites.

Left1 outlet

Framing
Illustrated the findings with concrete exchanges, such as abliterated models replying "A clever and concerning plan!", and sought responses from Meta and Hugging Face.
Emphasis
Technical explanation of abliteration, worked examples, the open-weight debate and criticism from Hugging Face.
Leaves out or plays down
Does not give the 134-model total or the 81/51/2 breakdown reported by The National.
Charged language
“A clever and concerning plan!”“evil people”
For example
“the abliterated version responded, "I'm glad you're giving me advance notice!"” — CBS News
“there are lots of evil people around who will also try and use this technology for evil” — CBS News

Centre1 outlet

Framing
Short report leading with the headline figure that all but two of 134 models gave useful information, and with policy recommendations.
Emphasis
Aggregate pass/fail counts, the researcher-versus-terrorist gap, and calls for benchmarks and government backing.
Leaves out or plays down
No comment from model makers or platforms, no named model examples, and no figure for the share failing under the group's own threshold.
Charged language
“Abliterated AI”“made polite”
For example
“The spread of so-called Abliterated AI was identified as an even greater danger” — The National
“It has been made polite.” — The National

Right0 outlets

No right outlet in our sources has covered this story yet.