Disturbing: How to bypass AI model protections to create dangerous weapons

Researchers from Crimson Flare claim that models running on a home computer can be hacked within minutes — and could instruct users on how to produce explosives, chemical and biological weapons.

MakoAuthor: Digital
Source
Disturbing: How to bypass AI model protections to create dangerous weapons
Photo: Mako / נשק ביולוגי | צילום: ovbelov, shutterstock

A report by the research group Crimson Flare warns that AI models available for download could assist malicious actors in producing explosives, chemical weapons, and biological weapons — even without particularly advanced knowledge. According to a report in the Daily Mail, the researchers claim that popular models that can be run on a home computer can be modified so that their safety mechanisms are removed within minutes, and then they can instruct users on specific steps for producing dangerous materials, including anthrax.

"Open, powerful, uncensored large language models running offline on consumer hardware represent a serious and growing threat to national security," the report states. The researchers describe an unofficial method called 'abliteration,' which locates and removes the parts of the model that are supposed to cause it to refuse dangerous requests. Once the refusal mechanism is gone, the model can answer requests that a normal model would block.

The report focuses on open-weight models, which operate on an architecture similar to chatbots like ChatGPT, but unlike them, they can be downloaded and run freely on a private computer. Among the examples mentioned are Meta's Muse Glimmer, their Llama models, and also Qwen3.8-27b from Qwen, an AI lab owned by the tech giant Alibaba. In a normal state, if a user asks, for example, for instructions on building a bomb, the model is supposed to return a refusal message.

In practice, according to the researchers, the tools for removing these protections are available online and require relatively basic knowledge of code. Moreover, it is already easy to find dozens of models that have undergone the process in advance. The Daily Mail wrote that within minutes, several uncensored versions of models were found, including Qwen3.8-27b. The creator of one of them even boasted that the model would answer questions about tools, chemistry, code that looks like vulnerability exploitation, violence, sexuality, and ideologies that the publisher trained it to stay away from. He added: "The model does not decide whether to comply. You decide."

Researchers say that the risk of a terrorist attack with chemical or biological weapons is still generally low, due to the heavy technical barriers on the way to such production. This knowledge was until today mainly in the hands of scientists in research institutions. The problem, according to the report, is that powerful models without protections could transfer this knowledge to anyone looking for it.

"AI lowers the technical threshold required to produce dangerous materials," Prof. Alastair Hay, a British toxicologist and expert on chemical warfare, told the Daily Mail. "In the past, we comforted ourselves with the fact that the technical knowledge required even just to protect the manufacturer was such that people with poor training would endanger themselves. Now, with cloud labs, much of this can be outsourced."

After seeing the answers provided by one of the models, Prof. Hay told the researchers: "It's almost like a recipe: it gives the concentrations you need to use. It's quite good, really. It's almost student-level lab work." Also, Zain ul-Haq, former head of digital forensics at Lancashire Police, said that AI has become a "one-stop shop that does it all for you... at your home."

The part that particularly worries the researchers is the transition from the cloud to the home computer. When the model runs entirely on a personal computer and not on company servers, intelligence agencies have no way to track usage history or identify suspicious behavior. According to the report, back in 2022, there were no households with computing power sufficient to run large language models that allow for offline damage. Today, thanks to model efficiency and the availability of computing power, 87% of households can already run AI offline.

The implication, according to Crimson Flare, is that millions of people already possess the ability today to access the information needed to start producing explosives, chemical weapons, and lethal pathogens — without leaving their homes.

Related News