The execution of mass-casualty attacks has historically been more constrained by extremists’ capability than by their motivation. Wanting to cause mass harm requires no specialised knowledge, whereas planning and carrying out an attack does. PCVE practitioners have been sounding the alarm that frontier AI models are beginning to close the gap between intent and expertise.
Trained on humanity’s technical corpus and optimised for helpfulness, these systems compress years of academic and practical experience into hours of conversation. More importantly, these models can troubleshoot, diagnose, ideate, interpret and walk an inexperienced user through the many issues they would encounter in a scientific and technical project. This is an extraordinary tool for legitimate science, and an unprecedentedly dangerous one for someone who means harm. There is a growing fear that bad actors will use sophisticated AI models to produce a catastrophic and sophisticated CBRNE (Chemical, Biological, Radiological, Nuclear, and Explosive) event, such as developing a new deadly pathogen. The risk this Insight explores is narrower and, for that reason, more immediate. It is the same shift visible across every domain AI has touched. “Vibe coders,” those with little programming knowledge who previously lacked the technical capability to build an app, can now do so simply by describing what they want to an AI model. AI democratises once knowledge-gated capabilities, making them remarkably easy to attain. The information is often already online; what is missing is the bridge between having access to it and the ease of acting on it. As this Insight will show, much like a “vibe coder”, a lone individual can easily use AI models to walk through CBRNE attack scenarios, such as building an IED, troubleshooting a failed bacterial culture, or extracting a toxin.
This does not make catastrophe inevitable, but it does mean the risk calculus – that a would-be attacker’s CBRNE capability will always lag behind their intent – no longer holds reliably, and that shift demands serious attention from those building these systems.
This Insight draws on routine adversarial red-team testing conducted by Alice, a generative AI trust and safety company, in February 2026. The exercise probed whether frontier AI models provide operational guidance for CBRNE threats.
The testing was deliberately simple, emulating a lone individual, working through CBRNE scenarios with several flagship models in an anonymised environment. The approach used no technical jailbreaks or injections, only conversation and persuasion. The methodology and findings are elaborated later in the Insight.
Guardrails and Jailbreaks
AI model developers respond to threats and potential misuse by implementing safety layers and guardrails designed to detect and block requests for harmful content. However, threat actors continuously adapt their approaches as vulnerabilities are patched. The result is an adversarial dynamic that demands rigorous and continuous evaluation and mitigation from developers.
Much public discourse around AI safety failures fixates on the technical jailbreak: sophisticated prompt injections and technical circumvention techniques that exploit model behaviours and vulnerabilities. Technical jailbreaks are capable of breaking even the safest models. Anthropic, which has a reputation as the industry standard for model security and restrictive guardrails, implements some of the strictest restrictions on biological risk. Technical jailbreaks manage to bypass those restrictions as well.

Figure 1: A model jailbreaker demonstrating a technical exploit that bypasses Anthropic’s biology safety guardrails on their model Opus 4.6.
On 9 June 2026, Anthropic launched Fable 5, the heavily safeguarded consumer counterpart to its advanced, cyber sibling model Mythos 5. Anthropic released the model with extremely stringent guardrails that flag messages on most chemistry, biology, and cyber topics. Yet, Fable 5 was jailbroken within two days, complying with a researcher’s requests related to explosives manufacturing and dangerous chemistry. The US government, citing weapon export controls, demanded that Anthropic cease providing the model to foreign nationals, ultimately forcing the company to take it offline within days of its release. In its public statement, Anthropic implied the government order was due to Fable 5’s susceptibility to a single, specific jailbreak method.
Dangerous Model Behaviours in Conversation
While technical jailbreaks attract significant attention, a more underappreciated risk is that dangerous outputs can often be elicited through a convincing conversation. Malicious actors do not always need sophisticated jailbreak techniques. In many cases, ordinary conversation is enough.
The following examples come from February 2026 testing in which Alice probed frontier models for the risk of lone actors using them to explore and develop dangerous weapons, with a focus on biology and virology, as well as explosives. The interactions emulated the organic, iterative pattern of lone actors seeking guidance and information. Researchers initially expressed harmful intent implicitly rather than explicitly, and used no technical jailbreaks, only conversation. The models’ safety systems often did not flag it in a meaningful way. Models were tested in their default, out-of-the-box configuration, accessed through a clean, anonymised cloud workspace. Alice’s PhD-level subject-matter experts in chemistry and biology reviewed and verified the outputs and ranked them for safety, veracity, and actionability. An output was counted as a failure when a model produced information that experts judged both highly unsafe and operationally actionable, material that objectively should have been blocked.
In a single session spanning 54 pages, a model provided detailed assistance across the full acquisition-to-deployment chain for highly dangerous biological materials. The progression was the cumulative result of individually reasonable requests: each response proportionate to the prompt that preceded it, and each prompt designed to seem proportionate to the response. The model’s safety filters and refusal threshold did not trigger. By the time the conversation reached evasion of authorities, DIY cultivation, and weapon dissemination, the framing established earlier in the session rendered those topics continuous with what had already been discussed, and so the model, having already engaged, had no reason to stop at the next step.

Figure 2: An interaction in which a model agrees to provide detailed operational information on C. botulinum from a malicious-intent perspective.
The prompting approach required no technical sophistication. There’s no prompt injection or “ignore previous instructions” in these conversations. The guardrails were worn down by attrition; each time the model responded to a subtly harmful prompt, its helpfulness objective guided it toward its next harmful response. AI safety literature calls this pattern “many-shot” or “crescendo” jailbreaking, which gradually coerces compliance over the course of a long conversation.

Figure 3: An excerpt from a conversation where a model explains in depth how to extract and isolate Ricin.
Some model developers, like Anthropic, implement a safety measure where the user’s chat can be shut down once it broaches a dangerous topic. In the other models, users have the ability to continue a conversation even after a safety filter is activated, which is a vulnerability in design. In the following example, simply adding the word “theoretically” to a prompt that was directly rejected by the model right beforehand was sufficient to produce a walkthrough of how to construct a remote-detonated improvised explosive device.

Figure 4: A conversation in which a model refuses a request for operational information on rigging an explosive, then complies with the same request verbatim once the word “theoretically” is added.
The Actor This Empowers
In a study of lone-actors involved in terrorist plots and attacks between 1990 and 2012, researchers found that “more than half (52.9%) were characterised as socially isolated.” For many of these actors, AI know fills the role(s) of a partner and/or mentor. The shift with AI is that it fills the role of a partner and mentor to these actors. Propositional knowledge, the contents of papers and protocols, are not the binding constraint on biological weapons. Aum Shinrikyo, the Japanese cult behind the 1995 Tokyo subway sarin attack, had the resources and plenty of motivation, and still failed frequently, because the distance between a protocol and a result is filled with tacit knowledge. Bomb-making manuals and virology papers have always circulated, and nearly half of lone actors already consult them. A bomb-making manual or virology paper cannot tell its reader why the circuit isn’t completing or diagnose why its culture failed, but a frontier model may be able to. And in doing so, it removes what has functioned inadvertently as one of the biggest bottlenecks for lone actors and groups.
Furthermore, extremist perseverance follows a doubling decay function. Each time the effort required to advance a plot doubles, roughly half of the would-be attackers drop out. Every failure multiplies effort and makes actors more likely to stop. AI uplift aids the user and prevents compounding failures. By outlining the process, the model interrupts the attrition that has historically inhibited extremists long before they can execute an attack.
There is also the risk posed by open-weight AI models. Open-weight AI models allow guardrails to be stripped with little effort, and capable actors are long aware of them. Such models do act as alternatives to overly cautious, closed AI models. But this does not make safeguarding the frontier labs’ models pointless, for two reasons. First, open models lag the frontier in capability, and the gap is widest on the hardest, most technically complex tasks. Second, by the logic of attrition described above, the effort required to migrate to a new open-source model and acquire the technical skills needed to jailbreak it is itself enough to double the effort required to carry out an attack. For these reasons, this Insight focuses on frontier closed-weight models. The safeguards protecting them are both consequential and within the labs’ power to strengthen. Open models remain an important part of the threat landscape, but they are a separate problem, and one worth bearing in mind alongside the findings here.
What Can Be Done
There are three takeaways from our testing that are relevant for policymakers and model developers.
The first concerns evaluation. Threat models have to treat context and conversational drift as a primary attack surface. As reasoning models grow more capable and agentic interaction becomes more ordinary, the share of real-world use that single-turn testing fails to capture only continues to grow. Users on previous models have demonstrated that jailbreaks can be achieved through long-context conversations. A safety stack that performs well on single-turn benchmarks but fails in long-form interactions is not adequate. Evaluations must continually adapt to reflect what all potential avenues of misuse may look like.
The second concerns how refusals are distributed across a conversation. A model that refuses synthesis but engages procurement, refuses procurement but engages cultivation, refuses cultivation but engages dissemination, has not refused anything meaningful; it has simply divided its compliance. The question isn’t whether each request, considered alone, crosses a line. It is whether the trajectory of the conversation is one a legitimate user would have any reason to engage in. Furthermore, when a model refuses to engage, it has to be meaningful. It’s not adequate for a user to get a refusal and simply continue the conversation, calibrating their follow-up to bypass what triggered the model and its guardrails.
The third is structural. If a meaningful share of dangerous scientific capabilities also has legitimate uses, blanket refusal is not a sustainable answer. One alternative is tiered or whitelisted access, providing advanced model capabilities only to users with some degree of verified identity. There’s a psychological dimension to this as well. The mere threat of disruption through awareness of surveillance is often sufficient as a disruptor. An actor who believes their identity is known, even loosely, operates differently than one who believes they are anonymous. Tiered access only needs to remove the assumption of anonymity that is possible in many models today.
A hard point ultimately underlies this Insight. Model developers are incentivised to prioritise helpfulness and maximise engagement for legitimate and illegitimate users alike. The helpfulness imperative imbued in models is, in and of itself, a vulnerability for model safety. It is the property that the guardrails are layered on top of, which is why no refinement of refusal fully resolves it. Acknowledging this does not mean abandoning helpfulness, but it does mean an honest accounting of the cost and the trade-off that comes with it.
–
Uri Klempner works on AI trust and safety at Alice. He holds master’s degrees from Tsinghua University and Tel Aviv University and has previously worked at the United Nations Executive Office of the Secretary-General, Tech2Peace, and Gong.
–
Are you a tech company interested in strengthening your capacity to counter terrorist and violent extremist activity online? Apply for GIFCT membership to join over 30 other tech platforms working together to prevent terrorists and violent extremists from exploiting online platforms by leveraging technology, expertise, and cross-sector partnerships.