← Back to Capabilities

AI Systems Provide Expert-Level Guidance on Dangerous Materials

Frontier AI models now score at or above PhD level on biosecurity and chemistry benchmarks. Anthropic activated AI Safety Level 3 protections for Claude Opus 4 in May 2025, citing inability to rule out CBRN uplift at threshold. The UK AI Security Institute reports AI-assisted novices were roughly five times more likely to produce feasible experimental protocols for viral recovery than internet-only controls.

Why this matters

Chemical, biological, radiological, and nuclear weapons are gated by a combination of information, materials, and tacit skill. For decades, biosecurity policy has assumed that information was the easiest barrier, because the others are physically scarce. AI systems challenge that assumption by making expert-level scientific guidance available to anyone with a prompt.

The capability is not hypothetical. The WMDP benchmark (3,668 questions across biosecurity, chemical security, and cybersecurity) is now standard. Frontier models score well above random and correlate strongly with general capability. Brent and McKelvey demonstrated in 2025 that multiple models - including Llama 3.1 405B, an open-weight model - could accurately guide users through a specific high-consequence recovery procedure using commercially obtainable synthetic DNA. The dangerous knowledge surface is partially in open weights already, challenging the assumption that lab-level safeguards are sufficient.

Evidence at this layer is deliberately withheld in detail by frontier labs for responsible disclosure. Anthropic's language - "cannot rule out ASL-3" - is itself an artefact of that regime. It means the capability may be at threshold but cannot be publicly confirmed. The open-weights finding matters because it changes the lever: refusal training on closed models cannot remove knowledge that is already present in models running on consumer hardware.

Documented incidents

View all documented incidents →

Evidence timeline

Discussed in Theory

2002 peer-reviewed 3,000+ citations

Existential Risks: Analyzing Human Extinction Scenarios. First rigorous philosophical treatment of anthropogenic existential risk as a category, specifically naming biotechnology and machine intelligence as the technological classes most likely to generate extinction-level events.

Bostrom, Journal of Evolution and Technology →
2018 peer-reviewed

Biodefense in the Age of Synthetic Biology. Establishes the formal Dual Use Research of Concern (DURC) framework adopted by US federal policy. Defines the analytic structure later used to reason about whether LLMs constitute a dual-use concern.

National Academies Press, DoD commissioned →
2020

The Precipice. Quantifies anthropogenic existential risk, giving roughly 1-in-30 odds of an engineered pandemic causing existential catastrophe this century. Frames why CBRN uplift from AI is a first-tier concern rather than a theoretical curiosity.

Ord, Bloomsbury →

Demonstrated in Lab

2024 peer-reviewed 200+ citations

WMDP Benchmark. 3,668-question multiple-choice proxy benchmark covering biosecurity, chemical security, and cybersecurity hazardous knowledge. Frontier models score well above random and correlate strongly with general capability. The authors deliberately filtered sensitive content before release.

Li, Pan, Gopal et al. (CAIS consortium), ICML 2024 →
2024

LAB-Bench: Measuring Capabilities of Language Models for Biology Research. 2,400+ questions across 8 practical biology-research categories. UK AISI reports frontier models now surpass PhD-holders on biology-knowledge proxy sets, and by 2025 were outperforming virology experts on lab-troubleshooting scenarios.

Laurent et al. (FutureHouse), arXiv 2024 →
2025

Contemporary AI Foundation Models Increase Biological Weapons Risk. Demonstrated that multiple frontier models (Llama 3.1 405B, GPT-4o, Claude 3.5 Sonnet) can accurately guide users through a specific high-consequence recovery procedure from commercially obtainable synthetic DNA. Challenges the assumption that tacit knowledge is an insurmountable barrier.

Brent & McKelvey (RAND), arXiv 2025 →

Demonstrated in Real World

May 2025

Launched Claude Opus 4 under AI Safety Level 3 standards, citing inability to rule out CBRN uplift at threshold. In a bioweapons-acquisition uplift trial, participants assisted by Opus 4 (safeguards removed, experimental condition) achieved a 2.53x capability increase over an internet-only control group on a planning task.

Anthropic, ASL-3 activation announcement →
2024-2025

Real-world wet-lab study found AI-assisted novices were roughly five times more likely to produce feasible experimental protocols for viral recovery than internet-only controls. Jailbreak resistance improved from ~10 minutes of expert effort to over 7 hours across a six-month window.

UK AISI Frontier AI Trends Report →
2024-2025

DeepMind Frontier Safety Framework v3 formally defines CBRN as a risk domain with staged Critical Capability Levels. OpenAI Preparedness Framework v2 lists Biological and Chemical as Tracked Categories with mature evaluations. Frontier-lab governance structures now treat CBRN uplift as standing operational concern, not hypothetical.

Frontier lab governance frameworks →

Strongest Counterargument

Textbook and protocol knowledge - even expert-level - is not operational capability. Biological and chemical weapons development is gated by physical bottlenecks: access to pathogens or precursors, specialised equipment, biocontainment, live-culture handling, and tacit knowledge acquired only through hands-on mentorship. Information uplift that does not cross this bottleneck does not meaningfully change the threat landscape.

Source: S.A. Norwood (2025) argues the design stage of bioweapons development is where AI might assist, but 'the transition to the physical world is a significant pinch point.' Consistent with Sonia Ben Ouagrham-Gormley's 'Barriers to Bioweapons' (2014) and RAND's initial 2024 red-team finding of no statistically significant improvement. A 2025 pre-registered wet-lab RCT (n=153) found mid-2025 LLMs did not substantially increase novice completion of complex laboratory procedures.

Why this deserves weight: The counterargument is empirically grounded. In silico benchmark performance has historically over-predicted real-world operational capability in this domain. Historical bioweapons programmes (Aum Shinrikyo, Soviet Biopreparat) failed despite motivated, resourced actors with more than textbook knowledge. The gap between 'model scores high on LAB-Bench' and 'attacker successfully causes mass harm' is real. However, the Brent & McKelvey (2025) work and AISI's 2025 wet-lab findings suggest the bottleneck is being eroded rather than fixed. The counterargument narrows the concern - it does not eliminate it.