AI Biology Benchmarks

Uplift Studies

A growing body of research evaluates the uplift AI models provide to differently resourced actors attempting biological tasks. Studies vary widely in methodology, scope, and findings. This section reviews the current evidence base from empirical studies, with and without wet lab components. For a detailed discussion of the methodological challenges and practical solutions in designing uplift studies, see (Paskov et al. 2026).

Key findings

  • Wet lab vs. in silico gap. The two completed wet lab studies (Active Site, LANL) found no statistically significant uplift on their primary endpoints, while in silico1 studies have found minimal (OpenAI/Gryphon, RAND, Meta) to substantial (Scale AI/SecureBio, Anthropic) uplift on knowledge and planning tasks.
  • Focus on novice uplift. Both wet lab and in silico studies primarily focus on uplifting novices without any lab experience; expert or semi-expert uplift remains underexplored.
  • Small sample sizes. Most studies only have a handful of participants in each group, making measurements noisy and underpowered. Medium-to-large effect sizes would, however, be detectable by studies similar to the Active Site RCT.
  • Missing methodological details. Company-internal uplift studies often report only high level approaches and results while omitting key study metrics.
  • Rapid model evolution. Each study can only test a snapshot of rapidly improving model capabilities. Findings from 2023–2024 (GPT-4/Claude 3/Gemini 1.5 era) are less informative about present-day risks.
  • Control group contamination. Participants in the internet-only control group can either actively access AI models on their own devices or might passively encounter AI-generated content, making “internet-only” conditions porous, especially in longer studies where compliance is harder to enforce.

“Measuring Mid-2025 LLM-Assistance on Novice Performance in Biology” (Hong et al. 2026)

  • Organization: Active Site (formerly Panoplia Laboratories)
  • Methodology: Pre-registered, investigator-blinded, randomized controlled trial (RCT)
  • Participants: 153 novices (134 STEM, 19 non-STEM/other)
  • AI models tested: Frontier AI models available in summer 2025 (Claude 4 series, Gemini 2.5 series, GPT-4 series, o3, o4-mini) vs. internet-only
  • Tasks: Three core tasks modeling a viral reverse genetics workflow (cell culture, cloning, virus production), and two accessory tasks (pipetting, RNA quantification)
  • Duration: 8 weeks (39 four-hour sessions)
  • Key findings: No significant uplift on the primary endpoint, core workflow completion, where success rate was 5.2% in the AI model group vs 6.6% in the internet-only group. A modest, non-significant benefit (~1.4-fold increase) was observed on individual tasks and the cell culture step showed higher success rates (68.8% vs 55.3%).

“Measuring skill-based uplift from AI in a real biological laboratory” (Romero-Severson et al. 2025)

  • Organization: Los Alamos National Laboratory
  • Methodology: Pilot observational study with randomized group assignment
  • Participants: 10 LANL employees with no prior wet lab experience
  • AI models tested: o1 vs. internet-only
  • Tasks: Bacterial transformation with a proinsulin-encoding plasmid, selective plating, colony growth verification, IPTG induction, mass spectrometry sample preparation
  • Duration: 3 days
  • Key findings: Results consistent with uplift but not statistically significant; completion rate of the first attempt was 60% in the AI model group vs 20% in the internet-only group. After two attempts: 80% vs 60%. Expert guidance was provided when participants were stuck.

Commissioned study on AI-enabled biological risk

  • Organization: UK AI Security Institute, Imperial College London
  • Scope: A randomized controlled trial (RCT) investigating how the use of AI models can make synthetic biology accessible to non-experts in a laboratory. The study is supposed to compare (1) Google-based internet search, (2) frontier AI model access, and (3) human expert assistance.
  • Status: Commissioned via public tender. No public results as of March 2026.

“LLM Novice Uplift on Dual-Use, In Silico Biology Tasks” (Zhang et al. 2026)

  • Organization: Scale AI, SecureBio, University of Oxford, UC Berkeley
  • Methodology: Mixed design – non-STEM participants alternated between AI model/internet arm, STEM participants were assigned once
  • Participants: 57 novices (47 STEM, 10 non-STEM)
  • AI models tested: OpenAI o3, o4-mini, Gemini 2.5 Pro, Gemini Deep Research, Claude 3.7 Sonnet, Claude Opus 4 vs. internet-only
  • Tasks: Eight benchmark task sets covering virology, molecular biology, and general biology (Long-Form Virology, ABC-Bench, World Class Biology, VCT, HPCT, MBCT, LAB-Bench, Humanity’s Last Exam)
  • Duration: Up to 13 hours
  • Key findings: Novices with AI models were 4.16x more accurate than controls. AI model groups exceeded expert baselines on 3/4 benchmarks. 89.6% of participants reported no difficulty overcoming safeguards. Standalone AI models often outperformed AI-assisted novices.

Amazon commissioned an independent uplift study from Nemesys Insights as part of the Nova 2.0 Lite Frontier Model Safety Framework evaluation (Krishna et al. 2026).

  • Organization: Amazon, Nemesys Insights
  • Methodology: Unspecified; described as a “large-scale, independent uplift study” and “rigorous red-teaming exercise”
  • Participants: ~800
  • AI models tested: Nova 2.0 Lite
  • Tasks: Unspecified, but included attack planning across chemical, biological, radiological, and nuclear domains
  • Duration: Unspecified
  • Key findings: Model remains below the overall CBRN Critical Capability Threshold. However, a meaningful uplift was identified in radiological attack planning, resulting in additionally deployed filters and monitoring.

As part of their biorisk evaluation efforts, Anthropic has conducted various internal, small-scale uplift trials across multiple Claude models with removed safeguards.

  • Methodology: Controlled trials with grading rubrics by Deloitte biodefense experts. Control groups use internet only; treatment groups use Claude with safeguards removed.
  • Participants and tasks: Varying between model generations, but all tasks were text-based.
  • Duration: 2–4 days (per trial)
  • Key findings:
    • Claude 3.7 Sonnet (Feb 2025): Participants from SepalAI and Anthropic drafted bioweapon acquisition plans. Highest-scoring group received 57±20% fraction score, below the 80% ASL-3 threshold. Conclusion: some productivity enhancement, but does not translate into significant real-world risk increase. (System card)
    • Claude Opus 4 (May 2025): ~18 participants from SepalAI, Mercor, and Anthropic drafted bioweapon acquisition plans. 2.53x uplift in plan quality over internet-only controls. Participants scored much higher with substantially fewer critical failures. Below the internal 5x threshold, but sufficiently close that Anthropic could not rule out ASL-3 risks. (System card)
    • Claude Opus 4.5 (Nov 2025): PhD-level experts drafted a step-by-step virus acquisition protocol. The model was meaningfully more helpful than previous models, leading to higher scores and fewer critical errors, but plans still contained critical errors yielding non-viable protocols. (System card)
    • Claude Opus 4.6 (Feb 2026): PhD-level experts drafted a step-by-step virus acquisition protocol. The model group achieved lower scores than in the identical Claude Opus 4.5 trial. In addition, Anthropic also ran a creative biology uplift trial, in which 20 molecular biology PhDs (split into an AI model vs. internet-only group) were given 20 hours over 3 days to produce reports on novel biological workflows. The AI model group achieved ~2x performance vs controls, but no single plan was broadly judged by experts as highly creative or likely to succeed. (System card)

“Building an early warning system for LLM-aided biological threat creation”

  • Organization: OpenAI, Gryphon Scientific (now part of Deloitte)
  • Methodology: Randomized controlled trial (RCT) with A/B testing
  • Participants: 100 (50 biology PhD experts, 50 undergraduate-level students)
  • AI models tested: GPT-4 (including version without safety guardrails) vs. internet-only
  • Tasks: Five tasks assessing end-to-end critical knowledge of biothreat creation, evaluated on accuracy, completeness, innovation, time, and difficulty
  • Duration: 20–30 minutes per task
  • Key findings: Mild but not statistically significant uplift, though the statistical analysis has been criticized. Expert accuracy rose from 6.00 to 6.88 on a 10-point scale; completeness from 5.50 to 6.32. The student group as well as the other metrics showed smaller differences.

“The Operational Risks of AI in Large-Scale Biological Attacks” (Mouton et al. 2024)

  • Organization: RAND Corporation
  • Methodology: Red team RCT with Delphi-technique expert grading
  • Participants: ~45 researchers with diverse backgrounds across 15 red teams (8 with internet+AI model, 4 internet-only, 3 specialized compositions)
  • AI models tested: Two unnamed frontier AI models from summer 2023 vs. internet-only
  • Tasks: Four realistic vignettes for planning large-scale biological attacks
  • Duration: 7 weeks (max 80 hours/participant)
  • Key findings: No statistically significant uplift. All plans scored between “untenable” and “problematic.” AI-assisted plans were statistically indistinguishable from internet-only plans.

Meta has conducted an internal in silico uplift trial as part of the Llama 3 release (Grattafiori et al. 2024).

  • Organization: Meta, in collaboration with CBRNE experts
  • Methodology: RCT with Delphi-technique expert grading. Preliminary study validated the design including a power analysis.
  • Participants: Teams of two, recruited based on scientific or operational expertise. Teams classified as low-skill (no formal training) or moderate-skill (some formal training and practical experience).
  • AI models tested: Llama 3 70B and Llama 3 405B (with web search, RAG over hundreds of relevant papers, and code execution) vs. internet-only
  • Tasks: Generating operational plans for biological or chemical attacks, covering agent acquisition, production, weaponization, and delivery across multiple scenarios. Plans were evaluated by domain experts across four attack stages on metrics including scientific accuracy, detail, detection avoidance, and probability of success.
  • Duration: 6 hours per scenario
  • Key findings: No significant uplift in any condition – aggregate or by subgroup (model size, chemical vs. biological). Meta concluded “low risk that release of Llama 3 models will increase ecosystem risk related to biological or chemical weapon attacks.”

References

Grattafiori, Aaron, Abhimanyu Dubey, Abhinav Jauhri, et al. 2024. The Llama 3 Herd of Models. https://arxiv.org/abs/2407.21783.
Hong, Shen Zhou, Alex Kleinman, Alyssa Mathiowetz, et al. 2026. Measuring Mid-2025 LLM-Assistance on Novice Performance in Biology. https://arxiv.org/abs/2602.16703.
Krishna, Satyapriya, Matteo Memelli, Tong Wang, et al. 2026. Evaluating Nova 2.0 Lite Model Under Amazon’s Frontier Model Safety Framework. https://arxiv.org/abs/2601.19134.
Mouton, Christopher A., Caleb Lucas, and Ella Guest. 2024. The Operational Risks of AI in Large-Scale Biological Attacks: Results of a Red-Team Study. https://www.rand.org/pubs/research_reports/RRA2977-2.html.
Paskov, Patricia, Kevin Wei, Shen Zhou Hong, et al. 2026. RCTs & Human Uplift Studies: Methodological Challenges and Practical Solutions for Frontier AI Evaluation. https://arxiv.org/abs/2603.11001.
Romero-Severson, Ethan Obie, Tara Harvey, Nick Generous, and Phillip M. Mach. 2025. Measuring Skill-Based Uplift from AI in a Real Biological Laboratory. https://arxiv.org/abs/2512.10960.
Zhang, Chen Bo Calvin, Christina Q. Knight, Nicholas Kruus, et al. 2026. LLM Novice Uplift on Dual-Use, in Silico Biology Tasks. https://arxiv.org/abs/2602.23329.

Footnotes

  1. Experiments performed on a computer as opposed to in a wet lab.