AI Biology Benchmarks

Biosecurity Evaluations on Model Reports

Listed below are the biosecurity-relevant evaluations that AI firms have conducted on their models, drawn from published model reports1. These typically sit within CBRN2 risk assessments. The aim is to show how often and how thoroughly firms test these capabilities. We focus on frontier models, since they set the upper bound on current capabilities.

Further resources:

AI Safety Policy: Frontier Model Safety Framework (February 10, 2025)

Amazon’s Frontier Model Safety Framework focuses on severe risks unique to frontier AI models, with Critical Capability Thresholds that emphasize quantifying uplift in CBRN attacks, offensive cyber operations, and autonomous AI R&D.

📄 Nova 2 technical report, Nova 2.0 Lite Frontier Model Safety Framework Evaluation

The following biosecurity-relevant evaluations were conducted on Nova 2:

  • WMDP (CAIS et al.)
  • LAB-Bench: ProtocolQA (FutureHouse)
  • BioLP-bench (Igor Ivanov)
  • Expert red teaming and uplift study (Nemesys Insights)

Safety assessment: Amazon concluded that Nova 2 “remains below the critical threshold for CBRN weapons proliferation, indicating that its deployment does not pose an elevated security risk.” A meaningful performance uplift in radiological attack planning was noted, resulting in the deployment of additional filters and monitoring.

📄 Nova Premier model report

The following biosecurity-relevant evaluations were conducted on Nova Premier:

  • WMDP (CAIS et al.)
  • LAB-Bench: ProtocolQA (FutureHouse)
  • BioLP-bench (Igor Ivanov)
  • Expert red teaming (Carnegie Mellon’s Gomes Group, Nemesys Insights, Deloitte)
  • Uplift studies (Nemesys Insights)

Safety assessment: Amazon concluded that Nova Premier “remains below the critical threshold for CBRN weapons proliferation, indicating that its deployment does not pose an elevated security risk.”

AI Safety Policy: Responsible Scaling Policy (RSP, Version 3, February 2025)

Anthropic evaluates their models extensively for chemical and biological capabilities, specifically on how their models may help groups create and deploy CB weapons. They focus most on biological risks.

📄 Claude Opus 4.8 system card

Anthropic limited Claude Opus 4.8’s CBRN testing to automated evaluations, since the model does not push the capability frontier beyond Claude Mythos Preview (which was determined not to cross the CB-2 novel-weapons threshold). No expert red-teaming sessions or uplift trials were conducted for this release.

The following biosecurity-relevant evaluations were conducted on Claude Opus 4.8:

  • Long-form Virology (SecureBio)
  • VCT (SecureBio)
  • ABC-Bench: DNA Synthesis Screening Evasion (SecureBio)
  • Black-box RNA sequence design (Anthropic)
  • AAV capsid packaging prediction (Anthropic)
  • BioMysteryBench (Anthropic)
  • LAB-Bench: FigQA (FutureHouse)

Safety assessment: Deployed under ASL-3 protections according to Anthropic’s Responsible Scaling Policy. Anthropic concluded that Opus 4.8 does not advance the capability frontier beyond Claude Mythos Preview and does not cross the CB-2 threshold for novel chemical/biological weapons production, so catastrophic CBRN risks remain low under current mitigations.

📄 Claude Opus 4.7 system card

The following biosecurity-relevant evaluations were conducted on Claude Opus 4.7:

  • Expert red teaming (Anthropic)
  • Long-form Virology (SecureBio)
  • VCT (SecureBio)
  • ABC-Bench: DNA Synthesis Screening Evasion (SecureBio)
  • Sequence-to-function modeling and design (Anthropic)
  • BioPipelineBench (Anthropic)
  • BioMysteryBench (Anthropic)
  • LAB-Bench: FigQA (FutureHouse)

Safety assessment: Deployed under ASL-3 protections according to Anthropic’s Responsible Scaling Policy. Opus 4.7 did not cross the CB-2 threshold for novel chemical/biological weapons production; red-teamers assessed it as broadly less concerning than Claude Mythos Preview, characterizing it as a competent aggregator of published information that requires constant steering.

📄 Claude Mythos Preview system card

The following biosecurity-relevant evaluations were conducted on Claude Mythos Preview:

  • Expert red teaming (Anthropic)
  • Catastrophic biology scenario construction uplift trial (Anthropic)
  • Sequence-to-function modeling and design (Anthropic)
  • Long-form Virology (SecureBio, Deloitte, Signature Science)
  • VCT (SecureBio)
  • ABC-Bench: DNA Synthesis Screening Evasion (SecureBio)
  • BioMysteryBench (Anthropic)
  • LAB-Bench: FigQA (FutureHouse)

Safety assessment: Offered to partners under ASL-3 protections according to Anthropic’s Responsible Scaling Policy. The median expert assessed Mythos Preview as a “force-multiplier” and the model reached CB-1 level uplift, but Anthropic determined it had not reached the CB-2 threshold for novel chemical/biological weapons production.

📄 Claude Sonnet 4.6 system card

Anthropic ran automated evaluations only for Claude Sonnet 4.6; no human uplift trials or expert red-teaming sessions were conducted.

The following biosecurity-relevant evaluations were conducted on Claude Sonnet 4.6:

  • Long-form Virology (SecureBio)
  • VCT (SecureBio)
  • ABC-Bench: DNA Synthesis Screening Evasion (SecureBio)
  • Creative biology automated evaluation (SecureBio)
  • Short-horizon computational biology (Anthropic, Faculty.ai)
  • BioPipelineBench (Anthropic)
  • BioMysteryBench (Anthropic)
  • LAB-Bench: FigQA (FutureHouse)

Safety assessment: Deployed under ASL-3 protections according to Anthropic’s Responsible Scaling Policy. Sonnet 4.6 performed below previously released models on all CBRN evaluations and did not cross the ASL-4 rule-out threshold on the short-horizon computational biology evaluation, so it does not cross the CBRN-4 threshold.

📄 Claude Opus 4.6 system card

The following biosecurity-relevant evaluations were conducted on Claude Opus 4:

  • Long-form Virology (SecureBio, Deloitte, Signature Science)
  • VCT (SecureBio)
  • ABC-Bench: DNA Synthesis Screening Evasion (SecureBio)
  • Creative biology uplift trial (Anthropic)
  • WCB (SecureBio)
  • Short-horizon computational biology (Anthropic, Faculty.ai)
  • BioMysteryBench (Anthropic)
  • ASL-4 virology uplift trial (Deloitte)
  • ASL-4 expert red-teaming (Anthropic)
  • ASL-4 red teaming with the US CAISI (US CAISI)

Safety assessment: ASL-3 according to Anthropic’s Responsible Scaling Policy. Not meriting ASL-4 safeguards, though “the margin for future rule-outs is narrowing”.

📄 Claude Opus 4.5 system card

The following biosecurity-relevant evaluations were conducted on Claude Opus 4:

  • Long-form Virology (SecureBio, Deloitte, Signature Science)
  • VCT (SecureBio)
  • ABC-Bench: DNA Synthesis Screening Evasion (SecureBio)
  • LAB-Bench: FigQA, ProtocolQA, SeqQA, CloningScenarios (FutureHouse)
  • WCB (SecureBio)
  • Short-horizon computational biology (Anthropic, Faculty.ai)
  • Bioinformatics Evaluations (Anthropic)
  • ASL-4 virology uplift trial (Deloitte)
  • ASL-4 expert red-teaming (Anthropic)
  • ASL-4 red teaming with the US CAISI (US CAISI)

Safety assessment: ASL-3 according to Anthropic’s Responsible Scaling Policy. Targeted safety evaluations confirmed risk profile remains consistent with Claude Opus 4, with biological risk well below CBRN-4 thresholds.

📄 Claude Opus 4.1 system card

Claude Opus 4.1 did not undergo comprehensive new RSP evaluations as it did not meet criteria for re-evaluation relative to Claude Opus 4.5. Therefore their assessment focus was on ASL-4 rule -out evaluations and relied on automated benchmarks and evaluations (no uplift trials or expert red-teaming):

  • Long-form Virology (SecureBio, Deloitte, Signature Science)
  • LAB-Bench subset (FutureHouse)
  • ABC-Bench subset (SecureBio)
  • ASL-4 rule-out evaluations:
    • Short-horizon computational biology (Anthropic, Faculty.ai)
    • WCB (SecureBio)

Safety assessment: ASL-3 CBRN risk according to Anthropic’s Responsible Scaling Policy. Targeted safety evaluations confirmed risk profile remains consistent with Claude Opus 4, with biological risk well below ASL-4 thresholds.

📄 Claude Opus 4 system card

The following biosecurity-relevant evaluations were conducted on Claude Opus 4:

  • Bioweapons Acquisition Uplift Trial (Anthropic, Deloitte, SepalAI, Mercor)
  • Expert Red Teaming (Deloitte)
  • Long-form Virology (SecureBio, Signature Science, Deloitte)
  • VCT (SecureBio)
  • Bioweapons Knowledge Questions (Deloitte)
  • ABC-Bench subset (SecureBio)
  • LAB-Bench: FigQA, ProtocolQA, SeqQA, CloningScenarios (FutureHouse)
  • ASL-4 rule-out evaluations:
    • creative biology (SecureBio)
    • short-horizon computational biology tasks (Anthropic, Faculty.ai)
    • ASL-4 expert redteaming (Anthropic)

Safety assessment: ASL-3 CBRN risk according to Anthropic’s Responsible Scaling Policy.

📄 Claude Haiku 4.5 system card

The following biosecurity-relevant evaluations were conducted on Claude Haiku 4.5:

  • LAB-Bench: ProtocolQA, SeqQA, CloningScenarios, FigQA (FutureHouse)
  • VCT (SecureBio)
  • Long-form Virology (SecureBio, Signature Science, Deloitte)
  • ABC-Bench: DNA Synthesis Screening Evasion (SecureBio)

Safety assessment: ASL-2 CBRN risk according to Anthropic’s Responsible Scaling Policy

📄 Claude 3.7 Sonnet model report

The following biosecurity-relevant evaluations were conducted on Claude 3.7 Sonnet:

  • Bioweapons Acquisition Uplift Trial (Anthropic)
  • Expert Red Teaming (Deloitte)
  • Long-form Virology Tasks (SecureBio, Deloitte)
  • Multimodal Virology (VCT) (SecureBio)
  • Bioweapons Knowledge Questions (Anthropic)
  • LAB-Bench: FigQA, ProtocolQA, SeqQA, CloningScenarios (FutureHouse)

Safety assessment: ASL-2 CBRN risk according to Anthropic’s Responsible Scaling Policy

📄 Claude 3.5 Sonnet model report

The following biosecurity-relevant evaluations were conducted on Claude 3.5 Sonnet:

  • CBRN evaluations (Anthropic)
  • Pre-deployment evaluation (UK AISI, US AISI)

Safety assessment: ASL-2 CBRN risk according to Anthropic’s Responsible Scaling Policy. Claude 3.5 Sonnet did not exceed preset thresholds of concern during testing and was classified as ASL-2, indicating it does not pose risk of catastrophic harm.

AI Safety Policy: No published safety framework

📄 DeepSeek V3 technical report

No biosecurity-specific evaluations were conducted. DeepSeek does not publish dedicated CBRN safety reports.

📄 DeepSeek-R1 technical report

No biosecurity-specific evaluations were conducted. DeepSeek makes no mention of CBRN or misuse testing. The only biology-related benchmarks run are GPQA Diamond and MMLU.

AI Safety Policy: Frontier Safety Framework (Version 3.0, September 22, 2025)

Google DeepMind’s Frontier Safety Framework (FSF) addresses multiple risk areas including CBRN, and defines critical capability levels (CCLs) for each risk area.

📄 Gemini 3.5 Flash model card

As Gemini 3.5 Flash does not have meaningful new capabilities with respect to Frontier Safety compared to Gemini 3.1 Pro, it did not receive a separate full CBRN evaluation and refers to the Gemini 3 Pro FSF Report.

Safety assessment: Gemini 3.5 Flash has not reached any Critical Capability Level (CCL). Based on Gemini 3.1 Pro results, it is unlikely to reach the CBRN Uplift 1 CCL according to Google DeepMind’s Frontier Safety Framework.

📄 Gemini 3.1 Pro model card

The Gemini 3.1 Pro model card only reports updated scored on general capability evals and refers to the Gemini 3 Pro Frontier Safety Framework Report for safety evaluation details.

Safety assessment: Gemini 3.1 Pro (Deep Think mode) has not reached the CBRN Uplift 1 Critical Capability Level.

📄 Gemini 3 Pro model card, Gemini 3 Pro FSF Report, Gemini 3 Flash model card

As Gemini 3 Flash is based on Gemini 3 Pro, it did not receive a separate safety evaluation but refers to the Gemini 3 Pro FSF Report.

The following biosecurity-relevant evaluations were conducted on Gemini 3 Pro:

  • VCT (early version, SecureBio)
  • LAB-Bench: ProtocolQA, CloningScenarios, SeqQA (FutureHouse)
  • WMDP (biology and chemistry splits)
  • Single- and multi-turn open-ended question evaluations (Google DeepMind)
  • Expert red-teaming (Google DeepMind)
  • External scenario-based red teaming exercise

Google DeepMind also considered preliminary results from the (then) ongoing wet lab uplift trial conducted by Active Site (formerly Panoplia Laboratories).

Safety assessment: Gemini 3 Pro has not reached the CBRN Uplift 1 Critical Capability Level.

📄 Gemini 2.5 Deep Think model card

The following biosecurity-relevant evaluations were conducted on Gemini 2.5 Deep Think:

  • Open-ended questions (biological, radiological, and nuclear domains) (Google DeepMind)
  • Close-ended multiple choice questions (MCQs) (Google DeepMind)
  • VMQA single-choice (SecureBio, an early version of VCT)
  • LAB-Bench: ProtocolQA, CloningScenarios, SeqQA (FutureHouse)
  • WMDP (biology and chemistry splits)
  • Expert red-teaming (Google DeepMind)
  • External safety testing

Safety assessment: Google could not rule out that Gemini 2.5 Deep Think has reached the CBRN Uplift 1 Critical Capability Level (CCL), so took a precautionary approach and launched it with RAND SL2 security mitigations. Early warning threshold reached for CBRN Uplift Level 1 according to Google DeepMind’s Frontier Safety Framework (FSF). The Gemini 3 Pro FSF Report confirms that the CCL was not met by Gemini 2.5 Deep Think.

📄 Gemini 2.5 Pro model card

The following biosecurity-relevant evaluations were conducted on Gemini 2.5 Pro:

  • Close-ended multiple choice questions (MCQs) (Google DeepMind)
  • Open-ended questions (biological, radiological, and nuclear domains) (Google DeepMind)
  • VMQA single-choice (SecureBio, an early version of VCT)
  • LAB-Bench: ProtocolQA, CloningScenarios, SeqQA (FutureHouse)
  • WMDP (biology and chemistry splits)

Other biology-related benchmarks that were run without reporting the results on the biology-specific subsets:

  • Humanity’s Last Exam
  • MMLU Lite
  • GPQA Diamond

Safety assessment: CBRN Uplift Level 1 CCL has not been reached according to Google DeepMind’s Frontier Safety Framework

📄 Gemini 1.5 Pro technical report

The following biosecurity-relevant evaluations were conducted on Gemini 1.5:

  • Dolomites (Google DeepMind)
  • STEM QA with Context (Qasper dataset)
  • Internal CBRN evaluations (open-ended adversarial prompts with domain-expert raters; closed-ended, knowledge-based multiple-choice questions)
  • External red-teaming (third-party organizations)

No safety assessment given

AI Safety Policy: Frontier AI Framework (February 3, 2025)

Meta’s Frontier AI Framework focuses on cybersecurity threats and risks from chemical and biological weapons, using threat modeling exercises and risk threshold establishment to maintain acceptable risk parameters.

📄 Llama 4 model report

The following biosecurity-relevant evaluations were conducted on Llama 4:

  • CBRNE (Chemical, Biological, Radiological, Nuclear, and Explosive) evaluations (Meta)
  • Red teaming exercises for content policy violations

Safety assessment: Meta assessed whether Llama 4 could meaningfully increase the capabilities of malicious actors to plan or carry out attacks using CBRNE weapons. ==but unfortunately didn’t report any conclusion after their evaluation

📄 Llama 3.3 model report

The following biosecurity-relevant evaluations were conducted on Llama 3.3:

  • Uplift testing for chemical/biological weapons planning (Meta)

Other biology-related benchmarks that were run without reporting the results on the biology-specific subsets:

  • MMLU
  • GPQA Diamond

📄 Llama 3.1 technical report

The following biosecurity-relevant evaluations were conducted on Llama 3.1:

  • CBRNE uplift testing for chemical/biological weapons planning (Meta)

Safety assessment: Meta concluded that “there is a low risk that release of Llama 3 models will increase ecosystem risk related to biological or chemical weapon attacks.”

📄 Prompt Guard model report

No biosecurity-specific evaluations were conducted. The model report makes no mention of any CBRN testing.

📄 Llama Guard 3 model report

No biosecurity-specific evaluations were conducted, but:

Llama Guard 3’s default unsafe_categories specifies “chemical weapons (ex: nerve gas), biological weapons (ex: anthrax), radiological weapons (ex: salted bombs), nuclear weapons (ex: atomic warheads), and high-yield explosive weapons (ex: cluster munitions)” as unsafe under the ‘Indiscriminate weapons’ category.

AI Safety Policy: No published safety framework (committed to Frontier AI Safety Commitments, February 2025)

📄 Mistral Large 3 documentation

No biosecurity-specific evaluations were conducted.

📄 Mistral Medium 3 blog post

No biosecurity-specific evaluations were conducted. The only biology-related benchmarks that were run, but without reporting the results on the biology-specific subsets:

  • GPQA Diamond
  • MMLU Pro

📄 Mistral Large blog post

No biosecurity-specific evaluations were conducted. The blog post makes no mention of CBRN or misuse. The only biology-related benchmark that was run, but without reporting the results on the biology-specific subsets:

  • MMLU

AI Safety Policy: No published safety framework (includes safety section in technical reports)

📄 Kimi K2.5 technical report

No biosecurity-specific evaluations were conducted. The only biology-related benchmarks run were GPQA Diamond, MMLU, and HLE.

📄 Kimi K2 technical report

No biosecurity-specific evaluations were conducted. MoonshotAI mentions chemical and biological weapons as part of the misuse red-teaming. The only biology-related benchmarks run were GPQA Diamond, MMLU, and HLE.

AI Safety Policy: Preparedness Framework (Version 2, April 15, 2025)

OpenAI recognizes and addresses CBRN threats and misuse cases, prioritizing biological capability evaluations due to the higher potential severity of biological threats relative to chemical ones.

📄 GPT-5.5 system card

The following biosecurity-relevant evaluations were conducted on GPT-5.5:

  • Multimodal Troubleshooting Virology (SecureBio, early version of VCT)
  • Open-ended ProtocolQA (FutureHouse)
  • Tacit knowledge and troubleshooting multiple-choice dataset (OpenAI & Gryphon Scientific)
  • TroubleshootingBench (OpenAI)
  • Hard-negative protein binding prediction (OpenAI)
  • DNA sequence design for transcription factor binding (OpenAI)
  • External evaluations, including ABC-Bench and ABLE agentic task-based evaluations (SecureBio)
  • National security scenario testing (US CAISI)

Safety assessment: “High” capability in the Biological and Chemical domain according to OpenAI’s Preparedness Framework, activating the associated Preparedness safeguards.

📄 GPT-5.4 Thinking system card

OpenAI does not do a full re-assessment of GPT-5.4 and only refers to the GPT-5.2 system card and the GPT-5 system card.

Safety assessment (from the GPT-5.2 system card): “High” CBRN capability according to OpenAI’s Preparedness Framework.

📄 GPT-5.2 system card update

The following biosecurity-relevant evaluations were conducted on GPT-5.2 for the system card update:

  • Multimodal Troubleshooting Virology (SecureBio, early version of VCT)
  • Open-ended ProtocolQA (FutureHouse)
  • Tacit knowledge and troubleshooting multiple-choice dataset (OpenAI & Gryphon Scientific)
  • TroubleshootingBench (OpenAI)

Safety assessment: “High” CBRN capability according to OpenAI’s Preparedness Framework.

📄 GPT-5.1 system card addendum

OpenAI does not do a full re-assessment of GPT-5.1 and only refers to the original GPT-5 system card.

Safety assessment (from the GPT-5 system card): “High” CBRN capability according to OpenAI’s Preparedness Framework.

📄 GPT-5 system card

The following biosecurity-relevant evaluations were conducted on GPT-5:

  • Long-form Biological Risk Questions (Gryphon Scientific)
  • Multimodal Troubleshooting Virology (SecureBio, early version of VCT)
  • Open-ended ProtocolQA (FutureHouse)
  • Tacit knowledge and troubleshooting multiple-choice dataset (OpenAI & Gryphon Scientific)
  • TroubleshootingBench (OpenAI)
  • External evaluations by SecureBio:
    • Manual red-teaming
    • VCT
    • HPCT
    • MBCT
    • WCB
    • ABC-Bench
    • Long-form virology plasmid design
  • Expert red-teaming for bioweaponization (OpenAI)

Safety assessment: “High” CBRN capability according to OpenAI’s Preparedness Framework. Though they note:

“We do not have definitive evidence that this model could meaningfully help a novice to create severe biological harm, our defined threshold for High capability, and the model remains on the cusp of being able to reach this capability. We are treating the model as such primarily to ensure organizational readiness for future updates to gpt-5-thinking, which may further increase capabilities.”

📄 GPT-4.5 system card

The following biosecurity-relevant evaluations were conducted on GPT-4.5:

  • Long-form Biological Risk Questions (Gryphon Scientific)
  • WMDP (CAIS et al.)
  • Open-ended ProtocolQA (FutureHouse)
  • Tacit knowledge and troubleshooting dataset (OpenAI, Gryphon Scientific)
  • BioLP-bench (Igor Ivanov)
  • VCT (SecureBio)

Safety assessment: “Medium” CBRN risk according to OpenAI’s Preparedness Framework.

“GPT-4.5 can help experts with the operational planning of reproducing a known biological threat, which meets our medium risk threshold. Because such experts already have significant domain expertise, this risk is limited, but the capability may provide a leading indicator of future developments.”

📄 GPT-4o system card

The following biosecurity-relevant evaluations were conducted on GPT-4o:

  • Uplift studies (biology experts and students)
  • Questions covering main stages in the biological threat creation process (ideation, acquisition, magnification, formulation, and release) (OpenAI, Gryphon Scientific)
  • Tacit knowledge and troubleshooting dataset (OpenAI, Gryphon Scientific)

Safety assessment: “Low” CBRN risk according to OpenAI’s Preparedness Framework.

📄 o1 model report

The following biosecurity-relevant evaluations were conducted on o1:

  • Long-form Biological Risk Questions (Gryphon Scientific)
  • Biological Tooling (Ranger)
  • Multimodal Troubleshooting Virology (SecureBio, early version of VCT)
  • ProtocolQA Open-Ended (FutureHouse)
  • BioLP-Bench (Igor Ivanov)
  • Tacit knowledge troubleshooting (multiple choice) and brainstorming (open ended) (Gryphon Scientific)
  • Structured expert probing campaign – chem-bio novel design (Signature Science)

Also mentioned in the model report but without reported results:

  • GPQA-Bio
  • WMDP-Bio/Chem
  • Organic chemistry molecular structure dataset
  • Synthetic biology translation dataset

Safety assessment: “Medium” CBRN risk according to OpenAI’s Preparedness Framework

📄 o3-mini model report

The following biosecurity-relevant evaluations were conducted on o3-mini:

  • Long-form Biological Risk Questions (Gryphon Scientific)
    • Expert Comparisons (OpenAI)
    • Expert Probing (OpenAI)
  • Biological Tooling (Ranger)
  • Multimodal Troubleshooting Virology (SecureBio, early version of VCT)
  • BioLP-bench (Igor Ivanov)
  • ProtocolQA Open-Ended (FutureHouse)
  • Tacit knowledge troubleshooting and brainstorming (Gryphon Scientific)
  • Structured expert probing campaign – chem-bio novel design (Signature Science)

Also mentioned in the model report but without reported results:

  • GPQA-Bio
  • WMDP-Bio/Chem
  • Organic chemistry molecular structure dataset
  • Synthetic biology translation dataset

Safety assessment: “Medium” CBRN risk according to OpenAI’s Preparedness Framework

📄 o3 and o4-mini model report

The following biosecurity-relevant evaluations were conducted on o3 and o4-mini:

  • Long-form Biological Risk Questions (Gryphon Scientific)
  • Multimodal Troubleshooting Virology (SecureBio, early version of VCT)
  • ProtocolQA Open-Ended (FutureHouse)
  • Tacit Knowledge and Troubleshooting (Gryphon Scientific)

Safety assessment: “Medium” CBRN risk according to OpenAI’s Preparedness Framework

AI Safety Policy: Risk Management Framework (August 20, 2025)

The xAI Risk Management Framework acknowledges CBRN risks and states that xAI intends to benchmark on VCT, WMDP, LAB-Bench, and BioLP-bench in the future.

📄 Grok 4.1 model card

The following biosecurity-relevant evaluations were conducted on Grok 4.1:

  • WMDP (CAIS et al.)
  • VCT text-only subset (SecureBio)
  • BioLP-bench (Igor Ivanov)
  • LAB-Bench: ProtocolQA, FigQA, CloningScenarios (FutureHouse)

Safety assessment: No formalized risk assessment was given, but xAI notes that “Grok 4.1 achieves broadly similar results to Grok 4.”

📄 Grok 4 Fast model card

The following biosecurity-relevant evaluations were conducted on Grok 4 Fast:

  • WMDP (CAIS et al.)
  • VCT text-only subset (SecureBio)
  • BioLP-bench (Igor Ivanov)

Safety assessment: No formalized risk assessment was given, and xAI notes that the risk remains unchanged from that of Grok 4, resulting in the deployment of the same filters.

📄 Grok 4 model card

The following biosecurity-relevant evaluations were conducted on Grok 4:

  • WMDP (CAIS et al.)
  • VCT text-only subset (SecureBio)
  • BioLP-bench (Igor Ivanov)

Safety assessment: No formalized risk assessment was given, but xAI notes that “due to Grok 4’s strong dual-use biological capabilities, we have deployed narrow, topically-focused filters across all product surfaces as an additional safeguard against bioweapons-related abuse.”

📄 Grok 3 announcement

No biosecurity-specific evaluations were conducted on Grok 3. The model report makes no direct mention of any CBRN or misuse testing.

Footnotes

  1. A model report (sometimes also called model card, system card, or safety report) is a document that provides key information about a machine learning model, typically including information about the model’s architecture, capabilities, and safety measures.

  2. Chemical, biological, radiological, and nuclear. Usually refers to threats and hazards posed by these categories of weapons.