Review of Anthropic’s Unredacted Chemical and Biological Risk Report: Claude Opus 4.6
SecureBio partnered with Anthropic to review the chemical and biological (CB) risk sections of their February 2026 Risk Report. For our review, Anthropic shared with us an unredacted version of this report and 110 pages of additional materials. Over two months, they responded to over seventy follow-up questions and requests for materials regarding updated model capabilities and safeguards.
Read our full review of Anthropic’s Unredacted Chemical and Biological Risk Report: Claude Opus 4.6.
Summary of our Review
Overall, we agree with Anthropic that the current risk of catastrophic outcomes that are substantially enabled by Claude Opus 4.6 is:
- “Very low but not negligible” for non-novel CB weapons production,1 and
- “Low risk, but with substantial uncertainty” for novel CB weapons production.
We highlight a few key findings in particular:
On refusal classifiers: We examined the training constitutions and robustness results for Anthropic’s refusal classifiers, their main CB risk safeguard. In aggregate, we found that the classifiers generally refuse on topics which we would have refused on, including 94.2% of hazardous prompts in SecureBio’s safeguard benchmark BioTIER-refuse. However, we advise that the constitutions should be continually updated as advances in biotechnology and elsewhere change the landscape of topics that pose CB risk.
On user access: We also reviewed Anthropic’s process for (a) vetting users for exemptions from refusal classifiers and (b) monitoring such users. At this time, we find it unlikely but possible that a threat actor would successfully acquire and use such an exemption to substantially uplift CB weapons production.
On jailbreaks: We emphasize that risks remain non-negligible: actors with extremely high jailbreaking expertise can still quickly find or make substantial progress on jailbreaks around the refusal classifiers (e.g. UK AISI’s BPJ) and remediation times can be long.
Because of the similarity of safeguards between Opus 4.6 and later versions of Opus 4 (i.e. Opus 4.7 and Opus 4.8), we apply our conclusions to those later Opus 4 models as well. However, we do not apply our conclusions to Fable 5, Mythos 5, or Opus 5 due to their differing capabilities, safeguards, and access controls; we leave those models beyond this review’s scope.2
External review criteria for Responsible Scaling Policy
The Responsible Scaling Policy v3.0, released February 2024 and applying to model releases starting with Claude Opus 4.6, described criteria for external reviews (Section 3.63). In fulfillment of Section 3.6.3, covering the contents of the external review, we summarize our findings as follows.
- We found adequacy of information to be generally sufficient, although minor parts of our assessment depended on information from redactions and follow-ups (e.g. reviewing the classifier guard training constitutions).
- We found analytical rigor to be generally sufficient.
- We agreed with Anthropic’s overall conclusion about CB risk levels as stated above, but have minor areas of disagreement, most importantly regarding the risk posed by actors with high jailbreaking expertise.
- We make several risk reduction recommendations, including one that Anthropic had already implemented since the Risk Report (Anthropic has removed zero data retention for Mythos 5 and Fable 5). The other recommendations concern the monitoring of classifier-exempt users and the bio-classifier constitutions.
The full report contains more details on the review process, methodology, and findings.
SecureBio did not receive funding from Anthropic for this work.
Footnotes
This risk assessment does not claim that risk of novel-CB weapons production is more likely than non-novel CB weapons production, but rather that the threat model of novel-CB production is fundamentally more uncertain.↩︎
At the time of writing, Fable 5 did route biology-related queries to Opus 4.8 (implying some similarity of safeguards and capabilities in CB areas), but we stay conservative in applying the conclusions of this review to Fable 5.↩︎
The RSP has since been updated.↩︎