AI Biology Benchmarks
About
FAQ
References
Brent, Roger, and T. Greg McKelvey Jr. 2025. Contemporary AI Foundation Models Increase Biological Weapons Risk. https://arxiv.org/abs/2506.13798.
Gopal, Anjali, Nathan Helm-Burger, Lennart Justen, et al. 2023. Will Releasing the Weights of Future Large Language Models Grant Widespread Access to Pandemic Agents? https://arxiv.org/abs/2310.18233.
Götting, Jasper, Pedro Medeiros, Jon G Sanders, et al. 2025. Virology Capabilities Test (VCT): A Multimodal Virology q&a Benchmark. https://arxiv.org/abs/2504.16137.
Grattafiori, Aaron, Abhimanyu Dubey, Abhinav Jauhri, et al. 2024. The Llama 3 Herd of Models. https://arxiv.org/abs/2407.21783.
Hendrycks, Dan, Collin Burns, Steven Basart, et al. 2021. Measuring Massive Multitask Language Understanding. https://arxiv.org/abs/2009.03300.
Ho, Anson, Jean-Stanislas Denain, David Atanasov, Samuel Albanie, and Rohin Shah. 2025. A Rosetta Stone for AI Benchmarks. https://arxiv.org/abs/2512.00193.
Hong, Shen Zhou, Alex Kleinman, Alyssa Mathiowetz, et al. 2026. Measuring Mid-2025 LLM-Assistance on Novice Performance in Biology. https://arxiv.org/abs/2602.16703.
Jin, Qiao, Bhuwan Dhingra, Zhengping Liu, William W. Cohen, and Xinghua Lu. 2019. PubMedQA: A Dataset for Biomedical Research Question Answering. https://arxiv.org/abs/1909.06146.
Krishna, Satyapriya, Matteo Memelli, Tong Wang, et al. 2026. Evaluating Nova 2.0 Lite Model Under Amazon’s Frontier Model Safety Framework. https://arxiv.org/abs/2601.19134.
Laurent, Jon M., Joseph D. Janizek, Michael Ruzo, et al. 2024. LAB-Bench: Measuring Capabilities of Language Models for Biology Research. https://arxiv.org/abs/2407.10362.
Li, Nathaniel, Alexander Pan, Anjali Gopal, et al. 2024. The WMDP Benchmark: Measuring and Reducing Malicious Use with Unlearning. https://arxiv.org/abs/2403.03218.
Marshall, Eleanor M., Pedro Medeiros, Peter Peneder, et al. 2026. BioTIER: A Refusal Benchmark for Targeted Biological Risk Mitigation. https://arxiv.org/abs/2607.14479.
Mouton, Christopher A., Caleb Lucas, and Ella Guest. 2024. The Operational Risks of AI in Large-Scale Biological Attacks: Results of a Red-Team Study. https://www.rand.org/pubs/research_reports/RRA2977-2.html.
Paskov, Patricia, Kevin Wei, Shen Zhou Hong, et al. 2026. RCTs & Human Uplift Studies: Methodological Challenges and Practical Solutions for Frontier AI Evaluation. https://arxiv.org/abs/2603.11001.
Peppin, Aidan, Anka Reuel, Stephen Casper, et al. 2024. The Reality of AI and Biorisk. https://arxiv.org/abs/2412.01946.
Phan, Long, Alice Gatti, Ziwen Han, et al. 2025. “Humanity’s Last Exam.” Nature, ahead of print. https://doi.org/10.1038/s41586-025-09962-4.
Rein, David, Betty Li Hou, Asa Cooper Stickland, et al. 2023. GPQA: A Graduate-Level Google-Proof q&a Benchmark. https://arxiv.org/abs/2311.12022.
Romero-Severson, Ethan Obie, Tara Harvey, Nick Generous, and Phillip M. Mach. 2025. Measuring Skill-Based Uplift from AI in a Real Biological Laboratory. https://arxiv.org/abs/2512.10960.
Soice, Emily H., Rafael Rocha, Kimberlee Cordova, Michael Specter, and Kevin M. Esvelt. 2023. Can Large Language Models Democratize Access to Dual-Use Biotechnology? https://arxiv.org/abs/2306.03809.
Zhang, Chen Bo Calvin, Christina Q. Knight, Nicholas Kruus, et al. 2026. LLM Novice Uplift on Dual-Use, in Silico Biology Tasks. https://arxiv.org/abs/2602.23329.