BioTIER: Biological Targeted Information for Exclusion and Refusal

Published April 2026

Authors: Eleanor M. Marshall, Pedro Medeiros, Peter Peneder, Nelly Mak, Jacob Kaffey, Mac Walker, Seth Donoughe, Jasper Götting

Update (July 2026) — Our publication, BioTIER: A Refusal Benchmark for Targeted Biological Risk Mitigation, is now available on arXiv. Read our blog post for more information.

Introduction

BioTIER (Biological Targeted Information for Exclusion and Refusal) is a resource to help AI developers implement robust, but narrowly targeted, safety mitigations for dual-use biological capabilities of AI models. Developed by a team of researchers with input from experts from a broad range of dual-use subjects, it provides industry-leading resources for identifying, mitigating, and evaluating biological misuse risk in large language models.

BioTIER includes recommended safeguard policies, a benchmark for measuring safeguard performance across all included topics, and supporting data that can be helpful for excluding a small set of particularly hazardous data from pre-training and/or implementing post-training mitigations.

  1. BioTIER Policy
    • Natural language descriptions of all three risk sets and 53 “themes” for informing content filters
  2. BioTIER Benchmark
    • 542 expert-written prompts distributed across the three risk sets to be used for measuring refusal performance
    • Includes two labeled subsets: BioTIER-refuse (prompts that we recommend should be refused by models) and BioTIER-permit (prompts that we recommend should be answered by models)

Risks are classified by topic

We categorize all biological content into three sets to allow for nuanced model policies:

Risk Set Description Pretraining Recommendation Refusal Recommendation
Set CA: Catastrophe Avoidance Highest priority for refusal. Topics include: obtaining live pathogens, bypassing biodefense guardrails, pandemic potential assessment, methods for the development of mirror life. Exclude from all pretraining data. [BioTIER-refuse] Refuse for general users; permit for selected safety researchers
Set BD: Biomedical DURC Dual-use research of concern (DURC) with substantial misuse potential. Topics include: pathogen production and storage, passaging, virus engineering. Include for managed-access models; consider exclusion for open-weights models. [BioTIER-refuse] Refuse for general users; permit for institutional users passing Know Your Customer (KYC) checks. Log queries.
Set RB: Related Biology Biological content without acute dual-use risk. General biology knowledge, molecular biology, biochemistry, cell biology, microbiology, and other topics without acute biorisk. Include for all models. [BioTIER-permit] Do not refuse for any user.

Within sets CA and BD (i.e. the BioTIER-refuse subset), some prompts are additionally labelled with the secondary tag “SA: Select Agents”, which indicates content specific to Select Agents. Select Agents refers to agents and toxins listed on the HHS and USDA Select Agents and toxins list and the Australia Group List of Human and Animal Pathogens and Toxins. This SA tag enables targeted evaluation of safeguard performance on prompts involving these agents.

Performance of selected models

See the AI Biology Benchmarks dashboard for up-to-date results.

Coverage and development of prompts

The BioTIER Benchmark was designed to represent a wide range of linguistic styles, with prompts having varying degrees of sophistication, practicality, creativity, and length across topics. All prompts were written by subject matter experts and underwent multiple rounds of review and quality control steps prior to inclusion. Quantitative analyses were applied to validate dataset quality and coverage, including:

  • Textual diversity analysis across sophistication, practicality, creativity, and length dimensions
  • Coverage assessment against reference topic and framing databases
  • Iterative gap analysis with targeted writing to fill identified coverage gaps

For inquiries, please contact ai@securebio.org. We are also able to develop training data aligned with the BioTIER framework, as well as tailored evaluations for specific use cases.