Benchmark Trends
Last Updated July 17, 2026
Performance of publicly available models on biosecurity-relevant benchmarks. All results generated by SecureBio.
The plot above shows the Bio Capabilities Index (BCI)1 of all evaluated models over time. BCI is an aggregate score that distills model performance across multiple biology benchmarks into a single comparable number—akin to the model’s “biology IQ”.
Method: The BCI follows the exact methodology of the ECI, which is described in the paper “A Rosetta Stone for AI Benchmarks” (Ho et al. 2025). The BCI fits an Item Response Theory (IRT) sigmoid model to the observed score matrix:
\text{score}(m, b) = \sigma\bigl(\alpha_b \cdot (C_m - D_b)\bigr)
where C_m is a model’s capability parameter, D_b is a benchmark’s difficulty, and \alpha_b is its discrimination (how well it separates strong from weak models). VCT is used as the anchor benchmark (D = 0, \alpha = 1). Parameters are fitted via least-squares with L2 regularisation.
Included benchmarks: VCT, HPCT, MBCT, WCB, ABC-Bench, ABLE, LB-CloningScenarios, LB-LitQA2, LB-ProtocolQA, GPQA-Bio, MMLU-Bio, PubMedQA, and WMDP-Bio. Models must have scores on at least 4 of these 13 benchmarks to receive a BCI score. All benchmark scores were obtained by SecureBio, no outside scores are used in the BCI calculation.
Interpretation: A higher BCI indicates a stronger aggregate biology capability. The scale is linearly anchored to GPT-4 = 70 and Claude 3.5 Sonnet = 100. The confidence intervals shown in the datapoints’ hoverbox indicate sensitivity of the score to the fitted parameters. The frontier line tracks the highest BCI achieved by any model over time.
Limitations: The BCI shows much wider confidence intervals compared to the ECI, primarily due to the low number of biology benchmarks determining the score. Furthermore, some of these benchmarks are close to saturation, thus contributing more noise. This will be partially mitigated by including additional benchmarks with a wider dynamic range in the future. However, the BCI will continue to be a somewhat noisy measurement of dual-use biology capabilities, and the ECI should be consulted to understand a model’s overall capabilities.
The plot above shows how the maximum accuracy on available biology benchmarks evolved over time. The accuracy is normalized to the human expert baseline (0%, red dashed line) and a perfect score (100%) to make performance trends comparable across benchmarks.
Using the drop-down menu above the legend, you can switch between different accuracy measurements (normalized to expert baseline + perfect score, normalized to just expert baseline, or raw accuracy scores). Hover over any data point for additional information, or click a benchmark in the legend to isolate it (click additional entries to combine). While a filter is active, a Reset button appears at the top-right of the plot to restore the full view.
Please note that this plot only shows the top performing model over time on each benchmark. If you can’t find a new model on the plot, it means that it did not outperform the previous top-performing model on that benchmark. You can find the scores of all models on all benchmarks on the Benchmark Details page.
The plot above compares the frontier performance of open-weight and closed-weight models on biosecurity-relevant benchmarks over time. Solid lines represent the best closed-weight model performance, while dashed lines represent the best open-weight model performance, for each benchmark. Same-colored lines correspond to the same benchmark.
Using the drop-down menu above the legend, you can switch between displaying the accuracy or the accuracy gap on the benchmarks. The “Accuracy gap” view shows the difference (in percentage points) between the closed-weight and open-weight frontier on each benchmark. A positive gap indicates that closed-weight models are scoring higher on the benchmark. Hover over any data point for additional information, or click a benchmark in the legend to isolate it (click additional entries to combine). While a filter is active, a Reset button appears at the top-right of the plot to restore the full view.
Overall trends
- On most biology tasks that have been measured, leading AI models perform at or above human-expert level, and have been doing so on most biology benchmarks since 2024. Examples include troubleshooting laboratory experiments, designing DNA sequences for working with viral pathogens, and writing code to drive protein-design tools and lab automation.
- The same models often refuse to answer biosecurity-related questions, but the safeguards have notable gaps between providers.
- It is still unclear how much these capabilities translate into real-world ability to cause biological harm, but there is ongoing work in the industry to make this more well-defined.
Three patterns stand out across the benchmarks:
- Capabilities are rising over time, in both general biology and the biosecurity-relevant areas tested.
- Some recent models score lower than older ones in their family. Companies that have added chemical, biological, radiological, and nuclear (CBRN) safeguards have published models that deliberately decline to answer many biology questions, pulling their scores below the upward trend.
- Even with refusals factored into both of the above, capabilities are still rising for questions that models will answer.
The rise is clearest on the Frontier performance tab above: leading models started passing human experts on most biology benchmarks in 2024. Both the long-run rise and the recent CBRN-safeguard dip show up in the Bio Capabilities Index (BCI), a single score that aggregates a model’s results across all benchmarks.
References
Footnotes
This index is inspired by and uses the same methodology as Epoch AI’s Capabilities Index (ECI), but only includes biosecurity-relevant benchmarks, whereas the ECI is a measure of the “general capability” of a model.↩