The Role of Evals in the Biorisk Evidence Hierarchy
We frequently hear questions along the lines of “what is the theory of change for benchmarks?“ or “evals aren’t measuring what we actually care about, so why the effort?“ Tl;dr: evals provide better evidence than just arguing from first principles, they are much cheaper and repeatable than wet lab uplift studies, and can help us to rule out capabilities.
A common critique of AI evals and benchmarks that try to measure AI-enabled biorisk is the lack of construct validity: “We want to know if AI increases biorisk and, if it does, convince the relevant stakeholders that the risk can be mitigated. But the benchmarks don’t measure the actual bottlenecks malicious actors face in the real-world. Therefore, to uncover the impact of AI on the biorisk landscape and convince the relevant people, real-world uplift studies that measure success in genuine wet labs are needed.”
An uplift study measures whether human participants are more effective and successful at a task when they have access to AI. I agree that uplift studies on biological tasks would be very strongest evidence for AI-enabled biorisk, but ultimately, there is a hierarchy of evidence, and evaluations occupy a specific and important position within it.
The Hierarchy

ChatGPT’s interpretation of this post after appealing to its inner artist.
Tier 1 — Theoretical arguments (we have been here)
Arguing that AI increases biorisk merely from first principles has been surprisingly effective, at least in the beginning. One could gesture towards the hurdles and failure modes of past bioweapon attempts, interview experts, perform a simple demonstration at the White House, run (super)forecasting exercises, point out the utility of AI in other domains like coding, and those who are security-minded will start to take the possibility of AI-enabled biorisk seriously and then begin to look for additional evidence. This is how I and many others ended up in the field. Theoretical arguments also justified investment in the higher tiers (as laid out below). They are very much load bearing and I regard the early work on arguing for AI-enabled biorisks extremely highly!
However, while easy and cheap to produce, such arguments alone are not sufficient to convince anyone who’s even slightly skeptical of domain-transfer claims or who needs a solid basis for making consequential decisions.
Tier 2 — Evals (we are here)
Evals1 are in silico estimators of an underlying model capability. They are imperfect and often noisy, but running a set of many high-quality evals can collectively point out even subtle differences between models.
This is a very broad tier that can encompass a wide variety of tools to evaluate a model: simple multiple-choice biology exams, open-ended research questions validated against experimental ground truths, experts manually interrogating models over many turns to test for the presence of rare knowledge, or complex agentic tasks in a simulated research environment. Such variation in content also comes with large differences in harm specificity, difficulty, and construct validity: for instance, an agentic benchmark testing the ability of a model to re-design an actual pathogen genome provides better and more convincing evidence than a multiple-choice test on college virology.
The steelman of eval critiques is not that they’re completely useless, it’s that they have several known failure modes that limit their meaningfulness:
Specification gaming: Because evals are becoming decision-relevant (e.g. for frontier AI safety policy tiers or regulatory thresholds), incentives for model developers change, and the gap between eval performance and underlying capability can widen rather than narrow. Goodhart’s law is obviously not a problem unique to evals, but the stakes here are unusually asymmetric. It is not hard to imagine that an AI firm could suppress performance on a dangerous capability eval while underlying capabilities continue to grow.
Underelicitation: Eval scores are typically lower bounds on capability, as factors like scaffolding (e.g. which tools, like web search, the model can use), fine-tuning (e.g. whether the model was additionally trained on domain-specific data), and prompt/context engineering (what other information the model sees before responding) can alter performance substantially. A motivated user with time and resources is likely able to elicit more performance from a model than a standard eval run does. Eval awareness and sandbagging (i.e. the model realizing it’s being evaluated and concealing capabilities) is another, increasingly relevant source of uncertainty around eval results that hasn’t yet been conclusively solved.
Leakage: One way in which eval scores might overestimate the underlying capability is leakage of the answers into the model training data and subsequent memorization. Private holdout sets and canary strings remain important to catch those instances.
Construct validity: A multiple-choice benchmark about textbook biology measures something, but what it’s measuring isn’t obviously providing a clear signal about the ability to coach someone through the successful creation of a pathogen in an actual lab. Some evals may only weakly approximate the real-world capability they’re meant to proxy, especially as models improve along orthogonal dimensions associated with solving benchmarks (e.g. picking up patterns in tests). Rigorous QC with realistic tasks grounded in real-world bottlenecks can improve the validity of an eval, but as things currently stand, the linkage between eval performance and harm potential is poorly understood.
Non-linear threshold effects: While performance on benchmarks and derived metrics tends to scale predictably, crossing certain capabilities thresholds can lead to saltatory real-world impacts (e.g. the meteoric rise of coding agents immediately after they became good enough for many real-world SWE tasks). It is hard—maybe impossible—to say in advance what these thresholds might be for biology and how benchmark performance relates to such sudden jumps in utility.
Infohazard tradeoffs: A highly valid eval should be harm-specific, i.e. measuring a specific threat model of concern as closely as possible. This might necessitate keeping the eval dataset fully private and foregoing external scrutiny. Private sharing with third-party auditors to enable reproduction and detailed writeups of the eval’s methodology mitigates this a little bit, but it remains a trade-off.
Saturation: Benchmarks have the unfortunate tendency to saturate, requiring the design and development of more difficult successors, but they typically still last a few model generations. This is fine, as it’s possible to establish an uninterrupted chain of model capability measurements as long as models and evals overlap, but it limits the depth of validation any single eval can receive before it’s retired. And even saturated evals can be meaningful indicators (more on that in a separate post).
Yet, despite all their flaws, evals provide the best cheap, comparable and repeatable evidence we can get. Developing a benchmark costs tens to hundreds of thousands of dollars (compared to millions for a well-powered uplift study) and running it on newly released models is a fraction of that cost (though possibly still significant when using heavy inference time scaling). In this, evals are akin to in vitro studies in drug discovery. An intermediate between pure mechanistic reasoning and full clinical trials born from inevitable trade-offs between validity and cost.
Last but not least, benchmarks can serve as rule-out evidence. A properly elicited model performing badly on an eval and clearly lacking certain knowledge increases our confidence that it poses no corresponding real-world risk. It is important to note, however, that this real-world biorisk creates a catastrophic asymmetry: the cost of a false-negative (an erroneous rule-out) is much, much higher than the cost of a false-positive.
Tier 3 — Real-world uplift studies (we are getting here)
Past evals, you run head-first into a cost cliff2. While it is possible to collect evidence in the real-world for cheap by recruiting 10 friends and renting a lab for a weekend, these small-scale experiments are hardly better than anecdata. A proper real-world uplift study with enough statistical power requires a lot of participants and time (read: money). However, done correctly, uplift RCTs do present the most convincing evidence for AI‘s impact on real-world lab success.
The ‘done correctly’ is an important point, though, and hinges on a multitude of factors, from the aforementioned number of participants, target demographic, task difficulty and harm specificity, trial duration, LLM utilization, porosity of the non-LLM control group to ubiquitous AI integrations, etc.
Getting these parameters right (e.g. via smaller pilots) and implementing them well is a difficult, time-consuming, and expensive endeavor that—as of now—safety researchers might repeat annually but not for every single model release or even model generation. Being slightly off on these parameters might produce false negatives and incorrect conclusions; results might be read as “no risk” when the actual read should’ve been “this design can’t see the risk.” Even a really well designed and executed uplift study might get dismissed or not counted as convincing evidence if the participant demographic and task type didn’t match somebody’s specific threat model of concern (e.g. you demonstrated some uplift in novices but the relevant stakeholder is mostly worried about semi-experts).
I say “we are getting here“ because in my opinion, of all the uplift studies done so far, only the recent Active Site RCT qualifies as a sufficiently large study. They are planning to run more follow ups to address other threat models and methodological pitfalls, so watch, fund, and innovate in this space!
Revisiting the rocky terrain between tiers 2 and 3
Assume you ran your large uplift study, found some effect, and two months later, the new generation of frontier models is released. Can you make any sensible claim about the impact on real-world risk of these models (besides “probably as good as the last gen“) without re-running everything? Maybe! You carefully approach the edge of the cost cliff you valiantly scaled and peer down to the vast plain of evals again.
If rich, qualitative data on LLM usage were collected during the uplift study, you should in principle be able to map the failures and AI-assisted successes of your study participants onto benchmarks measuring a corresponding capability. If model A failed to guide participants around a common methodological pitfall, and model B didn’t, is the same pattern detectable in a benchmark task about this method? Is the delta in overall uplift-effectiveness between models or model generations the same as the delta in the relevant benchmark results? Did you find interesting model differences and behaviors that don’t yet map onto a benchmark, but such a benchmark could be made? Did you encounter missing refusals to questions that should be refused in hindsight? All these can let you draw connections between evals and real-world uplift studies.
Finding these connections and correlations between uplift studies and in silico evals (“correlates of uplift”3) is extremely important as it provides you with an OOM cheaper estimator for the underlying model capability that you can tweak and run over and over again for every newly released model; and, crucially, also for pre-release testing and risk assessment.
Tier 4 — Actual AI-enabled biological incidents and near-misses (we hopefully never get here)
When I wrote earlier that uplift studies would be the most conclusive evidence, I wasn’t quite correct. Actual incidents or near-misses would really constitute conclusive evidence for the risk of AI-enabled bioterrorism.
I obviously hope that we will never see such deliberate (biological) mass harm enabled by AI, or even near-misses. But even if we do, it’s still not guaranteed to be conclusive: counterfactual attribution (would this have happened anyway even without AI?) will be very hard unless it’s an extremely clear-cut case. But my guess is that the most likely first offenders won’t be complete novices who have never seen a lab from the inside; they will be people who were almost capable of doing such a thing without AI already and merely got a little extra help. Proving that this was decisive and counterfactual sounds quite forensically challenging, and is hard for bioweapon attacks generally; the Amerithrax and Aum Shinrikyo investigations, for example, took years.
It doesn’t need spelling out, but we can’t just be reactive. We might get certainty ex-post, but that certainty could come with a catastrophically high price tag. We will need to get comfortable with higher ex-ante uncertainty than in other domains, e.g. drug approval, where the cost of a false negative is limited.
Not a proper Tier — Demos
I initially thought about positioning demos, such as individual LLM transcripts, as a tier between theoretical arguments and evals. And in some way, they can serve as an existence proof for a capability, or a convincing narrative of risk.
But ultimately, demos should be informed and backed by high-quality evidence, not replace it. So I view them more as a way to package, present, and communicate evidence from evals and real-world studies and make it more accessible and understandable.
Summary
In conclusion, I don’t claim that evals are pristine, unrefutable evidence for AI-enabled biorisk. They do, however, provide the best readily-available evidence, and merely waiting for higher-tier evidence has costs rarely priced-in by critics.
Theoretical arguments from first principles or analogies won’t convince the relevant stakeholders, but kickstarted initial research interest and funding.
Evals provide the best evidence-per-buck, even though their validity can be limited by various factors. They can work as rule-outs—with the caveat that rule-outs are only as good as the elicitation behind them.
Real-world uplift studies are most convincing and ecologically valid, but also slow and expensive.
Real-world studies should be correlated with solid evals which then function as cheap estimators of real-world uplift. A small number of well-designed uplift studies could plausibly anchor a much larger, much cheaper eval-based monitoring regime, a highly desirable outcome.
People who dismiss evals wholesale are implicitly demanding an evidence quality that the cost structure and release cadence of this field can’t sustain.
I want to thank Coleman Breen, Ben Mueller, Ryan Ritterson, Cassidy Nelson, and Claire Qureshi for their sensible comments and additions, and Seth Donoughe for the discussions that led to this post. All mistakes are mine.
Footnotes
The terminology can be a bit vague, but I like to think of “evals” (i.e. evaluations) as anything that measures a certain capability in silico. This can encompass both manual elicitation as well as “benchmarks”; sets of automatically run tasks yielding a quantitative result. I use evals and benchmarks somewhat interchangeably since the vast majority of evaluations are benchmarks. And importantly, I am talking about capability evaluations, not safeguard evaluations (that measure, e.g., the refusal behavior of a model).↩︎
Rather than the clinical analogy of evidence-based medicine introduced at the end, I like to think of it as an evidence landscape: The swampy forest of theoretical arguments (you think you’re making progress while dragging yourself through the mud in circles) giving way to the vast plains of evals with its rolling hills and shaded groves (inviting you to spend an eternity tweaking the wording of system prompts near a babbling stream). Towering above it all are the treacherous peaks of real-world uplift studies (difficult to reach but very rewarding once you’ve ascended them). Not sure how Tier 4 fits here, though; maybe the accursed obelisk on top of the highest summit, with swirling cumulonimbi overhead.↩︎
Analogous to the “correlates of protection”, easily measured immune parameters that are used to estimate clinical outcomes in vaccinology.↩︎