GPT-6 Astra Pre-Release Testing Report
Our team performed pre-release testing of GPT-6 Astra, OpenAI’s latest flagship model (publicly released on September 3, 2026), to evaluate its biological misuse-relevant capabilities. We tested versions with and without OpenAI’s system-level biological guardrails, enabling us to more deeply probe its underlying capabilities and evaluate its safeguards.
We received pre-release access to the final launch snapshots on Tuesday, August 25, and were provided six calendar days, including four business days, to assess these checkpoints and deliver our report. Within this access window, we focused on a prioritized set of evaluations that we agreed upon with OpenAI and excluded some evaluations due to saturation and time constraints. This assessment was therefore narrower in scope than our previous pre-release assessments.
We found GPT-6 Astra to be highly capable across all of the evaluations tested. Among our knowledge benchmarks, GPT-6 Astra exceeded every previously tested model on both our original Virology Capabilities Test (VCT) and VCT-v2, an updated version designed to be more shortcut-resistant (and used in pre-release testing for the first time here). The model checkpoint also performed within range of GPT-5.6 Sol on the Human Pathogen Capabilities Test (HPCT) and scored slightly lower than earlier models on World-Class Bio (WCB), though this decrease was driven largely by increased hedging rather than a knowledge gap. On the agentic benchmarks we tested, Astra modestly outperformed GPT-5.6 Sol on ReproBAIT and was the first model we’ve examined that was able to generate de novo protein designs meeting in silico designability thresholds. The railfree variant of GPT-6 Astra also correctly outlined a complex strategy to evade DNA synthesis screening that prior models had not successfully identified for our Advanced Screening Evasion (ASE) task. In terms of refusal behavior, the production version of GPT-6 Astra refused or blocked 82.9% of hazardous BioTIER queries (vs. 66% for GPT-5.6 Sol) while still complying with 95.3% of benign ones, and refused the screening evasion task 100% of the time.
Our full pre-deployment assessment report is available here. This post provides a high-level summary of the report.
Our approach to pre-release testing
As an independent evaluator, SecureBio conducts pre-release testing to help AI model developers understand and mitigate the biological misuse risks posed by AI. We measure capabilities with a battery of model benchmarks and manual, open-ended explorations of model behaviors by virology experts.
We aim to follow the AI Evaluator Forum (AEF) guidelines and the AEF-1 standard for transparent independent assessments, as well as SecureBio’s own principles and practices for model assessments. While OpenAI provided suggestions for eliciting maximum capabilities from their models, SecureBio retained autonomy over our evaluation methodology. The full report contains detailed descriptions of our methods and a review of this assessment against the AEF-1 standard.
Measuring capabilities
Our capabilities assessment included four approaches, each helping us understand a different facet of biological risk. We measure model capabilities using a suite of diverse evaluations.
Knowledge benchmarks: We first measured biosecurity-relevant knowledge and reasoning with three static “knowledge” benchmarks (VCT and its updated version VCT-v2, HPCT, and WCB), which test a model’s ability to provide correct answers to difficult questions drawn from biology research and wet lab work. Model scores are contextualized using scores from PhD-level experts.
Agentic benchmarks: We measured the ability to design and execute complex, multi-step workflows with two agentic evaluations: ReproBAIT, which tests how well an agent can independently reproduce published biological AI models from its own knowledge-base, and Advanced Screening Evasion, which tests an agent’s ability to evade DNA synthesis screening systems.
Refusal benchmark: We tested the refusal behavior of the safeguarded version of GPT-6 Astra on BioTIER, which tests a model’s refusal of hazardous biology queries and compliance with benign biology queries.
Manual assessment: Our team’s virology experts performed manual assessments to test the dual-use knowledge of the railfree Astra pre-release checkpoint and the willingness of the safeguarded Astra pre-release checkpoint to answer high-risk questions.
Knowledge benchmarks
GPT-6 Astra outperformed every model we have previously tested on the Virology Capabilities Test (VCT) and VCT-v2 knowledge benchmarks, with performance increases of 4.3 percentage points and 8 percentage points, respectively, relative to GPT-5.6 Sol, and more than 30 percentage points better than our human subject matter expert baselines.

Model performance on VCT (full set, multimodal). Accuracy scores (n ≥ 10 epochs) are plotted against model release date. The Astra Pre-Release Checkpoint (without system-level bio-classifiers) and the Astra Pre-Release Checkpoint (Railfree) are highlighted in the legend.
GPT-6 Astra’s score on the Human Pathogen Capabilities Test (HPCT) was very similar to the score achieved by GPT-5.6 Sol. On World-Class Bio (WCB), an open-response test of niche biological knowledge held by only a handful of specialists in the world, the highest-performing Astra checkpoint scored 61.6% ± 0.6%, about seven percentage points lower than GPT-5.6 Sol. Follow-up analyses of this unexpected result suggested that these lower scores were attributable to Astra’s greater tendency to hedge its answers (i.e., more questions marked as incorrect due to a failure to commit to the correct answer).
Agentic capabilities in biological workflows
GPT-6 Astra performed extremely well on bio-relevant agentic workflows, achieving the highest scores of any model we have evaluated on both agentic benchmarks that we tested.
On ReproBAIT, which asks an agent to independently reproduce a published biological AI tool from its scientific paper, GPT-6 Astra strongly outperformed GPT-5.6 Sol, reaching 80-83% of published performance, compared to 63-69% for GPT-5.6 Sol.
On Advanced Screening Evasion, which assesses an agent’s ability to evade commercial nucleic acid synthesis screening algorithms, the railfree variant of Astra was the top-scoring model that SecureBio has tested. It also was the first to achieve full credit for the “quality of evasion strategy” portion of the task by outlining a complex, known evasion strategy that prior models, including railfree GPT-5.6 Sol, had not successfully identified. Similarly to GPT-5.6 Sol, the Astra checkpoint with production safeguards refused the task 100% of the time.

Frontier model performance on Advanced Screening Evasion. Hollow markers show scores with refusals included as zeros, and filled markers show refusal-corrected scores. The Astra Pre-Release Checkpoint (Railfree) achieved the highest observed score among the evaluated leading closed-weight and open-weight models.
Refusal behavior
We assessed the biology safeguards of GPT-6 Astra using BioTIER. The BioTIER evaluation includes a set of benign prompts that we recommend answering (“BioTIER-permit”) and a set of dual-use and high-risk prompts that we recommend refusing (“BioTIER-refuse”); models should not provide substantive answers to high-risk prompts, but should answer benign ones. The current iteration of BioTIER does not evaluate the correctness or actionability of responses but focuses on query refusal. The safeguarded Astra pre-release checkpoint correctly refused 82.9% of BioTIER-refuse and correctly did not refuse 95.3% of BioTIER-permit. In comparison, GPT-5.6 Sol refused 66% of BioTIER-refuse prompts and did not refuse 99% of benign prompts. The refusal rates in the two BioTIER-refuse sets, Catastrophe Avoidance (CA) and Biomedical Dual-Use Research of Concern (BD), were similar but varied across themes (39%-100%).

Safeguarded Astra pre-release checkpoint refusal on BioTIER-refuse and compliance on BioTIER-permit by A. query eval component and B. query set. Mean refusal rate (%) over 10 epochs. Higher % = better. Note: 11 timeouts have been removed from analysis.
Manual assessment by virology experts
Our manual assessment of GPT-6 Astra was conducted by SecureBio expert virologists. Over 10 hours of interaction, our manual assessors tested the bounds of dual-use knowledge of the railfree Astra pre-release checkpoint and the willingness of the safeguarded Astra pre-release checkpoint to answer high-risk questions.
Consistent with previous model generations, the railfree Astra pre-release checkpoint provided comprehensive responses that highlighted key considerations and accurately referenced relevant literature, although it sometimes became overly verbose, to the point of requiring redirection. It was also able to make informative, biorisk-relevant experimental predictions in the absence of web search and tool-use.
The safeguarded GPT-6 Astra version was generally willing to help with meta-prompting and high-level conceptual discussions but refused to provide dangerous operational information along most of the directions probed; it also clearly indicated that specific information was out of bounds. GPT-6 Astra was also less likely than GPT-5.6 Sol to suggest specific payloads or adaptations that could be used to cause harm.
Summary
Pre-release testing is a necessary step to helping prevent biological misuse with frontier AI models. We thank the OpenAI Preparedness Team for the opportunity to conduct pre-release testing. We look forward to further engagement with frontier model firms as increases in model capabilities continue to highlight the growing need for independent third-party assessment.
Our full technical report, with all benchmark scores and full description of our methodology, is available on SecureBio’s website.