Kimi K3 Biology Capabilities Assessment
Kimi K3, the latest model from Moonshot AI, is the best-performing open-weights model on our knowledge benchmarks. K3 achieves comparable performance across our evaluations as closed-weights models did approximately 8.1 months ago12. It is also among the most likely to answer biohazardous prompts, based on our BioTIER safeguard evaluation3.
Kimi K3 was released via API on July 16 with weights uploaded for public use on July 27. This post covers our knowledge benchmarks and safeguard evaluations. Scores on our agentic benchmarks will be released shortly on our dashboard.
Kimi K3 is comparable to closed-weights models approximately eight months ago
On SecureBio’s Bio Capabilities Index (BCI) score, which aggregates multiple biology benchmarks into a single comparable number, Kimi scores a 142, placing it seventh among the models we have evaluated to date. K3 sits where the closed frontier stood roughly 8.1 months earlier (see Figure 1).

Performance on knowledge benchmarks
Our “knowledge” benchmarks cover a wide range of biology and biosecurity-relevant topics. Kimi K3 scores highly on all of these benchmarks, surpassing other open-weights models, but trailing the frontier closed-weights models (see Figure 2). We highlight a few findings below:
- Virology Capabilities Test (VCT): 46.7 ± 5.4%. 322 multimodal questions on practical, hard-to-Google virology wet-lab work, deliberately targeting methods with dual-use potential. Kimi K3 sits tied with Claude Opus 4.6 (46.7%) and just behind Gemini 3.5 Flash (47.1%).
- Molecular Biology Capabilities Test (MBCT): 58.4 ± 6.8%. 200 text-only questions on practical, hard-to-Google molecular biology methods essential to any lab. Among models scored on more than 50% of items, Kimi K3 ranks second of 62.4
- World Class Biology (WCB): 55.0 ± 10.0%. 96 open-ended, rubric-graded questions on rare biology knowledge held by only a handful of world-class experts. Kimi K3 sits just below Claude Opus 4.6 (56.5%) and GPT-5.2 Pro (56.3%).
- Human Pathogen Capabilities Test (HPCT): 58.5 ± 9.7%. 100 text-only questions on tacit practical knowledge behind four pathogen groups of especially high-concern. This is Kimi K3’s weakest relative result, trailing a number of models and landing around GPT-5 (59.2%). GLM-5.2 scored higher at 65.3%.

Compared to the open-weights models we have evaluated5, Kimi K3 takes the top open-weights score on MBCT, WCB and VCT. The next-highest open-weights results are GLM-5.2 at 55.5% on MBCT, Kimi K2.6 at 54.9% on WCB, and Kimi K2.6 at 43.6% on VCT: margins of 2.9, 0.1 and 3.1 points. HPCT is the exception, where GLM-5.2 leads at 65.3%.
Kimi K3’s biology safeguards trail behind closed-weights models
BioTIER measures whether models refuse hazardous biological queries while permitting benign ones. Kimi K3 refuses on only 26.9% of the questions in our BioTIER-refuse set. Top-performing closed-weights models score over 90% on BioTIER-refuse. Kimi K3 correctly permits 99.3% of questions on our BioTIER-permit set.

What this means for biosecurity
Kimi K3 puts the capabilities of the closed frontier eight months ago (circa December 2025) into weights that can be downloaded and fine-tuned, with safeguards that refuse just 26.9% of hazardous prompts, underscoring the need for better preparation. We will continue to track the estimated lag in capabilities between the closed and open-weights frontiers in future assessments.
Footnotes
The average BCI of Kimi K3 is 142 with a 95% confidence interval of [138, 150]. The 8.1-month gap between Kimi K3’s BCI and the closed frontier is a rough estimate calculated by mapping the average BCI to the last date at which the closed weights frontier was comparable ( i.e. when then the best BCI score was about 142). Using a simple linear interpolation, the “frontier BCI” was approximately 138 8.3 months ago; and approximately 150 7.1 months ago. When using a similar procedure with individual benchmarks, the duration gap between Kimi K3’s BCI and the closed frontier is much wider, ranging from 2 to 15 months.↩︎
For comparison, Epoch AI’s Capability Index (ECI ) places Kimi K3 roughly 4-6 months behind the closed frontier on general capabilities.↩︎
The BioTIER manuscript describes the full methodology.↩︎
The only model above it, Claude Opus 4, refused 48% of the benchmark and was scored on the remaining half.↩︎
This includes Kimi K2.5 and K2.6, GLM-5 and GLM-5.2, Qwen3.5-397B and Qwen3.7-Max, DeepSeek V4 Pro, and gpt-oss-120b.↩︎