English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

ABC-Bench: An Agentic Bio-Capabilities Benchmark for Biosecurity

Forum topic · 小凯 · 2026-06-11

Summary

ABC-Bench (Agentic Bio-Capabilities Benchmark) is a new benchmark suite designed to measure biosecurity-relevant capabilities of large language model (LLM) agents. Introduced by Andrew Bo Liu, Samira Nedungadi, Bryce Cai, Alex Kleinman, Harmon Bhasin, and Seth Donoughe (arXiv:2606.11150), the benchmark evaluates LLM agents on both benign and dual-use biology tasks: writing code to operate liquid handling robots, designing DNA fragments for in vitro assembly, and attempting to evade DNA synthesis screening. According to the paper, all tested LLM agents surpassed the median human expert baseline across the three task categories, while performance was weaker on tasks requiring novel bioinformatics reasoning. In wet-lab validation, DNA assembly scripts generated by OpenAI's o4-mini-high successfully assembled DNA with the expected sequences on an OpenTrons liquid-handling robot. The results suggest that agentic AI capabilities in biology are advancing quickly enough to shift the biosecurity risk landscape, motivating benchmarks like ABC-Bench to track and evaluate these capabilities for safety purposes. The work is relevant to AI safety researchers, biosecurity policy experts, and ML practitioners monitoring frontier-model risks.

Overview

ABC-Bench (Agentic Bio-Capabilities Benchmark) is a suite of tasks for measuring biosecurity-relevant capabilities of LLM agents, presented in arXiv:2606.11150 by Andrew Bo Liu, Samira Nedungadi, Bryce Cai, Alex Kleinman, Harmon Bhasin, and Seth Donoughe.

Motivation

LLMs are rapidly acquiring capabilities relevant to biological research, from literature synthesis to interpretation of experimental data. Increasingly, LLM agents can also perform in silico biology tasks that previously required experienced human biologists. These emerging AI capabilities offer new opportunities for scientific discovery and biomedical advances, but they also shift the landscape of biosecurity risks.

What ABC-Bench Measures

ABC-Bench evaluates LLM agents on both benign and dual-use biology tasks, including:

  • Writing code to operate liquid handling robots
  • Designing DNA fragments for in vitro assembly
  • Evading DNA synthesis screening
  • Key Findings

  • All tested LLM agents exceeded the median human expert baseline across the three task categories.
  • Performance was weaker on tasks requiring novel bioinformatics reasoning.
  • In wet-lab validation, assembly scripts generated by OpenAI's o4-mini-high successfully assembled DNA with the expected sequences on an OpenTrons liquid-handling robot.
  • Links

  • Paper: arXiv:2606.11150

Tags

#abc-bench#biosecurity#llm-agents#ai-safety#benchmark#machine-learning#dual-use-research

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177981091