BioSecBench-Function: A Verifiable Benchmark for Reasoning about Biological Function from Experimental Data

Read the full article See related articles

Discuss this preprint

Start a discussion What are Sciety discussions?

Listed in

This article is not in any list yet, why not save it to one of your lists.
Log in to save this article

Abstract

Inferring biological function from experimental data is central to understanding emerging pathogens and developing effective countermeasures, yet interpreting these data remains slow and expert-intensive. AI agents could help accelerate this process by reasoning across sequence, structural, and biophysical evidence. We present BioSecBench-Function, a verifiable benchmark for recovering biosecurity-relevant function from real biological data. The benchmark comprises 111 evaluations built from published datasets and graded deterministically against ground truth. We organize evaluations along two dimensions: threat axis (spanning seven biosecurity-relevant question types) and biological question (indicating whether the solution depends primarily on sequence, structure, or biophysical assay data). Across 7,326 runs from twenty-two model-harness configurations, Opus 5 under Claude Code led on endpoint pass rate at 50.3%, and Grok 4.6 under Grok Build led on overall pass rate at 44.1% when refusals counted as failures. Performance varied substantially across both model-harness configurations and task categories. Refusal rates differed sharply by provider, and cost was a poor predictor of accuracy: several configurations exceeded 40% pass rate at low cost. BioSecBench-Function provides a standard for measuring whether agents can be trusted to interpret what a new pathogen or variant does when the next outbreak arrives.

Article activity feed