official-release · benchmark creator

Paving the way for AI agents in biology

Anthropic · 2025-05-20

Relationship layer

Benchmark usage

This table records what the work did with each benchmark before attempting to normalize a run. Partial claims remain visible without being treated as comparable evaluations.

No BenchmarkUse relation is normalized for this legacy work yet. Existing EvaluationRuns remain available below.

Normalized evaluation runs

VirBench1 run

Open benchmark record →

Evaluation run

virbench-official-run

From Paving the way for AI agents in biology

virbench-v1-full-ncbi-virusv1.0
Scopefull · n=120
ShotsNot reported
Turnsmulti-turn
System prompt publicNot reported
Reasoning / effortNot reported
BrowserYes
InternetYes
DatabasesNCBI Virus
Code executionNot reported
ContainerNot reported
External toolsscientific-agent retrieval tools
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
RepeatsNot reported
Gradermanual verified-answer comparison · human review: yes
Statisticsmean accuracy
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Mean accuracyabsolutepercentproblem-weighted meanNot reported

No numeric result rows are published yet; the verified protocol remains useful.

Evidence

  • VirBench methods and results — supports scope, protocol.tools, protocol.grader, metrics