preprint · benchmark creator

BixBench: a Comprehensive Benchmark for LLM-based Agents in Computational Biology

FutureHouse · ScienceMachine · 2025-02-28

Relationship layer

Benchmark usage

This table records what the work did with each benchmark before attempting to normalize a run. Partial claims remain visible without being treated as comparable evaluations.

No BenchmarkUse relation is normalized for this legacy work yet. Existing EvaluationRuns remain available below.

Normalized evaluation runs

BixBench4 runs

Open benchmark record →

bixbench-v1-open-agentic-images-ten-runsvv1.0

Evaluated models / systems: Claude 3.5 Sonnet (BixBench version not reported), GPT-4o (BixBench version not reported)

Scopefull · n=296
Shotszero-shot task initialization
Turnsmulti-turn ReAct agent trajectory
System prompt publicYes
Reasoning / effortAviary ReActAgent with at most 25 steps
BrowserNo
InternetNot reported
DatabasesNot reported
Code executionYes
ContainerYes
External toolsedit cell, list workdir, and submit answer
Token budgetNot reported
Time / cost budgetmaximum 25 agent steps
Temperature1
SeedNot reported
Repeats10
Graderbinary LLM judgment against the ground-truth solution · model: Claude 3.5 Sonnet (exact version not reported) · human review: no
Statisticsaccuracy over all question-trajectory pairs; official plots use 95% Wilson intervals
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracyabsolutepercentquestion-trajectory-weighted mean over ten runs per questionbinary LLM judge

Results

ModelMetricValuen
GPT-4o (BixBench version not reported)Accuracy9 percent
Rounded open-answer headline in the paper text; ten trajectories per capsule, with plots/images allowed.
296
Claude 3.5 Sonnet (BixBench version not reported)Accuracy17 percent
Rounded open-answer headline in the paper text; ten trajectories per capsule, with plots/images allowed.
296

Evidence

  • page: PDF pp. 4–7, §§3.2.2–3.2.5 and §4.1 (Reports 296 questions, Docker environment, three tools, notebook execution, ten parallel analyses, plots/images condition, Claude judge, and aggregation.) — supports /benchmark_version, /scope, /protocol, /metrics
  • repository-path: bixbench/config.yaml at commit 6c28217959d5d7dd6f48c59894534fced7c6c040 (Records ReActAgent, temperature 1.0, maximum 25 steps, public prompt key, and total_questions 296.) — supports /protocol/reasoning, /protocol/system_prompt_public, /protocol/time_budget, /protocol/temperature
  • page: PDF pp. 1 and 6, abstract and §4.1; Figure 4 (Text explicitly reports Claude 3.5 Sonnet at 17% and GPT-4o at 9% in open-answer evaluation.) — supports /results
bixbench-v1-mcq-refusal-no-images-ten-runsvv1.0

Evaluated models / systems: Claude 3.5 Sonnet (BixBench version not reported), GPT-4o (BixBench version not reported)

Scopefull · n=296
Shotszero-shot task initialization
Turnsmulti-turn ReAct analysis followed by a separate MCQ call
System prompt publicYes
Reasoning / effortAviary ReActAgent with at most 25 steps and an instruction to avoid plots/images
BrowserNo
InternetNot reported
DatabasesNot reported
Code executionYes
ContainerYes
External toolsedit cell, list workdir, and submit answer
Token budgetNot reported
Time / cost budgetmaximum 25 agent steps
Temperature1
SeedNot reported
Repeats10
Gradersecond-LLM multiple-choice selection with Insufficient information refusal · model: Claude 3.5 Sonnet (exact version not reported) · human review: no
Statisticsmajority-vote accuracy and precision over ten trajectories
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracyabsolutepercentmajority vote over ten trajectories per questionexact selected option
Precisionabsolutepercentcorrect among questions not assigned refusalexact selected option

No numeric result rows are published yet; the verified protocol remains useful.

Evidence

  • figure: PDF pp. 7–8, Figure 5 caption and image-generation ablation discussion (States that the refusal option is present and agents are instructed not to produce images/plots.) — supports /benchmark_version, /scope, /protocol, /metrics
bixbench-v1-mcq-no-refusal-images-ten-runsvv1.0

Evaluated models / systems: Claude 3.5 Sonnet (BixBench version not reported), GPT-4o (BixBench version not reported)

Scopefull · n=296
Shotszero-shot task initialization
Turnsmulti-turn ReAct analysis followed by a separate MCQ call
System prompt publicYes
Reasoning / effortAviary ReActAgent with at most 25 steps
BrowserNo
InternetNot reported
DatabasesNot reported
Code executionYes
ContainerYes
External toolsedit cell, list workdir, and submit answer
Token budgetNot reported
Time / cost budgetmaximum 25 agent steps
Temperature1
SeedNot reported
Repeats10
Gradersecond-LLM forced multiple-choice selection · model: Claude 3.5 Sonnet (exact version not reported) · human review: no
Statisticsmajority vote over ten trajectories with accuracy
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracyabsolutepercentmajority vote over ten trajectories per questionexact selected option

No numeric result rows are published yet; the verified protocol remains useful.

Evidence

  • page: PDF pp. 6–7, Figure 4 and §4.1 (Reports the forced-answer ablation and ten-trajectory majority-vote accuracy.) — supports /benchmark_version, /scope, /protocol, /metrics
bixbench-v1-mcq-refusal-images-ten-runsvv1.0

Evaluated models / systems: Claude 3.5 Sonnet (BixBench version not reported), GPT-4o (BixBench version not reported)

Scopefull · n=296
Shotszero-shot task initialization
Turnsmulti-turn ReAct analysis followed by a separate MCQ call
System prompt publicYes
Reasoning / effortAviary ReActAgent with at most 25 steps
BrowserNo
InternetNot reported
DatabasesNot reported
Code executionYes
ContainerYes
External toolsedit cell, list workdir, and submit answer
Token budgetNot reported
Time / cost budgetmaximum 25 agent steps
Temperature1
SeedNot reported
Repeats10
Gradersecond-LLM multiple-choice selection with Insufficient information refusal · model: Claude 3.5 Sonnet (exact version not reported) · human review: no
Statisticsmajority vote over ten trajectories; accuracy and precision among non-refusal answers
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracyabsolutepercentmajority vote over ten trajectories per questionexact selected option
Precisionabsolutepercentcorrect among questions not assigned the refusal optionexact selected option

No numeric result rows are published yet; the verified protocol remains useful.

Evidence

  • page: PDF pp. 5–7, §3.2.4, Figure 4, and §4.1 (Defines the second-LLM MCQ conversion, refusal option, ten-run majority vote, accuracy, and precision.) — supports /benchmark_version, /scope, /protocol, /metrics