agentic-eval · audited-with-caveats · verified 2026-07-21

BixBench

A containerized benchmark of long-horizon bioinformatics analysis over real published notebooks and associated data, with open-answer and multiple-choice evaluation modes.

+3 more
Audited with caveats: 2 field(s) are marked provisional or conflicted. Warnings are shown next to affected values and these claims are excluded from unqualified comparisons.
Version and unit audit: the original v1.0 paper evaluated 296 questions from 53 capsules; current v1.5 has 205 questions. Upstream states 60 source notebooks/capsules, while the released rows reference 59 capsule UUIDs and the tagged tree stores 64 archives, so the 60 claim is shown as Conflicted. Exact zero-shot JSON values are registered; unlabeled agentic bars are not digitized.

Benchmark definition

What is counted

Version
v1.5
Total
205 (one question per row in the official v1.5 BixBench.jsonl)
Task formats
open-ended bioinformatics analysis; multiple choice
Capabilities
Data analysisCodingTool useScientific reasoning
Modalities
TextTableFigureDNA or RNA sequenceRaw omicsImageCode

Version history

VersionStatusRelease / as-ofTotalFormal tracks
v1.0
bixbench-v1-0
superseded2025-02-28296 (open-answer questions reported in the original creator preprint)None registered
v1.5
bixbench-v1-5
current2025-09-23205 (one question per row in the official v1.5 BixBench.jsonl)None registered

Tracks and subsets

IDCountBasisPartition?Notes
Referenced capsules in v1.5 rows
bixbench-v1-5-referenced-capsules
59unique capsule_uuid values referenced by the 205 rowsNoA source-artifact count, not a partition of questions.
Source notebooks stated in README
bixbench-v1-5-readme-notebooks
60Conflicted · highreal-world published Jupyter notebooks and related capsules stated in the official READMENoConflicts with 59 referenced capsule UUIDs and 64 capsule archives; upstream does not document how the three units correspond.
Capsule archives in v1.5 release
bixbench-v1-5-release-archives
64CapsuleFolder-*.zip paths in the official v1.5 dataset tagNoA stored-artifact count, not a partition of questions.
LLM-verifier questions
bixbench-v1-5-llm-verifier
83rows whose eval_mode is llm_verifierExclusive & exhaustiveMutually exclusive with the two deterministic verifier modes.
String-verifier questions
bixbench-v1-5-string-verifier
61rows whose eval_mode is str_verifierExclusive & exhaustiveMutually exclusive with the other verifier modes.
Range-verifier questions
bixbench-v1-5-range-verifier
61rows whose eval_mode is range_verifierExclusive & exhaustiveMutually exclusive with the other verifier modes.

Scientific Task Atlas

Scientific task classification

partial for v1.5. Official categories are multi-label domains rather than an exhaustive scientific-task taxonomy.

Scientific taskCoverageCountMappingEvidence
End-to-end computational analysisexplicitly-in-scope205 questions
one question per row in the official v1.5 BixBench.jsonl
official-taxonomy
high confidence
bixbench-evidence-v1-5-counts
Questions are grounded in containerized published analysis capsules.
Omics and cellular analysisobservedNot reported
v1.5 questions carrying official genomics, transcriptomics, epigenomics, single-cell, proteomics, or integrative-omics category labels.
official-taxonomy
high confidence
bixbench-evidence-taxonomy
Official multi-label domain categories overlap, so no additive omics question total is asserted.

Scientific coverage notes

DomainCoverageCountInterpretation
Genomicsobserved74v1.5 rows carrying the multi-label Genomics category; category counts overlap.
Transcriptomicsobserved69v1.5 rows carrying the multi-label Transcriptomics category; category counts overlap.
Epigenomicsobserved12v1.5 rows carrying the multi-label Epigenomics category; category counts overlap.
Single-cellobserved2v1.5 rows carrying the multi-label Single-Cell category; category counts overlap.
Proteomicsobserved4v1.5 rows carrying the multi-label Proteomics category; category counts overlap.
Multi-omicsobserved2v1.5 rows carrying the multi-label Integrative Omics category; category counts overlap.
Protein-protein bindingunknownNot reportedNo standalone protein-protein-binding count is published or encoded as a v1.5 category.
Protein-ligand bindingunknownNot reportedNo standalone protein-ligand-binding count is published or encoded as a v1.5 category.

Relationship registry

How works use this benchmark

Partial claims, non-evaluation uses, and third-party summaries stay visible without entering model comparisons.

Partial evaluation claims

evaluation

anthropic-life-sciences-bixbench

Partialunknown

Work: Claude for Life Sciences · source version anthropic-life-sciences-2025-10-20

Selection
not reported
Metrics
Not reported / not applicable
Linked runs
None

Not reported / unresolved: benchmark version; full, subset, or track scope; realized n; metric and aggregation; numeric results; prompt, tools, and budget; repeats and grader

Anthropic states that Sonnet 4.5 shows a similar improvement over Sonnet 4 on BixBench, but publishes no score or sufficiently specified evaluation setting. No value is inferred from the wording.

Evidence
  • section: Making Claude a better research partner, paragraph 2 (Names BixBench and compares Sonnet 4.5 with predecessor Sonnet 4, without settings, metric, or results.)
    Supports: /benchmark_version, /relation_type, /status, /model_ids, /scope, /metric_labels, /evaluation_run_ids, /reporting_gaps, /notes

Evaluation registry

Works and run settings

A setting change—scope, prompt, tools, budget, grader, or repeats—creates a separate run. Charts never cross a comparability group.

bixbench-v1-open-agentic-images-ten-runsvv1.0

Evaluated models / systems: Claude 3.5 Sonnet (BixBench version not reported), GPT-4o (BixBench version not reported)

Scopefull · n=296
Shotszero-shot task initialization
Turnsmulti-turn ReAct agent trajectory
System prompt publicYes
Reasoning / effortAviary ReActAgent with at most 25 steps
BrowserNo
InternetNot reported
DatabasesNot reported
Code executionYes
ContainerYes
External toolsedit cell, list workdir, and submit answer
Token budgetNot reported
Time / cost budgetmaximum 25 agent steps
Temperature1
SeedNot reported
Repeats10
Graderbinary LLM judgment against the ground-truth solution · model: Claude 3.5 Sonnet (exact version not reported) · human review: no
Statisticsaccuracy over all question-trajectory pairs; official plots use 95% Wilson intervals
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracyabsolutepercentquestion-trajectory-weighted mean over ten runs per questionbinary LLM judge

Results

ModelMetricValuen
GPT-4o (BixBench version not reported)Accuracy9 percent
Rounded open-answer headline in the paper text; ten trajectories per capsule, with plots/images allowed.
296
Claude 3.5 Sonnet (BixBench version not reported)Accuracy17 percent
Rounded open-answer headline in the paper text; ten trajectories per capsule, with plots/images allowed.
296

Evidence

  • page: PDF pp. 4–7, §§3.2.2–3.2.5 and §4.1 (Reports 296 questions, Docker environment, three tools, notebook execution, ten parallel analyses, plots/images condition, Claude judge, and aggregation.) — supports /benchmark_version, /scope, /protocol, /metrics
  • repository-path: bixbench/config.yaml at commit 6c28217959d5d7dd6f48c59894534fced7c6c040 (Records ReActAgent, temperature 1.0, maximum 25 steps, public prompt key, and total_questions 296.) — supports /protocol/reasoning, /protocol/system_prompt_public, /protocol/time_budget, /protocol/temperature
  • page: PDF pp. 1 and 6, abstract and §4.1; Figure 4 (Text explicitly reports Claude 3.5 Sonnet at 17% and GPT-4o at 9% in open-answer evaluation.) — supports /results
bixbench-v1-mcq-refusal-no-images-ten-runsvv1.0

Evaluated models / systems: Claude 3.5 Sonnet (BixBench version not reported), GPT-4o (BixBench version not reported)

Scopefull · n=296
Shotszero-shot task initialization
Turnsmulti-turn ReAct analysis followed by a separate MCQ call
System prompt publicYes
Reasoning / effortAviary ReActAgent with at most 25 steps and an instruction to avoid plots/images
BrowserNo
InternetNot reported
DatabasesNot reported
Code executionYes
ContainerYes
External toolsedit cell, list workdir, and submit answer
Token budgetNot reported
Time / cost budgetmaximum 25 agent steps
Temperature1
SeedNot reported
Repeats10
Gradersecond-LLM multiple-choice selection with Insufficient information refusal · model: Claude 3.5 Sonnet (exact version not reported) · human review: no
Statisticsmajority-vote accuracy and precision over ten trajectories
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracyabsolutepercentmajority vote over ten trajectories per questionexact selected option
Precisionabsolutepercentcorrect among questions not assigned refusalexact selected option

No numeric result rows are published yet; the verified protocol remains useful.

Evidence

  • figure: PDF pp. 7–8, Figure 5 caption and image-generation ablation discussion (States that the refusal option is present and agents are instructed not to produce images/plots.) — supports /benchmark_version, /scope, /protocol, /metrics
bixbench-v1-mcq-no-refusal-images-ten-runsvv1.0

Evaluated models / systems: Claude 3.5 Sonnet (BixBench version not reported), GPT-4o (BixBench version not reported)

Scopefull · n=296
Shotszero-shot task initialization
Turnsmulti-turn ReAct analysis followed by a separate MCQ call
System prompt publicYes
Reasoning / effortAviary ReActAgent with at most 25 steps
BrowserNo
InternetNot reported
DatabasesNot reported
Code executionYes
ContainerYes
External toolsedit cell, list workdir, and submit answer
Token budgetNot reported
Time / cost budgetmaximum 25 agent steps
Temperature1
SeedNot reported
Repeats10
Gradersecond-LLM forced multiple-choice selection · model: Claude 3.5 Sonnet (exact version not reported) · human review: no
Statisticsmajority vote over ten trajectories with accuracy
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracyabsolutepercentmajority vote over ten trajectories per questionexact selected option

No numeric result rows are published yet; the verified protocol remains useful.

Evidence

  • page: PDF pp. 6–7, Figure 4 and §4.1 (Reports the forced-answer ablation and ten-trajectory majority-vote accuracy.) — supports /benchmark_version, /scope, /protocol, /metrics
bixbench-v1-mcq-refusal-images-ten-runsvv1.0

Evaluated models / systems: Claude 3.5 Sonnet (BixBench version not reported), GPT-4o (BixBench version not reported)

Scopefull · n=296
Shotszero-shot task initialization
Turnsmulti-turn ReAct analysis followed by a separate MCQ call
System prompt publicYes
Reasoning / effortAviary ReActAgent with at most 25 steps
BrowserNo
InternetNot reported
DatabasesNot reported
Code executionYes
ContainerYes
External toolsedit cell, list workdir, and submit answer
Token budgetNot reported
Time / cost budgetmaximum 25 agent steps
Temperature1
SeedNot reported
Repeats10
Gradersecond-LLM multiple-choice selection with Insufficient information refusal · model: Claude 3.5 Sonnet (exact version not reported) · human review: no
Statisticsmajority vote over ten trajectories; accuracy and precision among non-refusal answers
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracyabsolutepercentmajority vote over ten trajectories per questionexact selected option
Precisionabsolutepercentcorrect among questions not assigned the refusal optionexact selected option

No numeric result rows are published yet; the verified protocol remains useful.

Evidence

  • page: PDF pp. 5–7, §3.2.4, Figure 4, and §4.1 (Defines the second-LLM MCQ conversion, refusal option, ten-run majority vote, accuracy, and precision.) — supports /benchmark_version, /scope, /protocol, /metrics

Evaluation run

bixbench-v1-5-agentic-mcq-no-refusal-images

From BixBench v1.5 dataset and evaluation release

bixbench-v1-5-agentic-mcq-no-refusal-images-five-replicasvv1.5

Evaluated models / systems: Claude 3.5 Sonnet 20241022, GPT-4o (BixBench version not reported)

Scopefull · n=205
Shotszero-shot task initialization
Turnsmulti-turn Aviary SimpleAgent trajectory followed by forced-choice MCQ postprocessing
System prompt publicYes
Reasoning / effortSimpleAgent with at most 20 steps
BrowserNo
InternetNot reported
DatabasesNot reported
Code executionYes
ContainerYes
External toolsnotebook editing and capsule workspace inspection
Token budgetNot reported
Time / cost budgetmaximum 20 agent steps
Temperature1
SeedNot reported
Repeats5
Graderforced multiple-choice conversion without refusal · human review: no
Statisticsaccuracy with 95% Wilson interval and majority-vote analysis
Contaminationpublic benchmark canary
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracyabsolutepercentquestion-replica-weighted meanexact selected option

No numeric result rows are published yet; the verified protocol remains useful.

Evidence

  • repository-path: image model configs, v1.5_paper_results.yaml, README.md, and scripts/run_agentic.sh at commit 49311180bdacb324c596f2e07596c126f2004008 (Defines images allowed, forced-choice condition, full dataset, agent settings, metric, and majority-vote postprocessing.) — supports /benchmark_version, /scope, /protocol, /metrics

Evaluation run

bixbench-v1-5-agentic-mcq-refusal-images

From BixBench v1.5 dataset and evaluation release

bixbench-v1-5-agentic-mcq-refusal-images-five-replicasvv1.5

Evaluated models / systems: Claude 3.5 Sonnet 20241022, GPT-4o (BixBench version not reported)

Scopefull · n=205
Shotszero-shot task initialization
Turnsmulti-turn Aviary SimpleAgent trajectory followed by MCQ postprocessing
System prompt publicYes
Reasoning / effortSimpleAgent with at most 20 steps
BrowserNo
InternetNot reported
DatabasesNot reported
Code executionYes
ContainerYes
External toolsnotebook editing and capsule workspace inspection
Token budgetNot reported
Time / cost budgetmaximum 20 agent steps
Temperature1
SeedNot reported
Repeats5
Gradermultiple-choice conversion with Insufficient information refusal · human review: no
Statisticsaccuracy with 95% Wilson interval and majority-vote analysis
Contaminationpublic benchmark canary
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracyabsolutepercentquestion-replica-weighted meanexact selected option

No numeric result rows are published yet; the verified protocol remains useful.

Evidence

  • repository-path: image model configs, v1.5_paper_results.yaml, README.md, and scripts/run_agentic.sh at commit 49311180bdacb324c596f2e07596c126f2004008 (Defines images allowed, refusal condition, full dataset, agent settings, metric, and majority-vote postprocessing.) — supports /benchmark_version, /scope, /protocol, /metrics

Evaluation run

bixbench-v1-5-agentic-mcq-refusal-no-images

From BixBench v1.5 dataset and evaluation release

bixbench-v1-5-agentic-mcq-refusal-no-images-five-replicasvv1.5

Evaluated models / systems: Claude 3.5 Sonnet 20241022, GPT-4o (BixBench version not reported)

Scopefull · n=205
Shotszero-shot task initialization
Turnsmulti-turn Aviary SimpleAgent trajectory followed by MCQ postprocessing
System prompt publicYes
Reasoning / effortSimpleAgent with at most 20 steps and an instruction to avoid plots/images
BrowserNo
InternetNot reported
DatabasesNot reported
Code executionYes
ContainerYes
External toolsnotebook editing and capsule workspace inspection
Token budgetNot reported
Time / cost budgetmaximum 20 agent steps
Temperature1
SeedNot reported
Repeats5
Gradermultiple-choice conversion with Insufficient information refusal · human review: no
Statisticsmajority-vote accuracy by vote count
Contaminationpublic benchmark canary
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracyabsolutepercentmajority vote over stated five replicas per questionexact selected option

No numeric result rows are published yet; the verified protocol remains useful.

Evidence

  • repository-path: 4o_no_image.yaml, claude_no_image.yaml, v1.5_paper_results.yaml, README.md, and scripts/run_agentic.sh at commit 49311180bdacb324c596f2e07596c126f2004008 (Defines avoid_images true, refusal condition, full dataset, agent settings, and majority-vote postprocessing.) — supports /benchmark_version, /scope, /protocol, /metrics

Evaluation run

bixbench-v1-5-agentic-open-images

From BixBench v1.5 dataset and evaluation release

bixbench-v1-5-agentic-open-images-five-replicasvv1.5

Evaluated models / systems: Claude 3.5 Sonnet 20241022, GPT-4o (BixBench version not reported)

Scopefull · n=205
Shotszero-shot task initialization
Turnsmulti-turn Aviary SimpleAgent trajectory
System prompt publicYes
Reasoning / effortSimpleAgent with at most 20 steps
BrowserNo
InternetNot reported
DatabasesNot reported
Code executionYes
ContainerYes
External toolsnotebook editing and capsule workspace inspection
Token budgetNot reported
Time / cost budgetmaximum 20 agent steps
Temperature1
SeedNot reported
Repeats5
Graderrow-specific open-answer verifier · model: GPT-4o for llm_verifier rows (exact endpoint not reported) · human review: no
Statisticsaccuracy with 95% Wilson interval
Contaminationpublic benchmark canary
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracyabsolutepercentquestion-replica-weighted meanverifier-specific

No numeric result rows are published yet; the verified protocol remains useful.

Evidence

  • repository-path: README.md; bixbench/run_configuration/4o_image.yaml and claude_image.yaml; scripts/run_agentic.sh at commit 49311180bdacb324c596f2e07596c126f2004008 (Reports 205-question v1.5, Docker, SimpleAgent, 20 steps, temperature 1, images allowed, model strings, and stated five replicas.) — supports /benchmark_version, /scope, /protocol, /metrics
  • figure: bixbench_results_comparison.png, Open-answer group (Official accuracy bars with 95% Wilson intervals and no printed scalar labels.) — supports /metrics

Evaluation run

bixbench-v1-5-zero-shot-mcq-no-refusal

From BixBench v1.5 dataset and evaluation release

bixbench-v1-5-zero-shot-mcq-no-refusal-temperature-onevv1.5

Evaluated models / systems: Claude 3.5 Sonnet (BixBench version not reported), GPT-4o (BixBench version not reported)

Scopefull · n=205
Shotszero-shot
Turnssingle-turn model call
System prompt publicYes
Reasoning / effortdefault model behavior
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNo
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
Temperature1
SeedNot reported
Repeats1
Graderexact normalized selected-option match · human review: no
Statisticsaccuracy, precision, and coverage over 205 forced-choice questions
Contaminationpublic benchmark canary
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
accuracyabsolutepercentcorrect over all questionsexact normalized option
precisionabsolutepercentcorrect over sure answersexact normalized option
coverageabsolutepercentsure answers over all questionsNot reported

Results

ModelMetricValuen
GPT-4o (BixBench version not reported)accuracy36.0975609756 percent
74 correct of 205.
205
GPT-4o (BixBench version not reported)precision36.0975609756 percent
74 correct among 205 sure answers.
205
GPT-4o (BixBench version not reported)coverage100 percent
205 sure answers of 205.
205
Claude 3.5 Sonnet (BixBench version not reported)accuracy34.1463414634 percent
70 correct of 205.
205
Claude 3.5 Sonnet (BixBench version not reported)precision34.1463414634 percent
70 correct among 205 sure answers.
205
Claude 3.5 Sonnet (BixBench version not reported)coverage100 percent
205 sure answers of 205.
205

Evidence

  • repository-path: scripts/run_zeroshot.sh, generate_zeroshot_evals.py, bixbench/zero_shot.py, and grade_outputs.py at commit 49311180bdacb324c596f2e07596c126f2004008 (Defines full-set zero-shot forced-choice calls, randomized options, temperature 1.0, and exact option grading.) — supports /benchmark_version, /scope, /protocol, /metrics
  • repository-path: bixbench-v1.5_results/zero_shot_baselines.json at commit 28909d842bc492ecd99bab303279afb29e3cb353 (Exact accuracy, precision, coverage, n_total, n_correct, and n_sure for both forced-choice baselines.) — supports /results

Evaluation run

bixbench-v1-5-zero-shot-mcq-refusal

From BixBench v1.5 dataset and evaluation release

bixbench-v1-5-zero-shot-mcq-refusal-temperature-onevv1.5

Evaluated models / systems: Claude 3.5 Sonnet (BixBench version not reported), GPT-4o (BixBench version not reported)

Scopefull · n=205
Shotszero-shot
Turnssingle-turn model call
System prompt publicYes
Reasoning / effortdefault model behavior
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNo
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
Temperature1
SeedNot reported
Repeats1
Graderexact normalized selected-option match with a sure/refusal flag · human review: no
Statisticsaccuracy, precision among non-refusals, and non-refusal coverage over 205 questions
Contaminationpublic benchmark canary
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
accuracyabsolutepercentcorrect over all questionsexact normalized option
precisionabsolutepercentcorrect over non-refusal answersexact normalized option
coverageabsolutepercentnon-refusal answers over all questionsNot reported

Results

ModelMetricValuen
GPT-4o (BixBench version not reported)accuracy3.9024390244 percent
8 correct of 205.
205
GPT-4o (BixBench version not reported)precision40 percent
8 correct among 20 non-refusals.
205
GPT-4o (BixBench version not reported)coverage9.756097561 percent
20 non-refusals of 205.
205
Claude 3.5 Sonnet (BixBench version not reported)accuracy8.2926829268 percent
17 correct of 205.
205
Claude 3.5 Sonnet (BixBench version not reported)precision38.6363636364 percent
17 correct among 44 non-refusals.
205
Claude 3.5 Sonnet (BixBench version not reported)coverage21.4634146341 percent
44 non-refusals of 205.
205

Evidence

  • repository-path: scripts/run_zeroshot.sh, generate_zeroshot_evals.py, bixbench/zero_shot.py, and grade_outputs.py at commit 49311180bdacb324c596f2e07596c126f2004008 (Defines full-set zero-shot MCQ calls, randomized options, refusal flag, temperature 1.0, and exact option grading.) — supports /benchmark_version, /scope, /protocol, /metrics
  • repository-path: bixbench-v1.5_results/zero_shot_baselines.json at commit 28909d842bc492ecd99bab303279afb29e3cb353 (Exact accuracy, precision, coverage, n_total, n_correct, and n_sure for both refusal-condition baselines.) — supports /results

Evaluation run

bixbench-v1-5-zero-shot-open

From BixBench v1.5 dataset and evaluation release

bixbench-v1-5-zero-shot-open-temperature-onevv1.5

Evaluated models / systems: Claude 3.5 Sonnet (BixBench version not reported), GPT-4o (BixBench version not reported)

Scopefull · n=205
Shotszero-shot
Turnssingle-turn model call
System prompt publicYes
Reasoning / effortdefault model behavior
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNo
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
Temperature1
SeedNot reported
Repeats1
Graderrow-specific string, numeric-range, or LLM verifier · model: GPT-4o for llm_verifier rows (exact endpoint not reported) · human review: no
Statisticsaccuracy over 205 questions, plus precision and coverage from correct/sure flags
Contaminationpublic benchmark canary
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
accuracyabsolutepercentcorrect over all 205 questionsverifier-specific
precisionabsolutepercentcorrect over sure predictionsverifier-specific
coverageabsolutepercentsure predictions over all 205 questionsNot reported

Results

ModelMetricValuen
GPT-4o (BixBench version not reported)accuracy2.9268292683 percent
6 correct of 205.
205
GPT-4o (BixBench version not reported)precision2.9268292683 percent
6 correct among 205 sure predictions.
205
GPT-4o (BixBench version not reported)coverage100 percent
205 sure predictions of 205.
205
Claude 3.5 Sonnet (BixBench version not reported)accuracy2.9268292683 percent
6 correct of 205.
205
Claude 3.5 Sonnet (BixBench version not reported)precision2.9268292683 percent
6 correct among 205 sure predictions.
205
Claude 3.5 Sonnet (BixBench version not reported)coverage100 percent
205 sure predictions of 205.
205

Evidence

  • repository-path: generate_zeroshot_evals.py, grade_outputs.py, bixbench/zero_shot.py, bixbench/graders.py, and scripts/run_zeroshot.sh at commit 49311180bdacb324c596f2e07596c126f2004008 (Defines full-dataset single calls, temperature 1.0, public prompts, no tools/files, and mixed open-answer verifiers.) — supports /benchmark_version, /scope, /protocol, /metrics
  • repository-path: bixbench-v1.5_results/zero_shot_baselines.json at commit 28909d842bc492ecd99bab303279afb29e3cb353 (Both models have n_total 205, n_correct 6, n_sure 205.) — supports /results

Comparable result views

Accuracy

bixbench-creator-paper · bixbench-v1-open-agentic-images-ten-runs

CSV ↓
Accessible data table
ModelValueComparability group
GPT-4o (BixBench version not reported)9bixbench-v1-open-agentic-images-ten-runs
Claude 3.5 Sonnet (BixBench version not reported)17bixbench-v1-open-agentic-images-ten-runs

accuracy

bixbench-v1-5-zero-shot-mcq-no-refusal · bixbench-v1-5-zero-shot-mcq-no-refusal-temperature-one

CSV ↓
Accessible data table
ModelValueComparability group
GPT-4o (BixBench version not reported)36.0975609756bixbench-v1-5-zero-shot-mcq-no-refusal-temperature-one
Claude 3.5 Sonnet (BixBench version not reported)34.1463414634bixbench-v1-5-zero-shot-mcq-no-refusal-temperature-one

precision

bixbench-v1-5-zero-shot-mcq-no-refusal · bixbench-v1-5-zero-shot-mcq-no-refusal-temperature-one

CSV ↓
Accessible data table
ModelValueComparability group
GPT-4o (BixBench version not reported)36.0975609756bixbench-v1-5-zero-shot-mcq-no-refusal-temperature-one
Claude 3.5 Sonnet (BixBench version not reported)34.1463414634bixbench-v1-5-zero-shot-mcq-no-refusal-temperature-one

coverage

bixbench-v1-5-zero-shot-mcq-no-refusal · bixbench-v1-5-zero-shot-mcq-no-refusal-temperature-one

CSV ↓
Accessible data table
ModelValueComparability group
GPT-4o (BixBench version not reported)100bixbench-v1-5-zero-shot-mcq-no-refusal-temperature-one
Claude 3.5 Sonnet (BixBench version not reported)100bixbench-v1-5-zero-shot-mcq-no-refusal-temperature-one

accuracy

bixbench-v1-5-zero-shot-mcq-refusal · bixbench-v1-5-zero-shot-mcq-refusal-temperature-one

CSV ↓
Accessible data table
ModelValueComparability group
GPT-4o (BixBench version not reported)3.9024390244bixbench-v1-5-zero-shot-mcq-refusal-temperature-one
Claude 3.5 Sonnet (BixBench version not reported)8.2926829268bixbench-v1-5-zero-shot-mcq-refusal-temperature-one

precision

bixbench-v1-5-zero-shot-mcq-refusal · bixbench-v1-5-zero-shot-mcq-refusal-temperature-one

CSV ↓
Accessible data table
ModelValueComparability group
GPT-4o (BixBench version not reported)40bixbench-v1-5-zero-shot-mcq-refusal-temperature-one
Claude 3.5 Sonnet (BixBench version not reported)38.6363636364bixbench-v1-5-zero-shot-mcq-refusal-temperature-one

coverage

bixbench-v1-5-zero-shot-mcq-refusal · bixbench-v1-5-zero-shot-mcq-refusal-temperature-one

CSV ↓
Accessible data table
ModelValueComparability group
GPT-4o (BixBench version not reported)9.756097561bixbench-v1-5-zero-shot-mcq-refusal-temperature-one
Claude 3.5 Sonnet (BixBench version not reported)21.4634146341bixbench-v1-5-zero-shot-mcq-refusal-temperature-one

accuracy

bixbench-v1-5-zero-shot-open · bixbench-v1-5-zero-shot-open-temperature-one

CSV ↓
Accessible data table
ModelValueComparability group
GPT-4o (BixBench version not reported)2.9268292683bixbench-v1-5-zero-shot-open-temperature-one
Claude 3.5 Sonnet (BixBench version not reported)2.9268292683bixbench-v1-5-zero-shot-open-temperature-one

precision

bixbench-v1-5-zero-shot-open · bixbench-v1-5-zero-shot-open-temperature-one

CSV ↓
Accessible data table
ModelValueComparability group
GPT-4o (BixBench version not reported)2.9268292683bixbench-v1-5-zero-shot-open-temperature-one
Claude 3.5 Sonnet (BixBench version not reported)2.9268292683bixbench-v1-5-zero-shot-open-temperature-one

coverage

bixbench-v1-5-zero-shot-open · bixbench-v1-5-zero-shot-open-temperature-one

CSV ↓
Accessible data table
ModelValueComparability group
GPT-4o (BixBench version not reported)100bixbench-v1-5-zero-shot-open-temperature-one
Claude 3.5 Sonnet (BixBench version not reported)100bixbench-v1-5-zero-shot-open-temperature-one

Evidence and change history

Source locators remain visible; expand an item to inspect the exact Registry fields it supports.

BixBench: a Comprehensive Benchmark for LLM-based Agents in Computational Biology · page: PDF pp. 1–2, title, abstract, and contributions (Names the benchmark, creators, initial publication, agentic data-analysis objective, and open/MCQ modes.) · Supports 7 fields

Open source →

  • /name
  • /organizations
  • /release_date
  • /kind
  • /summary
  • /capabilities
  • /task_formats
BixBench: a Comprehensive Benchmark for LLM-based Agents in Computational Biology · page: PDF pp. 2, 4, and 6, Contributions, §3.2.1, and §4.1 (Reports 53 analytical scenarios/capsules and 296 questions in the original release.) · Supports 3 fields

Open source →

  • /versions/0/task_counts/total
  • /versions/0/task_counts/basis
  • /versions/0/task_counts/subsets
bixbench-v1-5-dataset-resource · repository-path: BixBench.jsonl at commit 2f660cbcab36b87f5d49b997f71e67be94131889 (205 unique rows; verifier modes are 83 llm_verifier, 61 str_verifier, and 61 range_verifier.) · Supports 8 fields

Open source →

  • /latest_version
  • /task_counts/total
  • /task_counts/basis
  • /task_counts/subsets
  • /versions/1/task_counts/total
  • /versions/1/task_counts/basis
  • /versions/1/task_counts/subsets
  • /scientific_task_classification/entries/0
bixbench-v1-5-dataset-resource · repository-path: README.md, BixBench.jsonl, and v1.5 tree at commit 2f660cbcab36b87f5d49b997f71e67be94131889 (README claims 60 source notebooks; JSONL references 59 capsule UUIDs; the tagged tree stores 64 CapsuleFolder ZIP files.) · Supports 4 fields

Open source →

  • /task_counts/subsets
  • /task_counts/subsets/1/count
  • /versions/1/task_counts/subsets
  • /versions/1/task_counts/subsets/1/count
bixbench-v1-5-dataset-resource · repository-path: BixBench.jsonl category, question, and artifact fields (Multi-label category counts include Genomics 74, Transcriptomics 69, Epigenomics 12, Proteomics 4, Single-Cell 2, and Integrative Omics 2; rows also reference tabular, sequence, imaging, raw-data, and notebook analyses.) · Supports 4 fields

Open source →

  • /domains
  • /modalities
  • /coverage_notes
  • /scientific_task_classification/entries/1
bixbench-repository-resource · repository-path: README.md, LICENSE, bixbench/graders.py, and scripts at commit 49311180bdacb324c596f2e07596c126f2004008 (Apache-2.0 license and public Docker harness, prompts, graders, and download/run instructions.) · Supports 7 fields

Open source →

  • /access/level
  • /access/license
  • /access/tasks
  • /access/artifacts
  • /access/grader
  • /resources
  • /implementations

Unresolved field claims

  • /task_counts/subsets/1/countConflicted · high — The README states 60 published notebooks/capsules, while v1.5 rows reference 59 unique capsule UUIDs and the release tree contains 64 capsule ZIP files; the relation among these units is undocumented.
    Evidence: bixbench-evidence-capsule-count-conflict
  • /versions/1/task_counts/subsets/1/countConflicted · high — The README states 60 published notebooks/capsules, while v1.5 rows reference 59 unique capsule UUIDs and the release tree contains 64 capsule ZIP files; the relation among these units is undocumented.
    Evidence: bixbench-evidence-capsule-count-conflict

View source-level modification history on GitHub →