official-release · benchmark creator

BixBench v1.5 dataset and evaluation release

FutureHouse · ScienceMachine · 2025-09-26

Relationship layer

Benchmark usage

This table records what the work did with each benchmark before attempting to normalize a run. Partial claims remain visible without being treated as comparable evaluations.

No BenchmarkUse relation is normalized for this legacy work yet. Existing EvaluationRuns remain available below.

Normalized evaluation runs

BixBench7 runs

Open benchmark record →

Evaluation run

bixbench-v1-5-agentic-mcq-no-refusal-images

From BixBench v1.5 dataset and evaluation release

bixbench-v1-5-agentic-mcq-no-refusal-images-five-replicasvv1.5

Evaluated models / systems: Claude 3.5 Sonnet 20241022, GPT-4o (BixBench version not reported)

Scopefull · n=205
Shotszero-shot task initialization
Turnsmulti-turn Aviary SimpleAgent trajectory followed by forced-choice MCQ postprocessing
System prompt publicYes
Reasoning / effortSimpleAgent with at most 20 steps
BrowserNo
InternetNot reported
DatabasesNot reported
Code executionYes
ContainerYes
External toolsnotebook editing and capsule workspace inspection
Token budgetNot reported
Time / cost budgetmaximum 20 agent steps
Temperature1
SeedNot reported
Repeats5
Graderforced multiple-choice conversion without refusal · human review: no
Statisticsaccuracy with 95% Wilson interval and majority-vote analysis
Contaminationpublic benchmark canary
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracyabsolutepercentquestion-replica-weighted meanexact selected option

No numeric result rows are published yet; the verified protocol remains useful.

Evidence

  • repository-path: image model configs, v1.5_paper_results.yaml, README.md, and scripts/run_agentic.sh at commit 49311180bdacb324c596f2e07596c126f2004008 (Defines images allowed, forced-choice condition, full dataset, agent settings, metric, and majority-vote postprocessing.) — supports /benchmark_version, /scope, /protocol, /metrics

Evaluation run

bixbench-v1-5-agentic-mcq-refusal-images

From BixBench v1.5 dataset and evaluation release

bixbench-v1-5-agentic-mcq-refusal-images-five-replicasvv1.5

Evaluated models / systems: Claude 3.5 Sonnet 20241022, GPT-4o (BixBench version not reported)

Scopefull · n=205
Shotszero-shot task initialization
Turnsmulti-turn Aviary SimpleAgent trajectory followed by MCQ postprocessing
System prompt publicYes
Reasoning / effortSimpleAgent with at most 20 steps
BrowserNo
InternetNot reported
DatabasesNot reported
Code executionYes
ContainerYes
External toolsnotebook editing and capsule workspace inspection
Token budgetNot reported
Time / cost budgetmaximum 20 agent steps
Temperature1
SeedNot reported
Repeats5
Gradermultiple-choice conversion with Insufficient information refusal · human review: no
Statisticsaccuracy with 95% Wilson interval and majority-vote analysis
Contaminationpublic benchmark canary
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracyabsolutepercentquestion-replica-weighted meanexact selected option

No numeric result rows are published yet; the verified protocol remains useful.

Evidence

  • repository-path: image model configs, v1.5_paper_results.yaml, README.md, and scripts/run_agentic.sh at commit 49311180bdacb324c596f2e07596c126f2004008 (Defines images allowed, refusal condition, full dataset, agent settings, metric, and majority-vote postprocessing.) — supports /benchmark_version, /scope, /protocol, /metrics

Evaluation run

bixbench-v1-5-agentic-mcq-refusal-no-images

From BixBench v1.5 dataset and evaluation release

bixbench-v1-5-agentic-mcq-refusal-no-images-five-replicasvv1.5

Evaluated models / systems: Claude 3.5 Sonnet 20241022, GPT-4o (BixBench version not reported)

Scopefull · n=205
Shotszero-shot task initialization
Turnsmulti-turn Aviary SimpleAgent trajectory followed by MCQ postprocessing
System prompt publicYes
Reasoning / effortSimpleAgent with at most 20 steps and an instruction to avoid plots/images
BrowserNo
InternetNot reported
DatabasesNot reported
Code executionYes
ContainerYes
External toolsnotebook editing and capsule workspace inspection
Token budgetNot reported
Time / cost budgetmaximum 20 agent steps
Temperature1
SeedNot reported
Repeats5
Gradermultiple-choice conversion with Insufficient information refusal · human review: no
Statisticsmajority-vote accuracy by vote count
Contaminationpublic benchmark canary
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracyabsolutepercentmajority vote over stated five replicas per questionexact selected option

No numeric result rows are published yet; the verified protocol remains useful.

Evidence

  • repository-path: 4o_no_image.yaml, claude_no_image.yaml, v1.5_paper_results.yaml, README.md, and scripts/run_agentic.sh at commit 49311180bdacb324c596f2e07596c126f2004008 (Defines avoid_images true, refusal condition, full dataset, agent settings, and majority-vote postprocessing.) — supports /benchmark_version, /scope, /protocol, /metrics

Evaluation run

bixbench-v1-5-agentic-open-images

From BixBench v1.5 dataset and evaluation release

bixbench-v1-5-agentic-open-images-five-replicasvv1.5

Evaluated models / systems: Claude 3.5 Sonnet 20241022, GPT-4o (BixBench version not reported)

Scopefull · n=205
Shotszero-shot task initialization
Turnsmulti-turn Aviary SimpleAgent trajectory
System prompt publicYes
Reasoning / effortSimpleAgent with at most 20 steps
BrowserNo
InternetNot reported
DatabasesNot reported
Code executionYes
ContainerYes
External toolsnotebook editing and capsule workspace inspection
Token budgetNot reported
Time / cost budgetmaximum 20 agent steps
Temperature1
SeedNot reported
Repeats5
Graderrow-specific open-answer verifier · model: GPT-4o for llm_verifier rows (exact endpoint not reported) · human review: no
Statisticsaccuracy with 95% Wilson interval
Contaminationpublic benchmark canary
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracyabsolutepercentquestion-replica-weighted meanverifier-specific

No numeric result rows are published yet; the verified protocol remains useful.

Evidence

  • repository-path: README.md; bixbench/run_configuration/4o_image.yaml and claude_image.yaml; scripts/run_agentic.sh at commit 49311180bdacb324c596f2e07596c126f2004008 (Reports 205-question v1.5, Docker, SimpleAgent, 20 steps, temperature 1, images allowed, model strings, and stated five replicas.) — supports /benchmark_version, /scope, /protocol, /metrics
  • figure: bixbench_results_comparison.png, Open-answer group (Official accuracy bars with 95% Wilson intervals and no printed scalar labels.) — supports /metrics

Evaluation run

bixbench-v1-5-zero-shot-mcq-no-refusal

From BixBench v1.5 dataset and evaluation release

bixbench-v1-5-zero-shot-mcq-no-refusal-temperature-onevv1.5

Evaluated models / systems: Claude 3.5 Sonnet (BixBench version not reported), GPT-4o (BixBench version not reported)

Scopefull · n=205
Shotszero-shot
Turnssingle-turn model call
System prompt publicYes
Reasoning / effortdefault model behavior
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNo
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
Temperature1
SeedNot reported
Repeats1
Graderexact normalized selected-option match · human review: no
Statisticsaccuracy, precision, and coverage over 205 forced-choice questions
Contaminationpublic benchmark canary
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
accuracyabsolutepercentcorrect over all questionsexact normalized option
precisionabsolutepercentcorrect over sure answersexact normalized option
coverageabsolutepercentsure answers over all questionsNot reported

Results

ModelMetricValuen
GPT-4o (BixBench version not reported)accuracy36.0975609756 percent
74 correct of 205.
205
GPT-4o (BixBench version not reported)precision36.0975609756 percent
74 correct among 205 sure answers.
205
GPT-4o (BixBench version not reported)coverage100 percent
205 sure answers of 205.
205
Claude 3.5 Sonnet (BixBench version not reported)accuracy34.1463414634 percent
70 correct of 205.
205
Claude 3.5 Sonnet (BixBench version not reported)precision34.1463414634 percent
70 correct among 205 sure answers.
205
Claude 3.5 Sonnet (BixBench version not reported)coverage100 percent
205 sure answers of 205.
205

Evidence

  • repository-path: scripts/run_zeroshot.sh, generate_zeroshot_evals.py, bixbench/zero_shot.py, and grade_outputs.py at commit 49311180bdacb324c596f2e07596c126f2004008 (Defines full-set zero-shot forced-choice calls, randomized options, temperature 1.0, and exact option grading.) — supports /benchmark_version, /scope, /protocol, /metrics
  • repository-path: bixbench-v1.5_results/zero_shot_baselines.json at commit 28909d842bc492ecd99bab303279afb29e3cb353 (Exact accuracy, precision, coverage, n_total, n_correct, and n_sure for both forced-choice baselines.) — supports /results

Evaluation run

bixbench-v1-5-zero-shot-mcq-refusal

From BixBench v1.5 dataset and evaluation release

bixbench-v1-5-zero-shot-mcq-refusal-temperature-onevv1.5

Evaluated models / systems: Claude 3.5 Sonnet (BixBench version not reported), GPT-4o (BixBench version not reported)

Scopefull · n=205
Shotszero-shot
Turnssingle-turn model call
System prompt publicYes
Reasoning / effortdefault model behavior
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNo
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
Temperature1
SeedNot reported
Repeats1
Graderexact normalized selected-option match with a sure/refusal flag · human review: no
Statisticsaccuracy, precision among non-refusals, and non-refusal coverage over 205 questions
Contaminationpublic benchmark canary
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
accuracyabsolutepercentcorrect over all questionsexact normalized option
precisionabsolutepercentcorrect over non-refusal answersexact normalized option
coverageabsolutepercentnon-refusal answers over all questionsNot reported

Results

ModelMetricValuen
GPT-4o (BixBench version not reported)accuracy3.9024390244 percent
8 correct of 205.
205
GPT-4o (BixBench version not reported)precision40 percent
8 correct among 20 non-refusals.
205
GPT-4o (BixBench version not reported)coverage9.756097561 percent
20 non-refusals of 205.
205
Claude 3.5 Sonnet (BixBench version not reported)accuracy8.2926829268 percent
17 correct of 205.
205
Claude 3.5 Sonnet (BixBench version not reported)precision38.6363636364 percent
17 correct among 44 non-refusals.
205
Claude 3.5 Sonnet (BixBench version not reported)coverage21.4634146341 percent
44 non-refusals of 205.
205

Evidence

  • repository-path: scripts/run_zeroshot.sh, generate_zeroshot_evals.py, bixbench/zero_shot.py, and grade_outputs.py at commit 49311180bdacb324c596f2e07596c126f2004008 (Defines full-set zero-shot MCQ calls, randomized options, refusal flag, temperature 1.0, and exact option grading.) — supports /benchmark_version, /scope, /protocol, /metrics
  • repository-path: bixbench-v1.5_results/zero_shot_baselines.json at commit 28909d842bc492ecd99bab303279afb29e3cb353 (Exact accuracy, precision, coverage, n_total, n_correct, and n_sure for both refusal-condition baselines.) — supports /results

Evaluation run

bixbench-v1-5-zero-shot-open

From BixBench v1.5 dataset and evaluation release

bixbench-v1-5-zero-shot-open-temperature-onevv1.5

Evaluated models / systems: Claude 3.5 Sonnet (BixBench version not reported), GPT-4o (BixBench version not reported)

Scopefull · n=205
Shotszero-shot
Turnssingle-turn model call
System prompt publicYes
Reasoning / effortdefault model behavior
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNo
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
Temperature1
SeedNot reported
Repeats1
Graderrow-specific string, numeric-range, or LLM verifier · model: GPT-4o for llm_verifier rows (exact endpoint not reported) · human review: no
Statisticsaccuracy over 205 questions, plus precision and coverage from correct/sure flags
Contaminationpublic benchmark canary
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
accuracyabsolutepercentcorrect over all 205 questionsverifier-specific
precisionabsolutepercentcorrect over sure predictionsverifier-specific
coverageabsolutepercentsure predictions over all 205 questionsNot reported

Results

ModelMetricValuen
GPT-4o (BixBench version not reported)accuracy2.9268292683 percent
6 correct of 205.
205
GPT-4o (BixBench version not reported)precision2.9268292683 percent
6 correct among 205 sure predictions.
205
GPT-4o (BixBench version not reported)coverage100 percent
205 sure predictions of 205.
205
Claude 3.5 Sonnet (BixBench version not reported)accuracy2.9268292683 percent
6 correct of 205.
205
Claude 3.5 Sonnet (BixBench version not reported)precision2.9268292683 percent
6 correct among 205 sure predictions.
205
Claude 3.5 Sonnet (BixBench version not reported)coverage100 percent
205 sure predictions of 205.
205

Evidence

  • repository-path: generate_zeroshot_evals.py, grade_outputs.py, bixbench/zero_shot.py, bixbench/graders.py, and scripts/run_zeroshot.sh at commit 49311180bdacb324c596f2e07596c126f2004008 (Defines full-dataset single calls, temperature 1.0, public prompts, no tools/files, and mixed open-answer verifiers.) — supports /benchmark_version, /scope, /protocol, /metrics
  • repository-path: bixbench-v1.5_results/zero_shot_baselines.json at commit 28909d842bc492ecd99bab303279afb29e3cb353 (Both models have n_total 205, n_correct 6, n_sure 205.) — supports /results