suite · audited · verified 2026-07-21

BLADE

A cross-domain suite for discerning defensible analysis decisions and generating executable end-to-end analyses for open-ended scientific research questions; four of its twelve source questions are explicitly biological or ecological.

Task-unit and protocol audit: BLADE uses 12 source research-question/dataset pairs. Its two formal tracks contain 188 MCQs (20 conceptual-variable and 168 transformation questions) and 12 end-to-end analysis tasks scored against 536 ground-truth decisions (118 conceptual-variable, 246 transformation, and 172 modeling decisions). MCQ, one-turn generation, and ReAct settings are separated; four source questions are biological or ecological, and none is a protein, binding, or omics task.

Benchmark definition

What is counted

Version
arXiv v3
Total
12 (paired real-world research questions and datasets used as the source units for BLADE)
Task formats
multiple choice; open-ended end-to-end scientific analysis
Capabilities
KnowledgeClassificationData analysisCodingTool useScientific reasoning
Modalities
TextTableCode

Version history

VersionStatusRelease / as-ofTotalFormal tracks
arXiv v3
blade-arxiv-v3
current2025-11-1012 (paired real-world research questions and datasets used as the source units for BLADE)blade-mcq, blade-analysis-generation

Tracks and subsets

IDCountBasisPartition?Notes
Multiple-choice questions
blade-mcq-items
188individual decision-discrimination questions across the source datasetsNoA separate question unit, not a partition of the 12 research-question/dataset pairs.
Conceptual-variable MCQs
blade-mcq-conceptual
20individual conceptual-variable discrimination questionsNoOne component of the 188-question MCQ track.
Transformation MCQs
blade-mcq-transform
168individual transformation-decision discrimination questionsNoOne component of the 188-question MCQ track.
Ground-truth analysis decisions
blade-ground-truth-decisions
536unique defensible decisions in the expert ground-truth decision spaceNoA grader-reference unit, not a count of benchmark tasks.
Ground-truth conceptual-variable decisions
blade-ground-truth-conceptual
118conceptual-variable decisions in the ground-truth decision spaceNoTogether with transform and modeling decisions, partitions the 536 decision references rather than the 12 source tasks.
Ground-truth transform decisions
blade-ground-truth-transform
246discrete data-transformation decisions in the ground-truth decision spaceNoTogether with conceptual-variable and modeling decisions, partitions the 536 decision references.
Ground-truth modeling decisions
blade-ground-truth-modeling
172statistical-model and model-formula decisions in the ground-truth decision spaceNoTogether with conceptual-variable and transform decisions, partitions the 536 decision references.

Registered child tracks

BLADE End-to-End Analysis Generation

The BLADE track requiring a conceptual-variable specification, executable data-transformation function, and statistical-model function for each open-ended research question and dataset.

2 evaluation run(s)

BLADE Decision-Discrimination MCQ

The BLADE track for selecting the most or least justifiable conceptual-variable and data-transformation decisions for a research question and dataset.

1 evaluation run(s)

Scientific Task Atlas

Scientific task classification

partial for arXiv v3. BLADE is cross-domain; four source questions are biological or ecological.

Scientific taskCoverageCountMappingEvidence
End-to-end computational analysisexplicitly-in-scope12 problems
paired real-world research questions and datasets used as the source units for BLADE
official-track
high confidence
blade-evidence-current-counts
Coverage is suite-wide; only four source questions are in life science.

Scientific coverage notes

DomainCoverageCountInterpretation
Life scienceobserved4Registry mapping of four explicitly biological or ecological source questions in paper Table 3: AMTL, Crofoot, Panda_nuts, and Fish. BLADE itself is cross-domain.
Protein designnot-in-scope0None of the twelve official research questions concerns protein design or optimization.
Protein-protein bindingnot-in-scope0None of the twelve official research questions concerns protein-protein binding.
Protein-ligand bindingnot-in-scope0None of the twelve official research questions concerns protein-ligand binding.
Genomicsnot-in-scope0The official question inventory contains no genomics or omics-analysis task.
Transcriptomicsnot-in-scope0The official question inventory contains no transcriptomics task.
Single-cellnot-in-scope0The official question inventory contains no single-cell task.

Evaluation registry

Works and run settings

A setting change—scope, prompt, tools, budget, grader, or repeats—creates a separate run. Charts never cross a comparability group.

blade-arxiv-v3-generation-one-turn-40-temperature-08varXiv v3

Evaluated models / systems: Claude 3.5 Sonnet (claude-3-5-sonnet-20240620; BLADE), CodeLlama Instruct 7B (BLADE), DeepSeek-Coder Instruct 6.7B (BLADE), Gemini 1.5 Pro (BLADE snapshot not reported), GPT-3.5 Turbo (BLADE snapshot not reported), GPT-4o (BLADE snapshot not reported), Llama 3 70B (BLADE snapshot not reported), Llama 3 8B (BLADE snapshot not reported), Mixtral 8x22B (BLADE snapshot not reported)

Scopefull · n=12
Shotsone-shot
Turnssingle-turn
System prompt publicYes
Reasoning / effortdirect prompted generation without an agent loop
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNo
External toolsnone
Token budgetNot reported
Time / cost budgetNot reported
Temperature0.8
SeedNot reported
Repeats40
Graderhybrid automatic executable and decision matching · model: GPT-4o (exact snapshot not reported) · human review: no
StatisticsFor each dataset and decision type, average precision across all runs and coverage@10 feed a harmonic-mean F1; decision-type F1 values are weighted by ground-truth decision counts. Reported F1 is the mean of 1,000 bootstrap resamples with a percentile-derived 95% interval.
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Average precisionabsolutepercentmean of per-run precision followed by macro averaging across datasets for each decision typeGPT-4o semantic matching plus transformation value/graph matching
Coverage@10absolutepercentground-truth coverage of the union of ten sampled runs followed by macro averaging across datasets for each decision typerun errors count as zero-coverage generations
Decision-type weighted F1-scoreabsolutepercentweighted mean of per-decision-type harmonic means of average precision and coverage@101000-run bootstrap mean and 95% interval

Results

ModelMetricValuen
CodeLlama Instruct 7B (BLADE)Decision-type weighted F1-score16.8 percent
Forty one-turn generations per dataset; Table 2.
12
DeepSeek-Coder Instruct 6.7B (BLADE)Decision-type weighted F1-score33.9 percent
Forty one-turn generations per dataset; Table 2.
12
Llama 3 8B (BLADE snapshot not reported)Decision-type weighted F1-score29.6 percent
Forty one-turn generations per dataset; exact served snapshot not reported.
12
Llama 3 70B (BLADE snapshot not reported)Decision-type weighted F1-score36.3 percent
Forty one-turn generations per dataset; exact served snapshot not reported.
12
Mixtral 8x22B (BLADE snapshot not reported)Decision-type weighted F1-score40.1 percent
Forty one-turn generations per dataset; exact served snapshot not reported.
12
GPT-3.5 Turbo (BLADE snapshot not reported)Decision-type weighted F1-score30.5 percent
Forty one-turn generations per dataset; dated endpoint snapshot not reported.
12
GPT-4o (BLADE snapshot not reported)Decision-type weighted F1-score41.7 percent
Forty one-turn generations per dataset; dated endpoint snapshot not reported.
12
Gemini 1.5 Pro (BLADE snapshot not reported)Decision-type weighted F1-score41.1 percent
Forty one-turn generations per dataset; immutable endpoint snapshot not reported.
12
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620; BLADE)Decision-type weighted F1-score43.9 percent
Forty one-turn generations per dataset; Table 2.
12

Evidence

  • section: arXiv v3 §§4.2, 5–6 and Appendices A.6–A.7 (Defines the three submitted artifacts, one-shot one-turn prompt, temperature 0.8, 40 runs, GPT-4o-assisted evaluation, average precision, coverage@10, weighted F1, and 1,000-iteration bootstrap.) — supports /benchmark_version, /scope, /protocol, /metrics, /comparability_group
  • table: Table 2, One-turn Setting (Prints all nine decision-type weighted F1 point estimates and 95% confidence intervals.) — supports /model_ids, /results
  • repository-path: run_scripts/sh_run_one_turn.sh, blade_bench/baselines/run.py, blade_bench/baselines/lm/analysis.py, blade_bench/conf/llm_config.yml, and blade_bench/eval/ at commit 6118fa8d5007b91aa8c91c518182db82446a4547 (Confirms all 12 datasets, 40 runs, direct generation, public prompt/evaluator, and the available exact model strings.) — supports /model_ids, /protocol/shots, /protocol/turns, /protocol/system_prompt_public, /protocol/tools, /protocol/repeats, /protocol/grader
blade-arxiv-v3-generation-react-20-temperature-08varXiv v3

Evaluated models / systems: Claude 3.5 Sonnet (claude-3-5-sonnet-20240620; BLADE), Gemini 1.5 Pro (BLADE snapshot not reported), GPT-3.5 Turbo (BLADE snapshot not reported), GPT-4o (BLADE snapshot not reported), Mixtral 8x22B (BLADE snapshot not reported)

Scopefull · n=12
Shotsone-shot ReAct trajectory
Turnsmulti-turn
System prompt publicYes
Reasoning / effortReAct loop with full prior thought, action, and observation context
BrowserNo
InternetNo
DatabasesNo
Code executionYes
ContainerYes
External toolscomputational notebook cell execution
Token budgetNot reported
Time / cost budgetmaximum 10 agent steps
Temperature0.8
SeedNot reported
Repeats20
Graderhybrid automatic executable and decision matching · model: GPT-4o (exact snapshot not reported) · human review: no
StatisticsFor each dataset and decision type, average precision across all runs and coverage@10 feed a harmonic-mean F1; decision-type F1 values are weighted by ground-truth decision counts. Reported F1 is the mean of 1,000 bootstrap resamples with a percentile-derived 95% interval.
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Average precisionabsolutepercentmean of per-run precision followed by macro averaging across datasets for each decision typeGPT-4o semantic matching plus transformation value/graph matching
Coverage@10absolutepercentground-truth coverage of the union of ten sampled runs followed by macro averaging across datasets for each decision typerun errors count as zero-coverage generations
Decision-type weighted F1-scoreabsolutepercentweighted mean of per-decision-type harmonic means of average precision and coverage@101000-run bootstrap mean and 95% interval

Results

ModelMetricValuen
Mixtral 8x22B (BLADE snapshot not reported)Decision-type weighted F1-score40.8 percent
Twenty ReAct trajectories per dataset; exact served snapshot not reported.
12
GPT-3.5 Turbo (BLADE snapshot not reported)Decision-type weighted F1-score37.2 percent
Twenty ReAct trajectories per dataset; dated endpoint snapshot not reported.
12
GPT-4o (BLADE snapshot not reported)Decision-type weighted F1-score44.8 percent
Twenty ReAct trajectories per dataset; dated endpoint snapshot not reported.
12
Gemini 1.5 Pro (BLADE snapshot not reported)Decision-type weighted F1-score40.1 percent
Twenty ReAct trajectories per dataset; immutable endpoint snapshot not reported.
12
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620; BLADE)Decision-type weighted F1-score43.1 percent
Twenty ReAct trajectories per dataset; Table 2.
12

Evidence

  • section: arXiv v3 §§5–6 and Appendices A.5–A.7 (Defines the ReAct notebook, one trajectory example, Python 3.10 environment, ten steps, temperature 0.8, 20 runs, GPT-4o-assisted evaluation, metrics, and 1,000-iteration bootstrap.) — supports /benchmark_version, /scope, /protocol, /metrics, /comparability_group
  • table: Table 2, Agent Setting (Prints all five ReAct decision-type weighted F1 point estimates and 95% confidence intervals.) — supports /model_ids, /results
  • repository-path: run_scripts/sh_run_agent.sh, blade_bench/baselines/agent/react_agent.py, blade_bench/baselines/run.py, blade_bench/conf/llm_config.yml, and blade_bench/eval/ at commit 6118fa8d5007b91aa8c91c518182db82446a4547 (Confirms all 12 datasets, 20 runs, ten-step ReAct execution, public prompt/evaluator, and available exact model strings.) — supports /model_ids, /protocol/shots, /protocol/turns, /protocol/system_prompt_public, /protocol/reasoning, /protocol/tools, /protocol/time_budget, /protocol/repeats, /protocol/grader

Evaluation run

blade-creator-decision-mcq

From BLADE: Benchmarking Language Model Agents for Data-Driven Science

blade-arxiv-v3-mcq-zero-shot-temperature-0varXiv v3

Evaluated models / systems: Claude 3.5 Sonnet (claude-3-5-sonnet-20240620; BLADE), CodeLlama Instruct 7B (BLADE), DeepSeek-Coder Instruct 6.7B (BLADE), Gemini 1.5 Pro (BLADE snapshot not reported), GPT-3.5 Turbo (BLADE snapshot not reported), GPT-4o (BLADE snapshot not reported), Llama 3 70B (BLADE snapshot not reported), Llama 3 8B (BLADE snapshot not reported), Mixtral 8x22B (BLADE snapshot not reported)

Scopefull · n=188
Shotszero-shot
Turnssingle-turn
System prompt publicYes
Reasoning / effortdirect multiple-choice completion
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNo
External toolsnone
Token budgetNot reported
Time / cost budgetNot reported
Temperature0
SeedNot reported
Repeats1
Graderdeterministic exact option match · human review: no
Statisticsquestion-weighted accuracy with 95% confidence intervals
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracyabsolutepercentcorrect choices divided by all 188 MCQsexact option identifier

No numeric result rows are published yet; the verified protocol remains useful.

Evidence

  • section: arXiv v3 §§4.1 and 6, Figure 3, and Appendix A.6 Figure 19 (Reports 188 MCQs, nine evaluated model families, temperature zero, accuracy with 95% intervals, and the public prompt.) — supports /benchmark_version, /model_ids, /scope, /protocol/temperature, /protocol/statistical, /metrics, /comparability_group
  • repository-path: run_mcq.py, run_scripts/sh_run_mcq.sh, blade_bench/baselines/lm/mcq.py, blade_bench/eval/datamodel/run_mcq.py, and blade_bench/conf/llm_config.yml at commit 6118fa8d5007b91aa8c91c518182db82446a4547 (Confirms one direct prompt per item, no tool loop, exact-choice grading, complete public prompt, 20/168 question components, and available exact model strings.) — supports /scope, /protocol/shots, /protocol/turns, /protocol/system_prompt_public, /protocol/reasoning, /protocol/tools, /protocol/repeats, /protocol/grader, /model_ids

Comparable result views

Decision-type weighted F1-score

blade-creator-paper · blade-arxiv-v3-generation-one-turn-40-temperature-08

CSV ↓
Accessible data table
ModelValueComparability group
CodeLlama Instruct 7B (BLADE)16.8blade-arxiv-v3-generation-one-turn-40-temperature-08
DeepSeek-Coder Instruct 6.7B (BLADE)33.9blade-arxiv-v3-generation-one-turn-40-temperature-08
Llama 3 8B (BLADE snapshot not reported)29.6blade-arxiv-v3-generation-one-turn-40-temperature-08
Llama 3 70B (BLADE snapshot not reported)36.3blade-arxiv-v3-generation-one-turn-40-temperature-08
Mixtral 8x22B (BLADE snapshot not reported)40.1blade-arxiv-v3-generation-one-turn-40-temperature-08
GPT-3.5 Turbo (BLADE snapshot not reported)30.5blade-arxiv-v3-generation-one-turn-40-temperature-08
GPT-4o (BLADE snapshot not reported)41.7blade-arxiv-v3-generation-one-turn-40-temperature-08
Gemini 1.5 Pro (BLADE snapshot not reported)41.1blade-arxiv-v3-generation-one-turn-40-temperature-08
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620; BLADE)43.9blade-arxiv-v3-generation-one-turn-40-temperature-08

Decision-type weighted F1-score

blade-creator-react · blade-arxiv-v3-generation-react-20-temperature-08

CSV ↓
Accessible data table
ModelValueComparability group
Mixtral 8x22B (BLADE snapshot not reported)40.8blade-arxiv-v3-generation-react-20-temperature-08
GPT-3.5 Turbo (BLADE snapshot not reported)37.2blade-arxiv-v3-generation-react-20-temperature-08
GPT-4o (BLADE snapshot not reported)44.8blade-arxiv-v3-generation-react-20-temperature-08
Gemini 1.5 Pro (BLADE snapshot not reported)40.1blade-arxiv-v3-generation-react-20-temperature-08
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620; BLADE)43.1blade-arxiv-v3-generation-react-20-temperature-08

Evidence and change history

Source locators remain visible; expand an item to inspect the exact Registry fields it supports.

BLADE: Benchmarking Language Model Agents for Data-Driven Science · section: Title, abstract, and §§1–4 (Defines BLADE, the twelve source research questions/datasets, its decision-based purpose, and the two task types.) · Supports 9 fields

Open source →

  • /name
  • /aliases
  • /summary
  • /organizations
  • /release_date
  • /kind
  • /capabilities
  • /modalities
  • /task_formats
blade-arxiv-v3-resource · section: arXiv v3 §§1, 4.1–4.2, 6 and Appendix A.2 (Reports 12 paired questions/datasets, 188 MCQs (20 conceptual-variable and 168 transform), and 536 ground-truth decisions (118 conceptual-variable, 246 transform, and 172 modeling).) · Supports 8 fields

Open source →

  • /latest_version
  • /task_counts
  • /task_counts/total
  • /task_counts/basis
  • /task_counts/subsets
  • /versions/0/task_counts
  • /versions/0/formal_tracks
  • /scientific_task_classification/entries/0
blade-arxiv-v3-resource · table: Table 3, complete twelve-question inventory (AMTL, Crofoot, and Panda_nuts are labeled Evolutionary Biology; Fish is ecological/health-and-well-being. No official question is a protein, binding, genomics, transcriptomics, or single-cell task.) · Supports 3 fields

Open source →

  • /domains
  • /coverage_notes
  • /access/biosafety_notes
blade-repository-resource · repository-path: README.md, pyproject.toml, LICENSE, blade_bench/datasets/LICENSE, blade_bench/datasets/*, and blade_bench/eval/* at commit 6118fa8d5007b91aa8c91c518182db82446a4547 (Package 0.1.1; Apache-2.0 code; ODC-By-1.0 data; public task, annotation, baseline, and grader files.) · Supports 7 fields

Open source →

  • /access/level
  • /access/tasks
  • /access/artifacts
  • /access/grader
  • /access/license
  • /resources
  • /implementations
blade-repository-resource · repository-path: blade_bench/datasets/*/mcq_dataset.json and run_mcq.py at commit 6118fa8d5007b91aa8c91c518182db82446a4547 (The public files contain 20 conceptual-variable and 168 transformation MCQs across 11 source datasets, totaling 188.) · Supports 3 fields

Open source →

  • /task_counts/subsets/0/count
  • /task_counts/subsets/1/count
  • /task_counts/subsets/2/count

View source-level modification history on GitHub →