track · audited · verified 2026-07-21

BLADE End-to-End Analysis Generation

The BLADE track requiring a conceptual-variable specification, executable data-transformation function, and statistical-model function for each open-ended research question and dataset.

Benchmark definition

What is counted

Version
arXiv v3
Total
12 (paired research questions and datasets requiring a complete generated analysis)
Task formats
open-ended end-to-end scientific analysis
Capabilities
Data analysisCodingTool useScientific reasoning
Modalities
TextTableCode

Version history

VersionStatusRelease / as-ofTotalFormal tracks
arXiv v3
blade-generation-arxiv-v3
current2025-11-1012 (paired research questions and datasets requiring a complete generated analysis)None registered

Tracks and subsets

IDCountBasisPartition?Notes
Ground-truth analysis decisions
blade-generation-ground-truth-decisions
536unique defensible decisions used by the grader across all source tasksNoA grader-reference unit, not a partition of the 12 generated-analysis tasks.
Conceptual-variable decisions
blade-generation-conceptual-decisions
118conceptual-variable decisions in the expert ground truthNoA decision-reference count.
Transform decisions
blade-generation-transform-decisions
246discrete transformation decisions in the expert ground truthNoA decision-reference count.
Modeling decisions
blade-generation-modeling-decisions
172statistical-model and formula decisions in the expert ground truthNoA decision-reference count.

Scientific Task Atlas

Scientific task classification

complete for arXiv v3. Formal executable-analysis generation track.

Scientific taskCoverageCountMappingEvidence
End-to-end computational analysisexplicitly-in-scope12 problems
paired research questions and datasets requiring a complete generated analysis
official-track
high confidence
blade-generation-evidence-counts
The track grades complete analysis decisions and execution.

Scientific coverage notes

DomainCoverageCountInterpretation
Life scienceobserved4Four of the twelve official source questions are explicitly biological or ecological: AMTL, Crofoot, Panda_nuts, and Fish.
Protein designnot-in-scope0None of the twelve analysis-generation tasks concerns protein design.
Protein-protein bindingnot-in-scope0None of the twelve analysis-generation tasks concerns protein-protein binding.
Protein-ligand bindingnot-in-scope0None of the twelve analysis-generation tasks concerns protein-ligand binding.
Genomicsnot-in-scope0None of the twelve analysis-generation tasks is an omics-analysis task.

Evaluation registry

Works and run settings

A setting change—scope, prompt, tools, budget, grader, or repeats—creates a separate run. Charts never cross a comparability group.

blade-arxiv-v3-generation-one-turn-40-temperature-08varXiv v3

Evaluated models / systems: Claude 3.5 Sonnet (claude-3-5-sonnet-20240620; BLADE), CodeLlama Instruct 7B (BLADE), DeepSeek-Coder Instruct 6.7B (BLADE), Gemini 1.5 Pro (BLADE snapshot not reported), GPT-3.5 Turbo (BLADE snapshot not reported), GPT-4o (BLADE snapshot not reported), Llama 3 70B (BLADE snapshot not reported), Llama 3 8B (BLADE snapshot not reported), Mixtral 8x22B (BLADE snapshot not reported)

Scopefull · n=12
Shotsone-shot
Turnssingle-turn
System prompt publicYes
Reasoning / effortdirect prompted generation without an agent loop
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNo
External toolsnone
Token budgetNot reported
Time / cost budgetNot reported
Temperature0.8
SeedNot reported
Repeats40
Graderhybrid automatic executable and decision matching · model: GPT-4o (exact snapshot not reported) · human review: no
StatisticsFor each dataset and decision type, average precision across all runs and coverage@10 feed a harmonic-mean F1; decision-type F1 values are weighted by ground-truth decision counts. Reported F1 is the mean of 1,000 bootstrap resamples with a percentile-derived 95% interval.
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Average precisionabsolutepercentmean of per-run precision followed by macro averaging across datasets for each decision typeGPT-4o semantic matching plus transformation value/graph matching
Coverage@10absolutepercentground-truth coverage of the union of ten sampled runs followed by macro averaging across datasets for each decision typerun errors count as zero-coverage generations
Decision-type weighted F1-scoreabsolutepercentweighted mean of per-decision-type harmonic means of average precision and coverage@101000-run bootstrap mean and 95% interval

Results

ModelMetricValuen
CodeLlama Instruct 7B (BLADE)Decision-type weighted F1-score16.8 percent
Forty one-turn generations per dataset; Table 2.
12
DeepSeek-Coder Instruct 6.7B (BLADE)Decision-type weighted F1-score33.9 percent
Forty one-turn generations per dataset; Table 2.
12
Llama 3 8B (BLADE snapshot not reported)Decision-type weighted F1-score29.6 percent
Forty one-turn generations per dataset; exact served snapshot not reported.
12
Llama 3 70B (BLADE snapshot not reported)Decision-type weighted F1-score36.3 percent
Forty one-turn generations per dataset; exact served snapshot not reported.
12
Mixtral 8x22B (BLADE snapshot not reported)Decision-type weighted F1-score40.1 percent
Forty one-turn generations per dataset; exact served snapshot not reported.
12
GPT-3.5 Turbo (BLADE snapshot not reported)Decision-type weighted F1-score30.5 percent
Forty one-turn generations per dataset; dated endpoint snapshot not reported.
12
GPT-4o (BLADE snapshot not reported)Decision-type weighted F1-score41.7 percent
Forty one-turn generations per dataset; dated endpoint snapshot not reported.
12
Gemini 1.5 Pro (BLADE snapshot not reported)Decision-type weighted F1-score41.1 percent
Forty one-turn generations per dataset; immutable endpoint snapshot not reported.
12
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620; BLADE)Decision-type weighted F1-score43.9 percent
Forty one-turn generations per dataset; Table 2.
12

Evidence

  • section: arXiv v3 §§4.2, 5–6 and Appendices A.6–A.7 (Defines the three submitted artifacts, one-shot one-turn prompt, temperature 0.8, 40 runs, GPT-4o-assisted evaluation, average precision, coverage@10, weighted F1, and 1,000-iteration bootstrap.) — supports /benchmark_version, /scope, /protocol, /metrics, /comparability_group
  • table: Table 2, One-turn Setting (Prints all nine decision-type weighted F1 point estimates and 95% confidence intervals.) — supports /model_ids, /results
  • repository-path: run_scripts/sh_run_one_turn.sh, blade_bench/baselines/run.py, blade_bench/baselines/lm/analysis.py, blade_bench/conf/llm_config.yml, and blade_bench/eval/ at commit 6118fa8d5007b91aa8c91c518182db82446a4547 (Confirms all 12 datasets, 40 runs, direct generation, public prompt/evaluator, and the available exact model strings.) — supports /model_ids, /protocol/shots, /protocol/turns, /protocol/system_prompt_public, /protocol/tools, /protocol/repeats, /protocol/grader
blade-arxiv-v3-generation-react-20-temperature-08varXiv v3

Evaluated models / systems: Claude 3.5 Sonnet (claude-3-5-sonnet-20240620; BLADE), Gemini 1.5 Pro (BLADE snapshot not reported), GPT-3.5 Turbo (BLADE snapshot not reported), GPT-4o (BLADE snapshot not reported), Mixtral 8x22B (BLADE snapshot not reported)

Scopefull · n=12
Shotsone-shot ReAct trajectory
Turnsmulti-turn
System prompt publicYes
Reasoning / effortReAct loop with full prior thought, action, and observation context
BrowserNo
InternetNo
DatabasesNo
Code executionYes
ContainerYes
External toolscomputational notebook cell execution
Token budgetNot reported
Time / cost budgetmaximum 10 agent steps
Temperature0.8
SeedNot reported
Repeats20
Graderhybrid automatic executable and decision matching · model: GPT-4o (exact snapshot not reported) · human review: no
StatisticsFor each dataset and decision type, average precision across all runs and coverage@10 feed a harmonic-mean F1; decision-type F1 values are weighted by ground-truth decision counts. Reported F1 is the mean of 1,000 bootstrap resamples with a percentile-derived 95% interval.
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Average precisionabsolutepercentmean of per-run precision followed by macro averaging across datasets for each decision typeGPT-4o semantic matching plus transformation value/graph matching
Coverage@10absolutepercentground-truth coverage of the union of ten sampled runs followed by macro averaging across datasets for each decision typerun errors count as zero-coverage generations
Decision-type weighted F1-scoreabsolutepercentweighted mean of per-decision-type harmonic means of average precision and coverage@101000-run bootstrap mean and 95% interval

Results

ModelMetricValuen
Mixtral 8x22B (BLADE snapshot not reported)Decision-type weighted F1-score40.8 percent
Twenty ReAct trajectories per dataset; exact served snapshot not reported.
12
GPT-3.5 Turbo (BLADE snapshot not reported)Decision-type weighted F1-score37.2 percent
Twenty ReAct trajectories per dataset; dated endpoint snapshot not reported.
12
GPT-4o (BLADE snapshot not reported)Decision-type weighted F1-score44.8 percent
Twenty ReAct trajectories per dataset; dated endpoint snapshot not reported.
12
Gemini 1.5 Pro (BLADE snapshot not reported)Decision-type weighted F1-score40.1 percent
Twenty ReAct trajectories per dataset; immutable endpoint snapshot not reported.
12
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620; BLADE)Decision-type weighted F1-score43.1 percent
Twenty ReAct trajectories per dataset; Table 2.
12

Evidence

  • section: arXiv v3 §§5–6 and Appendices A.5–A.7 (Defines the ReAct notebook, one trajectory example, Python 3.10 environment, ten steps, temperature 0.8, 20 runs, GPT-4o-assisted evaluation, metrics, and 1,000-iteration bootstrap.) — supports /benchmark_version, /scope, /protocol, /metrics, /comparability_group
  • table: Table 2, Agent Setting (Prints all five ReAct decision-type weighted F1 point estimates and 95% confidence intervals.) — supports /model_ids, /results
  • repository-path: run_scripts/sh_run_agent.sh, blade_bench/baselines/agent/react_agent.py, blade_bench/baselines/run.py, blade_bench/conf/llm_config.yml, and blade_bench/eval/ at commit 6118fa8d5007b91aa8c91c518182db82446a4547 (Confirms all 12 datasets, 20 runs, ten-step ReAct execution, public prompt/evaluator, and available exact model strings.) — supports /model_ids, /protocol/shots, /protocol/turns, /protocol/system_prompt_public, /protocol/reasoning, /protocol/tools, /protocol/time_budget, /protocol/repeats, /protocol/grader

Comparable result views

Decision-type weighted F1-score

blade-creator-paper · blade-arxiv-v3-generation-one-turn-40-temperature-08

CSV ↓
Accessible data table
ModelValueComparability group
CodeLlama Instruct 7B (BLADE)16.8blade-arxiv-v3-generation-one-turn-40-temperature-08
DeepSeek-Coder Instruct 6.7B (BLADE)33.9blade-arxiv-v3-generation-one-turn-40-temperature-08
Llama 3 8B (BLADE snapshot not reported)29.6blade-arxiv-v3-generation-one-turn-40-temperature-08
Llama 3 70B (BLADE snapshot not reported)36.3blade-arxiv-v3-generation-one-turn-40-temperature-08
Mixtral 8x22B (BLADE snapshot not reported)40.1blade-arxiv-v3-generation-one-turn-40-temperature-08
GPT-3.5 Turbo (BLADE snapshot not reported)30.5blade-arxiv-v3-generation-one-turn-40-temperature-08
GPT-4o (BLADE snapshot not reported)41.7blade-arxiv-v3-generation-one-turn-40-temperature-08
Gemini 1.5 Pro (BLADE snapshot not reported)41.1blade-arxiv-v3-generation-one-turn-40-temperature-08
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620; BLADE)43.9blade-arxiv-v3-generation-one-turn-40-temperature-08

Decision-type weighted F1-score

blade-creator-react · blade-arxiv-v3-generation-react-20-temperature-08

CSV ↓
Accessible data table
ModelValueComparability group
Mixtral 8x22B (BLADE snapshot not reported)40.8blade-arxiv-v3-generation-react-20-temperature-08
GPT-3.5 Turbo (BLADE snapshot not reported)37.2blade-arxiv-v3-generation-react-20-temperature-08
GPT-4o (BLADE snapshot not reported)44.8blade-arxiv-v3-generation-react-20-temperature-08
Gemini 1.5 Pro (BLADE snapshot not reported)40.1blade-arxiv-v3-generation-react-20-temperature-08
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620; BLADE)43.1blade-arxiv-v3-generation-react-20-temperature-08

Evidence and change history

Source locators remain visible; expand an item to inspect the exact Registry fields it supports.

blade-generation-paper-resource · section: arXiv v3 §§3–5 and Appendix A.2 (Defines twelve full-analysis tasks and 536 decision references: 118 conceptual-variable, 246 transform, and 172 modeling decisions.) · Supports 16 fields

Open source →

  • /name
  • /summary
  • /kind
  • /parent_id
  • /organizations
  • /release_date
  • /latest_version
  • /capabilities
  • /modalities
  • /task_formats
  • /task_counts
  • /task_counts/total
  • /task_counts/basis
  • /task_counts/subsets
  • /versions/0/task_counts
  • /scientific_task_classification/entries/0
blade-generation-repository-resource · repository-path: run_gen_analyses.py, run_get_eval.py, run_scripts/, blade_bench/baselines/, and blade_bench/eval/ at commit 6118fa8d5007b91aa8c91c518182db82446a4547 (Public one-turn/ReAct runners, ten-step notebook sandbox, code execution, GPT-4o-assisted conversion/matching, structural matching, and result aggregation.) · Supports 5 fields

Open source →

  • /access
  • /access/level
  • /access/license
  • /resources
  • /implementations
blade-generation-paper-resource · table: Table 3 (Four explicitly biological/ecological source questions and no protein, binding, or omics analysis task.) · Supports 3 fields

Open source →

  • /domains
  • /coverage_notes
  • /access/biosafety_notes

View source-level modification history on GitHub →