BLADE End-to-End Analysis Generation
The BLADE track requiring a conceptual-variable specification, executable data-transformation function, and statistical-model function for each open-ended research question and dataset.
2 evaluation run(s)
suite · audited · verified 2026-07-21
A cross-domain suite for discerning defensible analysis decisions and generating executable end-to-end analyses for open-ended scientific research questions; four of its twelve source questions are explicitly biological or ecological.
Benchmark definition
| Version | Status | Release / as-of | Total | Formal tracks |
|---|---|---|---|---|
arXiv v3blade-arxiv-v3 | current | 2025-11-10 | 12 (paired real-world research questions and datasets used as the source units for BLADE) | blade-mcq, blade-analysis-generation |
| ID | Count | Basis | Partition? | Notes |
|---|---|---|---|---|
Multiple-choice questionsblade-mcq-items | 188 | individual decision-discrimination questions across the source datasets | No | A separate question unit, not a partition of the 12 research-question/dataset pairs. |
Conceptual-variable MCQsblade-mcq-conceptual | 20 | individual conceptual-variable discrimination questions | No | One component of the 188-question MCQ track. |
Transformation MCQsblade-mcq-transform | 168 | individual transformation-decision discrimination questions | No | One component of the 188-question MCQ track. |
Ground-truth analysis decisionsblade-ground-truth-decisions | 536 | unique defensible decisions in the expert ground-truth decision space | No | A grader-reference unit, not a count of benchmark tasks. |
Ground-truth conceptual-variable decisionsblade-ground-truth-conceptual | 118 | conceptual-variable decisions in the ground-truth decision space | No | Together with transform and modeling decisions, partitions the 536 decision references rather than the 12 source tasks. |
Ground-truth transform decisionsblade-ground-truth-transform | 246 | discrete data-transformation decisions in the ground-truth decision space | No | Together with conceptual-variable and modeling decisions, partitions the 536 decision references. |
Ground-truth modeling decisionsblade-ground-truth-modeling | 172 | statistical-model and model-formula decisions in the ground-truth decision space | No | Together with conceptual-variable and transform decisions, partitions the 536 decision references. |
The BLADE track requiring a conceptual-variable specification, executable data-transformation function, and statistical-model function for each open-ended research question and dataset.
2 evaluation run(s)
The BLADE track for selecting the most or least justifiable conceptual-variable and data-transformation decisions for a research question and dataset.
1 evaluation run(s)
Scientific Task Atlas
partial for arXiv v3. BLADE is cross-domain; four source questions are biological or ecological.
| Scientific task | Coverage | Count | Mapping | Evidence |
|---|---|---|---|---|
| End-to-end computational analysis | explicitly-in-scope | 12 problems paired real-world research questions and datasets used as the source units for BLADE | official-track high confidence | blade-evidence-current-countsCoverage is suite-wide; only four source questions are in life science. |
| Domain | Coverage | Count | Interpretation |
|---|---|---|---|
| Life science | observed | 4 | Registry mapping of four explicitly biological or ecological source questions in paper Table 3: AMTL, Crofoot, Panda_nuts, and Fish. BLADE itself is cross-domain. |
| Protein design | not-in-scope | 0 | None of the twelve official research questions concerns protein design or optimization. |
| Protein-protein binding | not-in-scope | 0 | None of the twelve official research questions concerns protein-protein binding. |
| Protein-ligand binding | not-in-scope | 0 | None of the twelve official research questions concerns protein-ligand binding. |
| Genomics | not-in-scope | 0 | The official question inventory contains no genomics or omics-analysis task. |
| Transcriptomics | not-in-scope | 0 | The official question inventory contains no transcriptomics task. |
| Single-cell | not-in-scope | 0 | The official question inventory contains no single-cell task. |
Evaluation registry
A setting change—scope, prompt, tools, budget, grader, or repeats—creates a separate run. Charts never cross a comparability group.
Evaluation run
From BLADE: Benchmarking Language Model Agents for Data-Driven Science
Evaluated models / systems: Claude 3.5 Sonnet (claude-3-5-sonnet-20240620; BLADE), CodeLlama Instruct 7B (BLADE), DeepSeek-Coder Instruct 6.7B (BLADE), Gemini 1.5 Pro (BLADE snapshot not reported), GPT-3.5 Turbo (BLADE snapshot not reported), GPT-4o (BLADE snapshot not reported), Llama 3 70B (BLADE snapshot not reported), Llama 3 8B (BLADE snapshot not reported), Mixtral 8x22B (BLADE snapshot not reported)
| Metric | Kind / baseline | Unit | Aggregation | Threshold / tolerance |
|---|---|---|---|---|
| Average precision | absolute | percent | mean of per-run precision followed by macro averaging across datasets for each decision type | GPT-4o semantic matching plus transformation value/graph matching |
| Coverage@10 | absolute | percent | ground-truth coverage of the union of ten sampled runs followed by macro averaging across datasets for each decision type | run errors count as zero-coverage generations |
| Decision-type weighted F1-score | absolute | percent | weighted mean of per-decision-type harmonic means of average precision and coverage@10 | 1000-run bootstrap mean and 95% interval |
| Model | Metric | Value | n |
|---|---|---|---|
| CodeLlama Instruct 7B (BLADE) | Decision-type weighted F1-score | 16.8 percent Forty one-turn generations per dataset; Table 2. | 12 |
| DeepSeek-Coder Instruct 6.7B (BLADE) | Decision-type weighted F1-score | 33.9 percent Forty one-turn generations per dataset; Table 2. | 12 |
| Llama 3 8B (BLADE snapshot not reported) | Decision-type weighted F1-score | 29.6 percent Forty one-turn generations per dataset; exact served snapshot not reported. | 12 |
| Llama 3 70B (BLADE snapshot not reported) | Decision-type weighted F1-score | 36.3 percent Forty one-turn generations per dataset; exact served snapshot not reported. | 12 |
| Mixtral 8x22B (BLADE snapshot not reported) | Decision-type weighted F1-score | 40.1 percent Forty one-turn generations per dataset; exact served snapshot not reported. | 12 |
| GPT-3.5 Turbo (BLADE snapshot not reported) | Decision-type weighted F1-score | 30.5 percent Forty one-turn generations per dataset; dated endpoint snapshot not reported. | 12 |
| GPT-4o (BLADE snapshot not reported) | Decision-type weighted F1-score | 41.7 percent Forty one-turn generations per dataset; dated endpoint snapshot not reported. | 12 |
| Gemini 1.5 Pro (BLADE snapshot not reported) | Decision-type weighted F1-score | 41.1 percent Forty one-turn generations per dataset; immutable endpoint snapshot not reported. | 12 |
| Claude 3.5 Sonnet (claude-3-5-sonnet-20240620; BLADE) | Decision-type weighted F1-score | 43.9 percent Forty one-turn generations per dataset; Table 2. | 12 |
Evaluation run
From BLADE: Benchmarking Language Model Agents for Data-Driven Science
Evaluated models / systems: Claude 3.5 Sonnet (claude-3-5-sonnet-20240620; BLADE), Gemini 1.5 Pro (BLADE snapshot not reported), GPT-3.5 Turbo (BLADE snapshot not reported), GPT-4o (BLADE snapshot not reported), Mixtral 8x22B (BLADE snapshot not reported)
| Metric | Kind / baseline | Unit | Aggregation | Threshold / tolerance |
|---|---|---|---|---|
| Average precision | absolute | percent | mean of per-run precision followed by macro averaging across datasets for each decision type | GPT-4o semantic matching plus transformation value/graph matching |
| Coverage@10 | absolute | percent | ground-truth coverage of the union of ten sampled runs followed by macro averaging across datasets for each decision type | run errors count as zero-coverage generations |
| Decision-type weighted F1-score | absolute | percent | weighted mean of per-decision-type harmonic means of average precision and coverage@10 | 1000-run bootstrap mean and 95% interval |
| Model | Metric | Value | n |
|---|---|---|---|
| Mixtral 8x22B (BLADE snapshot not reported) | Decision-type weighted F1-score | 40.8 percent Twenty ReAct trajectories per dataset; exact served snapshot not reported. | 12 |
| GPT-3.5 Turbo (BLADE snapshot not reported) | Decision-type weighted F1-score | 37.2 percent Twenty ReAct trajectories per dataset; dated endpoint snapshot not reported. | 12 |
| GPT-4o (BLADE snapshot not reported) | Decision-type weighted F1-score | 44.8 percent Twenty ReAct trajectories per dataset; dated endpoint snapshot not reported. | 12 |
| Gemini 1.5 Pro (BLADE snapshot not reported) | Decision-type weighted F1-score | 40.1 percent Twenty ReAct trajectories per dataset; immutable endpoint snapshot not reported. | 12 |
| Claude 3.5 Sonnet (claude-3-5-sonnet-20240620; BLADE) | Decision-type weighted F1-score | 43.1 percent Twenty ReAct trajectories per dataset; Table 2. | 12 |
Evaluation run
From BLADE: Benchmarking Language Model Agents for Data-Driven Science
Evaluated models / systems: Claude 3.5 Sonnet (claude-3-5-sonnet-20240620; BLADE), CodeLlama Instruct 7B (BLADE), DeepSeek-Coder Instruct 6.7B (BLADE), Gemini 1.5 Pro (BLADE snapshot not reported), GPT-3.5 Turbo (BLADE snapshot not reported), GPT-4o (BLADE snapshot not reported), Llama 3 70B (BLADE snapshot not reported), Llama 3 8B (BLADE snapshot not reported), Mixtral 8x22B (BLADE snapshot not reported)
| Metric | Kind / baseline | Unit | Aggregation | Threshold / tolerance |
|---|---|---|---|---|
| Accuracy | absolute | percent | correct choices divided by all 188 MCQs | exact option identifier |
No numeric result rows are published yet; the verified protocol remains useful.
blade-creator-paper · blade-arxiv-v3-generation-one-turn-40-temperature-08
| Model | Value | Comparability group |
|---|---|---|
| CodeLlama Instruct 7B (BLADE) | 16.8 | blade-arxiv-v3-generation-one-turn-40-temperature-08 |
| DeepSeek-Coder Instruct 6.7B (BLADE) | 33.9 | blade-arxiv-v3-generation-one-turn-40-temperature-08 |
| Llama 3 8B (BLADE snapshot not reported) | 29.6 | blade-arxiv-v3-generation-one-turn-40-temperature-08 |
| Llama 3 70B (BLADE snapshot not reported) | 36.3 | blade-arxiv-v3-generation-one-turn-40-temperature-08 |
| Mixtral 8x22B (BLADE snapshot not reported) | 40.1 | blade-arxiv-v3-generation-one-turn-40-temperature-08 |
| GPT-3.5 Turbo (BLADE snapshot not reported) | 30.5 | blade-arxiv-v3-generation-one-turn-40-temperature-08 |
| GPT-4o (BLADE snapshot not reported) | 41.7 | blade-arxiv-v3-generation-one-turn-40-temperature-08 |
| Gemini 1.5 Pro (BLADE snapshot not reported) | 41.1 | blade-arxiv-v3-generation-one-turn-40-temperature-08 |
| Claude 3.5 Sonnet (claude-3-5-sonnet-20240620; BLADE) | 43.9 | blade-arxiv-v3-generation-one-turn-40-temperature-08 |
blade-creator-react · blade-arxiv-v3-generation-react-20-temperature-08
| Model | Value | Comparability group |
|---|---|---|
| Mixtral 8x22B (BLADE snapshot not reported) | 40.8 | blade-arxiv-v3-generation-react-20-temperature-08 |
| GPT-3.5 Turbo (BLADE snapshot not reported) | 37.2 | blade-arxiv-v3-generation-react-20-temperature-08 |
| GPT-4o (BLADE snapshot not reported) | 44.8 | blade-arxiv-v3-generation-react-20-temperature-08 |
| Gemini 1.5 Pro (BLADE snapshot not reported) | 40.1 | blade-arxiv-v3-generation-react-20-temperature-08 |
| Claude 3.5 Sonnet (claude-3-5-sonnet-20240620; BLADE) | 43.1 | blade-arxiv-v3-generation-react-20-temperature-08 |
Source locators remain visible; expand an item to inspect the exact Registry fields it supports.
/name/aliases/summary/organizations/release_date/kind/capabilities/modalities/task_formats/latest_version/task_counts/task_counts/total/task_counts/basis/task_counts/subsets/versions/0/task_counts/versions/0/formal_tracks/scientific_task_classification/entries/0/domains/coverage_notes/access/biosafety_notes/access/level/access/tasks/access/artifacts/access/grader/access/license/resources/implementations/task_counts/subsets/0/count/task_counts/subsets/1/count/task_counts/subsets/2/count