track · audited · verified 2026-07-21
BLADE Decision-Discrimination MCQ
The BLADE track for selecting the most or least justifiable conceptual-variable and data-transformation decisions for a research question and dataset.
Benchmark definition
What is counted
- Version
- arXiv v3
- Total
- 188 (individual multiple-choice decision-discrimination questions)
- Task formats
- multiple choice
- Capabilities
KnowledgeClassificationData analysisScientific reasoning
Version history
| Version | Status | Release / as-of | Total | Formal tracks |
|---|
arXiv v3
blade-mcq-arxiv-v3 | current | 2025-11-10 | 188 (individual multiple-choice decision-discrimination questions) | None registered |
Tracks and subsets
| ID | Count | Basis | Partition? | Notes |
|---|
Conceptual-variable MCQs
blade-mcq-conceptual-items | 20 | individual questions asking which conceptual variable is most or least justifiable | Exclusive & exhaustive | Mutually exclusive with transformation MCQs. |
Transformation MCQs
blade-mcq-transform-items | 168 | individual questions asking which transformation is most or least justifiable | Exclusive & exhaustive | Mutually exclusive with conceptual-variable MCQs. |
Source datasets represented
blade-mcq-source-datasets | 11 | official dataset directories containing mcq_dataset.json at the pinned commit | No | A source-dataset count, not a partition of questions; the soccer dataset is part of analysis generation but has no MCQ file. |
Scientific Task Atlas
Scientific task classification
complete for arXiv v3. Formal decision-discrimination track.
| Scientific task | Coverage | Count | Mapping | Evidence |
|---|
| Scientific evidence interpretation | explicitly-in-scope | 188 questions individual multiple-choice decision-discrimination questions | official-track high confidence | blade-mcq-evidence-counts Questions test defensibility of scientific analysis decisions. |
Scientific coverage notes
| Domain | Coverage | Count | Interpretation |
|---|
| Life science | observed | Not reported | Biological source datasets are present, but the paper does not publish an MCQ count by scientific domain. |
| Protein design | not-in-scope | 0 | No official BLADE source research question concerns protein design. |
| Protein-protein binding | not-in-scope | 0 | No official BLADE source research question concerns protein-protein binding. |
| Protein-ligand binding | not-in-scope | 0 | No official BLADE source research question concerns protein-ligand binding. |
Evaluation registry
Works and run settings
A setting change—scope, prompt, tools, budget, grader, or repeats—creates a separate run. Charts never cross a comparability group.
blade-arxiv-v3-mcq-zero-shot-temperature-0varXiv v3
Evaluated models / systems: Claude 3.5 Sonnet (claude-3-5-sonnet-20240620; BLADE), CodeLlama Instruct 7B (BLADE), DeepSeek-Coder Instruct 6.7B (BLADE), Gemini 1.5 Pro (BLADE snapshot not reported), GPT-3.5 Turbo (BLADE snapshot not reported), GPT-4o (BLADE snapshot not reported), Llama 3 70B (BLADE snapshot not reported), Llama 3 8B (BLADE snapshot not reported), Mixtral 8x22B (BLADE snapshot not reported)
Scopefull · n=188
Shotszero-shot
Turnssingle-turn
System prompt publicYes
Reasoning / effortdirect multiple-choice completion
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNo
External toolsnone
Token budgetNot reported
Time / cost budgetNot reported
Temperature0
SeedNot reported
Repeats1
Graderdeterministic exact option match · human review: no
Statisticsquestion-weighted accuracy with 95% confidence intervals
ContaminationNot reported
Metrics, results, and full protocol
Metrics
| Metric | Kind / baseline | Unit | Aggregation | Threshold / tolerance |
|---|
| Accuracy | absolute | percent | correct choices divided by all 188 MCQs | exact option identifier |
No numeric result rows are published yet; the verified protocol remains useful.
Evidence
- section: arXiv v3 §§4.1 and 6, Figure 3, and Appendix A.6 Figure 19 (Reports 188 MCQs, nine evaluated model families, temperature zero, accuracy with 95% intervals, and the public prompt.) — supports /benchmark_version, /model_ids, /scope, /protocol/temperature, /protocol/statistical, /metrics, /comparability_group
- repository-path: run_mcq.py, run_scripts/sh_run_mcq.sh, blade_bench/baselines/lm/mcq.py, blade_bench/eval/datamodel/run_mcq.py, and blade_bench/conf/llm_config.yml at commit 6118fa8d5007b91aa8c91c518182db82446a4547 (Confirms one direct prompt per item, no tool loop, exact-choice grading, complete public prompt, 20/168 question components, and available exact model strings.) — supports /scope, /protocol/shots, /protocol/turns, /protocol/system_prompt_public, /protocol/reasoning, /protocol/tools, /protocol/repeats, /protocol/grader, /model_ids
Evidence and change history
Source locators remain visible; expand an item to inspect the exact Registry fields it supports.
blade-mcq-paper-resource · section: arXiv v3 §4.1, §6, Figure 3, and Appendix A.6 (Defines the MCQ task and reports 188 questions: 20 conceptual-variable and 168 transformation questions.) · Supports 16 fields
Open source →
/name/summary/kind/parent_id/organizations/release_date/latest_version/capabilities/modalities/task_formats/task_counts/task_counts/total/task_counts/basis/task_counts/subsets/versions/0/task_counts/scientific_task_classification/entries/0
blade-mcq-repository-resource · repository-path: blade_bench/datasets/*/mcq_dataset.json, blade_bench/baselines/lm/mcq.py, run_mcq.py, and blade_bench/eval/datamodel/run_mcq.py at commit 6118fa8d5007b91aa8c91c518182db82446a4547 (Public files reproduce 20/168/188 across eleven source datasets and implement a single prompt call plus exact-choice accuracy.) · Supports 6 fields
Open source →
/task_counts/access/access/level/access/license/resources/implementations
blade-mcq-paper-resource · table: Table 3 (The source-question inventory includes biological/ecological datasets but no protein, binding, or omics question; MCQ-by-domain counts are not reported.) · Supports 3 fields
Open source →
/domains/coverage_notes/access/biosafety_notes
View source-level modification history on GitHub →