BLADE Decision-Discrimination MCQ
The BLADE track for selecting the most or least justifiable conceptual-variable and data-transformation decisions for a research question and dataset.
Scientific task · Scientific workflow
Interpret or reconcile scientific text
科学证据解读
scientific-evidence-interpretationCoverage
The BLADE track for selecting the most or least justifiable conceptual-variable and data-transformation decisions for a research question and dataset.
A practical biology-research suite of 2,457 multiple-choice questions across eight broad categories and 31 versioned task files, with public and private contamination-monitoring splits.
Multiple-choice interpretation and multi-element reasoning over scientific figures shown without captions or paper context.
Literature-retrieval questions whose answers require findings in full research papers rather than titles or abstracts.
Retrieval and interpretation questions answerable from paper supplementary text or PDF tables.
Lookup, calculation, and reasoning questions over table images extracted from scientific papers.
Each row keeps its original unit and basis. Rows with different units or overlapping mappings are never added.
| Benchmark | Mapped task | Coverage | Count | Version | Evidence |
|---|---|---|---|---|---|
| Anthropic Scientific Figure Interpretation Eval root: anthropic-key-life-sciences-evals | Scientific evidence interpretation official-track · high | explicitly-in-scope | Not reported private scientific figure interpretation tasks | reported-2026-01-11 as of 2026-01-11 | anthropic-scientific-figure-evidence |
| BLADE Decision-Discrimination MCQ root: blade | Scientific evidence interpretation official-track · high | explicitly-in-scope | 188 questions individual multiple-choice decision-discrimination questions | arXiv v3 | blade-mcq-evidence-counts |
| LAB-Bench root: lab-bench | Scientific evidence interpretation official-track · high | explicitly-in-scope | Not reported FigQA, LitQA2, SuppQA, and TableQA questions. | repository-998a8e0 | lab-bench-evidence-paper |
| LAB-Bench FigQA root: lab-bench | Scientific evidence interpretation official-track · high | explicitly-in-scope | 226 questions questions across public and private splits | repository-998a8e0 | lab-bench-figqa-evidence-paper |
| LAB-Bench LitQA2 root: lab-bench | Scientific evidence interpretation official-track · high | explicitly-in-scope | 248 questions questions across public and private splits | repository-998a8e0 | lab-bench-litqa2-evidence-paper |
| LAB-Bench SuppQA root: lab-bench | Scientific evidence interpretation official-track · high | explicitly-in-scope | 102 questions questions across public and private splits | repository-998a8e0 | lab-bench-suppqa-evidence-paper |
| LAB-Bench TableQA root: lab-bench | Scientific evidence interpretation official-track · high | explicitly-in-scope | 305 questions questions across public and private splits | repository-998a8e0 | lab-bench-tableqa-evidence-paper |
Runs are included only for benchmark records mapped here (and formal child tracks when a mapped suite is the root). A task mapping does not imply that every run isolates this task.