Scientific task · Scientific workflow

Scientific evidence interpretation

Interpret or reconcile scientific text

科学证据解读

Experimental systemScientific workflow

Definition and search aliases

Permanent ID
scientific-evidence-interpretation
Aliases
figure interpretation, table interpretation, literature synthesis
Deprecated aliases
None
Hierarchy
Leaf task under Scientific workflow and systems analysis

Coverage

2 benchmark families cover this task

trackcomplete

BLADE Decision-Discrimination MCQ

The BLADE track for selecting the most or least justifiable conceptual-variable and data-transformation decisions for a research question and dataset.

Scientific evidence interpretation
suitepartial

LAB-Bench

A practical biology-research suite of 2,457 multiple-choice questions across eight broad categories and 31 versioned task files, with public and private contamination-monitoring splits.

Scientific evidence interpretation
trackcomplete

LAB-Bench FigQA

Multiple-choice interpretation and multi-element reasoning over scientific figures shown without captions or paper context.

Scientific evidence interpretation
trackcomplete

LAB-Bench LitQA2

Literature-retrieval questions whose answers require findings in full research papers rather than titles or abstracts.

Scientific evidence interpretation
trackcomplete

LAB-Bench SuppQA

Retrieval and interpretation questions answerable from paper supplementary text or PDF tables.

Scientific evidence interpretation
trackcomplete

LAB-Bench TableQA

Lookup, calculation, and reasoning questions over table images extracted from scientific papers.

Scientific evidence interpretation

Evidence-backed count claims

Each row keeps its original unit and basis. Rows with different units or overlapping mappings are never added.

BenchmarkMapped taskCoverageCountVersionEvidence
Anthropic Scientific Figure Interpretation Eval
root: anthropic-key-life-sciences-evals
Scientific evidence interpretation
official-track · high
explicitly-in-scopeNot reported
private scientific figure interpretation tasks
reported-2026-01-11
as of 2026-01-11
anthropic-scientific-figure-evidence
BLADE Decision-Discrimination MCQ
root: blade
Scientific evidence interpretation
official-track · high
explicitly-in-scope188 questions
individual multiple-choice decision-discrimination questions
arXiv v3blade-mcq-evidence-counts
LAB-Bench
root: lab-bench
Scientific evidence interpretation
official-track · high
explicitly-in-scopeNot reported
FigQA, LitQA2, SuppQA, and TableQA questions.
repository-998a8e0lab-bench-evidence-paper
LAB-Bench FigQA
root: lab-bench
Scientific evidence interpretation
official-track · high
explicitly-in-scope226 questions
questions across public and private splits
repository-998a8e0lab-bench-figqa-evidence-paper
LAB-Bench LitQA2
root: lab-bench
Scientific evidence interpretation
official-track · high
explicitly-in-scope248 questions
questions across public and private splits
repository-998a8e0lab-bench-litqa2-evidence-paper
LAB-Bench SuppQA
root: lab-bench
Scientific evidence interpretation
official-track · high
explicitly-in-scope102 questions
questions across public and private splits
repository-998a8e0lab-bench-suppqa-evidence-paper
LAB-Bench TableQA
root: lab-bench
Scientific evidence interpretation
official-track · high
explicitly-in-scope305 questions
questions across public and private splits
repository-998a8e0lab-bench-tableqa-evidence-paper

Official evaluations connected to these benchmarks

Runs are included only for benchmark records mapped here (and formal child tracks when a mapped suite is the root). A task mapping does not imply that every run isolates this task.