agentic-eval · audited-with-caveats · verified 2026-07-28
scBench
Agentic evaluation suite for data-grounded single-cell analysis across diverse sequencing technologies and workflow stages.
Audited with caveats: 1 field(s) are marked provisional or conflicted. Warnings are shown next to affected values and these claims are excluded from unqualified comparisons.
Benchmark definition
What is counted
- Version
- repository-195-evaluations
- Total
- 195 (Official repository README at commit 0bc34032bfa402dad29fc40b4cf10ea8fc03193e)
- Task formats
- natural-language data-analysis task with structured JSON output
- Capabilities
Data analysisCodingTool useScientific reasoning
- Modalities
TextRaw omicsCode
Version history
| Version | Status | Release / as-of | Total | Formal tracks |
|---|
initial-release
scbench-initial-release-version | superseded | 2026-02-09 | 394 (Explicit overall benchmark total in the abstract and Table 1) | None registered |
repository-195-evaluations
scbench-repository-195-version | current | 2026-06-10 | 195 (Official repository README at commit 0bc34032bfa402dad29fc40b4cf10ea8fc03193e) | None registered |
Scientific Task Atlas
Scientific task classification
partial for repository-195-evaluations. No Scientific Task claim passed independent high-confidence verification; task mapping remains pending a targeted official-source audit.
The source names only a broad direction; no more specific leaf task can be assigned without inference.
Relationship registry
How works use this benchmark
Partial claims, non-evaluation uses, and third-party summaries stay visible without entering model comparisons.
Partial evaluation claims
evaluation
scbench-evaluating-ai-agents-on-single-cell-rna-seq-an-scbench-2-use
Partialunknown
Work: scBench: Evaluating AI Agents on Single-Cell RNA-seq Analysis · source version scbench-evaluating-ai-agents-on-single-cell-rna-seq-an-arxiv-v1
- Selection
- not reported
- Metrics
- Not reported / not applicable
- Linked runs
- None
Not reported / unresolved: Exact API snapshots and model release dates are not reported.; Shot count and model reasoning settings are not reported.; No token budget is reported.; Latency confidence intervals are not reported.; benchmark version; realized n/scope; metric; numeric result; prompt and tools; grader and repeats
Owner-reviewed conservative publication: the creator evaluation is retained only as a partial relationship; conflicted settings and outcomes are omitted pending manual reconciliation.
Evidence
- table: Table 2
Supports: /relation_type - table: Table 2
Supports: /benchmark_id - table: Table 2
Supports: /model_ids - table: Table 2
Supports: /model_ids - table: Table 2
Supports: /model_ids - table: Table 2
Supports: /model_ids - table: Table 2
Supports: /model_ids - table: Table 2
Supports: /model_ids - table: Table 2
Supports: /model_ids - table: Table 2
Supports: /model_ids
evaluation
system-card-claude-opus-5-scbench-3-use
Partialunknown · n=195
Work: System Card: Claude Opus 5 · source version system-card-claude-opus-5-2026-07-24
- Selection
- not reported
- Metrics
- Score
- Linked runs
- None
Not reported / unresolved: Exact benchmark version is not reported.; Per-workflow problem counts are not reported.; Prompt, shots, reasoning settings, budget, seed, repeats, grader, and human review are not reported.; Score definition and aggregation are not reported.; benchmark version
AI-assisted double-pass extraction; values are limited to independently supported claims.
Evidence
- section: Section 8.17.2 LatchBio Bioinformatics
Supports: /relation_type - section: Section 8.17.2 LatchBio Bioinformatics
Supports: /benchmark_id - section: Section 8.17.2 LatchBio Bioinformatics
Supports: /scope - section: Section 8.17.2 LatchBio Bioinformatics
Supports: /scope - figure: Figure 8.17.6.A, LatchBio Bioinformatics panel
Supports: /metric_labels - section: Section 8.17.2, SingleCellBench
Supports: /model_ids - section: Section 8.17.2, SingleCellBench
Supports: /model_ids - section: Section 8.17.2, SingleCellBench
Supports: /model_ids - section: Section 8.17.2, SingleCellBench
Supports: /model_ids
Creation, training, validation, or model-selection uses
benchmark creation
scbench-evaluating-ai-agents-on-single-cell-rna-seq-an-scbench-1-use
Non-evaluationunknown
Work: scBench: Evaluating AI Agents on Single-Cell RNA-seq Analysis · source version scbench-evaluating-ai-agents-on-single-cell-rna-seq-an-arxiv-v1
- Selection
- not applicable
- Models
- Not reported / not applicable
- Metrics
- Not reported / not applicable
- Linked runs
- None
Not reported / unresolved: No explicit benchmark version is reported.; The paper identifies the official repository but does not report its software license.
AI-assisted double-pass extraction; values are limited to independently supported claims.
Evidence
- section: Abstract
Supports: /relation_type - section: Abstract
Supports: /benchmark_id
Evaluation registry
Works and run settings
A setting change—scope, prompt, tools, budget, grader, or repeats—creates a separate run. Charts never cross a comparability group.
No normalized evaluation run is published yet. Creator evidence is still attached below.
Evidence and change history
Source locators remain visible; expand an item to inspect the exact Registry fields it supports.
scBench: Evaluating AI Agents on Single-Cell RNA-seq Analysis · section: Abstract; Sections 4.1–5 · Supports 12 fields
Open source →
/name/aliases/summary/kind/organizations/release_date/domains/capabilities/modalities/task_formats/access/level/access/license
scBench: Evaluating AI Agents on Single-Cell RNA-seq Analysis · table: Table 1 · Supports 1 field
Open source →
scBench: Evaluating AI Agents on Single-Cell RNA-seq Analysis · section: Section 5 · Supports 2 fields
Open source →
/resources/implementations
scBench: Evaluating AI Agents on Single-Cell RNA-seq Analysis · page: Title page · Supports 1 field
Open source →
scBench: Evaluating AI Agents on Single-Cell RNA-seq Analysis · page: Title page · Supports 2 fields
Open source →
/latest_version/versions/0
scBench: Evaluating AI Agents on Single-Cell RNA-seq Analysis · table: Table 7 (continued) · Supports 1 field
Open source →
/versions/0/task_counts/subsets
scbench-official-repository-resource · repository-path: README.md#benchmark-structure (Pinned official README reports 195 evaluations and describes the current withheld benchmark snapshot.) · Supports 5 fields
Open source →
/latest_version/task_counts/total/task_counts/basis/task_counts/subsets/versions/1
scbench-official-repository-resource · repository-path: README.md#license (Pinned official README states the repository license.) · Supports 1 field
Open source →
scbench-anthropic-opus-5-system-card-resource · section: Section 8.17.2, LatchBio Bioinformatics (Provider-used evaluation label; not represented as a creator-preferred benchmark name.) · Supports 1 field
Open source →
Unresolved field claims
/versions/0/task_counts/subsetsConflicted · high — The owner approved the independently supported root total while all conflicted inventory subcounts were excluded from publication.
Evidence: scbench-automated-count-conflict-evidence
View source-level modification history on GitHub →