agentic-eval · audited-with-caveats · verified 2026-07-28

scBench

Agentic evaluation suite for data-grounded single-cell analysis across diverse sequencing technologies and workflow stages.

Audited with caveats: 1 field(s) are marked provisional or conflicted. Warnings are shown next to affected values and these claims are excluded from unqualified comparisons.

Benchmark definition

What is counted

Version
repository-195-evaluations
Total
195 (Official repository README at commit 0bc34032bfa402dad29fc40b4cf10ea8fc03193e)
Task formats
natural-language data-analysis task with structured JSON output
Capabilities
Data analysisCodingTool useScientific reasoning
Modalities
TextRaw omicsCode

Version history

VersionStatusRelease / as-ofTotalFormal tracks
initial-release
scbench-initial-release-version
superseded2026-02-09394 (Explicit overall benchmark total in the abstract and Table 1)None registered
repository-195-evaluations
scbench-repository-195-version
current2026-06-10195 (Official repository README at commit 0bc34032bfa402dad29fc40b4cf10ea8fc03193e)None registered

Scientific Task Atlas

Scientific task classification

partial for repository-195-evaluations. No Scientific Task claim passed independent high-confidence verification; task mapping remains pending a targeted official-source audit.

The source names only a broad direction; no more specific leaf task can be assigned without inference.

Relationship registry

How works use this benchmark

Partial claims, non-evaluation uses, and third-party summaries stay visible without entering model comparisons.

Partial evaluation claims

evaluation

scbench-evaluating-ai-agents-on-single-cell-rna-seq-an-scbench-2-use

Partialunknown

Work: scBench: Evaluating AI Agents on Single-Cell RNA-seq Analysis · source version scbench-evaluating-ai-agents-on-single-cell-rna-seq-an-arxiv-v1

Selection
not reported
Metrics
Not reported / not applicable
Linked runs
None

Not reported / unresolved: Exact API snapshots and model release dates are not reported.; Shot count and model reasoning settings are not reported.; No token budget is reported.; Latency confidence intervals are not reported.; benchmark version; realized n/scope; metric; numeric result; prompt and tools; grader and repeats

Owner-reviewed conservative publication: the creator evaluation is retained only as a partial relationship; conflicted settings and outcomes are omitted pending manual reconciliation.

Evidence
  • table: Table 2
    Supports: /relation_type
  • table: Table 2
    Supports: /benchmark_id
  • table: Table 2
    Supports: /model_ids
  • table: Table 2
    Supports: /model_ids
  • table: Table 2
    Supports: /model_ids
  • table: Table 2
    Supports: /model_ids
  • table: Table 2
    Supports: /model_ids
  • table: Table 2
    Supports: /model_ids
  • table: Table 2
    Supports: /model_ids
  • table: Table 2
    Supports: /model_ids

evaluation

system-card-claude-opus-5-scbench-3-use

Partialunknown · n=195

Work: System Card: Claude Opus 5 · source version system-card-claude-opus-5-2026-07-24

Selection
not reported
Metrics
Score
Linked runs
None

Not reported / unresolved: Exact benchmark version is not reported.; Per-workflow problem counts are not reported.; Prompt, shots, reasoning settings, budget, seed, repeats, grader, and human review are not reported.; Score definition and aggregation are not reported.; benchmark version

AI-assisted double-pass extraction; values are limited to independently supported claims.

Evidence
  • section: Section 8.17.2 LatchBio Bioinformatics
    Supports: /relation_type
  • section: Section 8.17.2 LatchBio Bioinformatics
    Supports: /benchmark_id
  • section: Section 8.17.2 LatchBio Bioinformatics
    Supports: /scope
  • section: Section 8.17.2 LatchBio Bioinformatics
    Supports: /scope
  • figure: Figure 8.17.6.A, LatchBio Bioinformatics panel
    Supports: /metric_labels
  • section: Section 8.17.2, SingleCellBench
    Supports: /model_ids
  • section: Section 8.17.2, SingleCellBench
    Supports: /model_ids
  • section: Section 8.17.2, SingleCellBench
    Supports: /model_ids
  • section: Section 8.17.2, SingleCellBench
    Supports: /model_ids

Creation, training, validation, or model-selection uses

benchmark creation

scbench-evaluating-ai-agents-on-single-cell-rna-seq-an-scbench-1-use

Non-evaluationunknown

Work: scBench: Evaluating AI Agents on Single-Cell RNA-seq Analysis · source version scbench-evaluating-ai-agents-on-single-cell-rna-seq-an-arxiv-v1

Selection
not applicable
Models
Not reported / not applicable
Metrics
Not reported / not applicable
Linked runs
None

Not reported / unresolved: No explicit benchmark version is reported.; The paper identifies the official repository but does not report its software license.

AI-assisted double-pass extraction; values are limited to independently supported claims.

Evidence
  • section: Abstract
    Supports: /relation_type
  • section: Abstract
    Supports: /benchmark_id

Evaluation registry

Works and run settings

A setting change—scope, prompt, tools, budget, grader, or repeats—creates a separate run. Charts never cross a comparability group.

No normalized evaluation run is published yet. Creator evidence is still attached below.

Evidence and change history

Source locators remain visible; expand an item to inspect the exact Registry fields it supports.

scBench: Evaluating AI Agents on Single-Cell RNA-seq Analysis · section: Abstract; Sections 4.1–5 · Supports 12 fields

Open source →

  • /name
  • /aliases
  • /summary
  • /kind
  • /organizations
  • /release_date
  • /domains
  • /capabilities
  • /modalities
  • /task_formats
  • /access/level
  • /access/license
scBench: Evaluating AI Agents on Single-Cell RNA-seq Analysis · table: Table 1 · Supports 1 field

Open source →

  • /versions/0/task_counts
scBench: Evaluating AI Agents on Single-Cell RNA-seq Analysis · section: Section 5 · Supports 2 fields

Open source →

  • /resources
  • /implementations
scBench: Evaluating AI Agents on Single-Cell RNA-seq Analysis · page: Title page · Supports 1 field

Open source →

  • /resources/0
scBench: Evaluating AI Agents on Single-Cell RNA-seq Analysis · page: Title page · Supports 2 fields

Open source →

  • /latest_version
  • /versions/0
scBench: Evaluating AI Agents on Single-Cell RNA-seq Analysis · table: Table 7 (continued) · Supports 1 field

Open source →

  • /versions/0/task_counts/subsets
scbench-official-repository-resource · repository-path: README.md#benchmark-structure (Pinned official README reports 195 evaluations and describes the current withheld benchmark snapshot.) · Supports 5 fields

Open source →

  • /latest_version
  • /task_counts/total
  • /task_counts/basis
  • /task_counts/subsets
  • /versions/1
scbench-official-repository-resource · repository-path: README.md#license (Pinned official README states the repository license.) · Supports 1 field

Open source →

  • /access/license
scbench-anthropic-opus-5-system-card-resource · section: Section 8.17.2, LatchBio Bioinformatics (Provider-used evaluation label; not represented as a creator-preferred benchmark name.) · Supports 1 field

Open source →

  • /aliases/0

Unresolved field claims

  • /versions/0/task_counts/subsetsConflicted · high — The owner approved the independently supported root total while all conflicted inventory subcounts were excluded from publication.
    Evidence: scbench-automated-count-conflict-evidence

View source-level modification history on GitHub →