preprint · benchmark creator

scBench: Evaluating AI Agents on Single-Cell RNA-seq Analysis

LatchBio · 2026-02-09

Relationship layer

Benchmark usage

This table records what the work did with each benchmark before attempting to normalize a run. Partial claims remain visible without being treated as comparable evaluations.

benchmark creation

scbench-evaluating-ai-agents-on-single-cell-rna-seq-an-scbench-1-use

Non-evaluationunknown

Benchmark: scBench · version Not reported

Selection
not applicable
Models
Not reported / not applicable
Metrics
Not reported / not applicable
Linked runs
None

Not reported / unresolved: No explicit benchmark version is reported.; The paper identifies the official repository but does not report its software license.

AI-assisted double-pass extraction; values are limited to independently supported claims.

Evidence
  • section: Abstract
    Supports: /relation_type
  • section: Abstract
    Supports: /benchmark_id

evaluation

scbench-evaluating-ai-agents-on-single-cell-rna-seq-an-scbench-2-use

Partialunknown

Benchmark: scBench · version Not reported

Selection
not reported
Metrics
Not reported / not applicable
Linked runs
None

Not reported / unresolved: Exact API snapshots and model release dates are not reported.; Shot count and model reasoning settings are not reported.; No token budget is reported.; Latency confidence intervals are not reported.; benchmark version; realized n/scope; metric; numeric result; prompt and tools; grader and repeats

Owner-reviewed conservative publication: the creator evaluation is retained only as a partial relationship; conflicted settings and outcomes are omitted pending manual reconciliation.

Evidence
  • table: Table 2
    Supports: /relation_type
  • table: Table 2
    Supports: /benchmark_id
  • table: Table 2
    Supports: /model_ids
  • table: Table 2
    Supports: /model_ids
  • table: Table 2
    Supports: /model_ids
  • table: Table 2
    Supports: /model_ids
  • table: Table 2
    Supports: /model_ids
  • table: Table 2
    Supports: /model_ids
  • table: Table 2
    Supports: /model_ids
  • table: Table 2
    Supports: /model_ids

Normalized evaluation runs

This source has no normalized model run. It may be a creator-only source or a partial/non-evaluation benchmark use.