agentic-eval · audited-with-caveats · verified 2026-07-22

SpatialBench

A benchmark of deterministic, verifiable agentic problems derived from real spatial-transcriptomics workflows, testing whether agents can manipulate data and recover key biological results.

+2 more
Audited with caveats: 1 field(s) are marked provisional or conflicted. Warnings are shown next to affected values and these claims are excluded from unqualified comparisons.
Version separation: paper v2 reports 146 problems, while the pinned current repository snapshot contains 159 evaluations. Results and comparability groups never cross these versions. The paper's printed seven task-category counts and five platform counts each sum to 147 rather than its stated 146 total, so those historical partitions remain visibly Conflicted.

Benchmark definition

What is counted

Version
repo-159-5042c4f
Total
159 (evaluations in the official repository snapshot at commit 5042c4f3ee597da1590650c7b894d068ae968e26)
Task formats
containerized spatial-biology analysis problem; deterministic graded agent episode
Capabilities
Data analysisCodingTool useScientific reasoning
Modalities
TextTableFigureRaw omicsImageCode

Version history

VersionStatusRelease / as-ofTotalFormal tracks
paper-v2
spatialbench-paper-v2
superseded2026-01-05146 (problems in the complete arXiv v2 benchmark inventory)None registered
repo-159-5042c4f
spatialbench-repo-159-5042c4f
current2026-06-10159 (evaluations in the official repository snapshot at commit 5042c4f3ee597da1590650c7b894d068ae968e26)None registered

Tracks and subsets

IDCountBasisPartition?Notes
Cell typing
spatialbench-159-cell-typing
45official category_results.json n_evalsExclusive & exhaustiveCategory partition of the 159 evaluations.
Clustering
spatialbench-159-clustering
3official category_results.json n_evalsExclusive & exhaustiveCategory partition of the 159 evaluations.
Differential expression
spatialbench-159-differential-expression
40official category_results.json n_evalsExclusive & exhaustiveCategory partition of the 159 evaluations.
Dimensionality reduction
spatialbench-159-dimensionality-reduction
11official category_results.json n_evalsExclusive & exhaustiveCategory partition of the 159 evaluations.
Normalization
spatialbench-159-normalization
7official category_results.json n_evalsExclusive & exhaustiveCategory partition of the 159 evaluations.
Quality control
spatialbench-159-qc
17official category_results.json n_evalsExclusive & exhaustiveCategory partition of the 159 evaluations.
Spatial analysis
spatialbench-159-spatial-analysis
36official category_results.json n_evalsExclusive & exhaustiveCategory partition of the 159 evaluations.
AtlasXOmics
spatialbench-159-atlasxomics
32official platform_results.json n_evalsExclusive & exhaustivePlatform partition of the same 159 evaluations; never added to category counts.
Curio
spatialbench-159-curio
34official platform_results.json n_evalsExclusive & exhaustivePlatform partition of the same 159 evaluations; never added to category counts.
MERFISH / Vizgen
spatialbench-159-merfish
33official platform_results.json merfish n_evalsExclusive & exhaustiveThe result artifact uses merfish while the README names the vendor/platform as Vizgen.
Visium
spatialbench-159-visium
30official platform_results.json n_evalsExclusive & exhaustivePlatform partition of the same 159 evaluations; never added to category counts.
Xenium
spatialbench-159-xenium
30official platform_results.json n_evalsExclusive & exhaustivePlatform partition of the same 159 evaluations; never added to category counts.

Scientific Task Atlas

Scientific task classification

partial for repo-159-5042c4f · as of 2026-06-10. Four official categories map directly to existing leaf tasks. Dimensionality reduction, normalization, and QC remain visible as benchmark subsets without inventing new task terms in this intake.

Scientific taskCoverageCountMappingEvidence
Cell-type annotationexplicitly-in-scope45 problems
official category_results.json n_evals
official-taxonomy
high confidence
spatialbench-evidence-current-counts
Official Cell Typing category.
Cell-state clusteringexplicitly-in-scope3 problems
official category_results.json n_evals
official-taxonomy
high confidence
spatialbench-evidence-current-counts
Official Clustering category.
Differential expression analysisexplicitly-in-scope40 problems
official category_results.json n_evals
official-taxonomy
high confidence
spatialbench-evidence-current-counts
Official Differential Expression category.
Spatial omics analysisexplicitly-in-scope36 problems
official category_results.json n_evals
official-taxonomy
high confidence
spatialbench-evidence-current-counts
Official Spatial Analysis category.

Scientific coverage notes

DomainCoverageCountInterpretation
Spatial omicsexplicitly-in-scope159Every current evaluation is drawn from a spatial transcriptomics workflow.
Transcriptomicsexplicitly-in-scope159The official release describes all five platforms as spatial transcriptomics technologies.
Single-cellobservedNot reportedCell typing and clustering are explicit categories, but a standalone single-cell-only count is not published.

Relationship registry

How works use this benchmark

Partial claims, non-evaluation uses, and third-party summaries stay visible without entering model comparisons.

Normalized evaluations

evaluation

spatialbench-preprint-evaluation

Normalizedfull · n=146

Work: SpatialBench: Can Agents Analyze Real-World Spatial Biology Data? · source version spatialbench-preprint-v2

Selection
not applicable · all 146 paper-v2 evaluations
Metrics
Accuracy, Steps, Latency, Cost

Base, Claude Code, and Latch harnesses are normalized as separate runs and comparability groups.

Evidence
  • table: arXiv v2 Tables 1 and 4; §§2.2, 2.5, 3.5–3.7 (Model and harness results over all 146 evaluations.)
    Supports: /benchmark_version, /relation_type, /status, /model_ids, /scope, /metric_labels, /evaluation_run_ids, /notes

evaluation

spatialbench-repository-evaluation

Normalizedfull · n=159

Work: SpatialBench 159-evaluation repository snapshot · source version spatialbench-repository-release-2026-06-10

Selection
not applicable · all 159 evaluations
Metrics
Accuracy, Cost, Duration

mini-swe-agent, Claude Code, Pi, and OpenAI Codex are separate normalized runs and comparability groups.

Evidence
  • repository-path: METHODS.md and results/model_results.csv at commit 5042c4f3ee597da1590650c7b894d068ae968e26 (Full 159-evaluation methods and harness-specific result rows.)
    Supports: /benchmark_version, /relation_type, /status, /model_ids, /scope, /metric_labels, /evaluation_run_ids, /notes

Partial evaluation claims

evaluation

system-card-claude-opus-5-spatialbench-2-use

Partialunknown · n=115

Work: System Card: Claude Opus 5 · source version system-card-claude-opus-5-2026-07-24

Selection
not reported · Externally validated problems
Metrics
Score
Linked runs
None

Not reported / unresolved: The source does not identify an official benchmark version or artifact revision for the Verified qualifier.; Prompt, shots, reasoning settings, budget, seed, repeats, grader, and human review are not reported.; Score definition and aggregation are not reported.; benchmark version

AI-assisted double-pass extraction; values are limited to independently supported claims.

Evidence
  • section: Section 8.17.2 LatchBio Bioinformatics
    Supports: /relation_type
  • section: Section 8.17.2 LatchBio Bioinformatics
    Supports: /benchmark_id
  • section: Section 8.17.2 LatchBio Bioinformatics
    Supports: /scope
  • section: Section 8.17.2 LatchBio Bioinformatics
    Supports: /scope
  • section: Section 8.17.2 LatchBio Bioinformatics
    Supports: /scope
  • section: Section 8.17.2 LatchBio Bioinformatics
    Supports: /scope
  • section: Section 8.17.2 LatchBio Bioinformatics
    Supports: /metric_labels
  • section: Section 8.17.2, SpatialBench Verified
    Supports: /model_ids
  • section: Section 8.17.2, SpatialBench Verified
    Supports: /model_ids
  • section: Section 8.17.2, SpatialBench Verified
    Supports: /model_ids
  • section: Section 8.17.2, SpatialBench Verified
    Supports: /model_ids

Creation, training, validation, or model-selection uses

benchmark creation

spatialbench-preprint-creation

Non-evaluationunknown

Work: SpatialBench: Can Agents Analyze Real-World Spatial Biology Data? · source version spatialbench-preprint-v2

Selection
not applicable
Models
Not reported / not applicable
Metrics
Not reported / not applicable
Linked runs
None

Creator preprint defining the 146-problem paper-v2 benchmark snapshot.

Evidence
  • page: arXiv v2 abstract and §§2.1, 3.1–3.4 (Introduces and constructs SpatialBench.)
    Supports: /benchmark_version, /relation_type, /status, /scope, /notes

benchmark creation

spatialbench-repository-creation

Non-evaluationunknown

Work: SpatialBench 159-evaluation repository snapshot · source version spatialbench-repository-release-2026-06-10

Selection
not applicable
Models
Not reported / not applicable
Metrics
Not reported / not applicable
Linked runs
None

Creator-maintained revised 159-evaluation benchmark snapshot fixed to the registered commit.

Evidence
  • repository-path: README.md and CHANGELOG.md at commit 5042c4f3ee597da1590650c7b894d068ae968e26 (Defines the revised 159-evaluation version.)
    Supports: /benchmark_version, /relation_type, /status, /scope, /notes

External result summaries

external result summary

anthropic-spatialbench-external-summary

External summaryfull · n=146

Work: Advancing Claude in healthcare and the life sciences · source version anthropic-healthcare-life-sciences-2026-01-11

Selection
not applicable · all 146 paper-v2 problems described in the chart
Metrics
Accuracy
Linked runs
None

Not reported / unresolved: Anthropic did not conduct or claim an independent rerun; harness settings are inherited from the cited LatchBio source rather than reported as an Anthropic protocol

The exact rounded chart values are explicitly attributed to LatchBio SpatialBench. This relation is a third-party result summary and is never counted as an Anthropic self-evaluation.

Evidence
  • figure: SpatialBench: Spatial biology analysis by LatchBio (Caption says Source: LatchBio SpatialBench and 146 verifiable problems across five platforms and seven task categories.)
    Supports: /benchmark_version, /relation_type, /status, /model_ids, /scope, /metric_labels, /evaluation_run_ids, /reporting_gaps, /notes

Evaluation registry

Works and run settings

A setting change—scope, prompt, tools, budget, grader, or repeats—creates a separate run. Charts never cross a comparability group.

Evaluation run

spatialbench-paper-v2-base

From SpatialBench: Can Agents Analyze Real-World Spatial Biology Data?

spatialbench-paper-v2-basevpaper-v2

Evaluated models / systems: Claude Opus 4.5, Claude Sonnet 4.5, Gemini 2.5 Pro, GPT-5.1, GPT-5.2, Grok-4, Grok-4.1 (SpatialBench paper label)

Scopefull · n=146
ShotsNot reported
Turnsmulti-turn agent
System prompt publicNo
Reasoning / effortslightly modified Mini-SWE-Bench base harness
BrowserNot reported
InternetNot reported
DatabasesNot reported
Code executionYes
ContainerYes
External toolsinteractive compute and local workspace
Token budgetNot reported
Time / cost budgetmaximum 100 agent steps; exact timeout not reported
TemperatureNot reported
Seedindependent random seeds
Repeats3
Graderfive deterministic grader families · human review: no
Statisticstwo-stage evaluation-weighted mean with t-distribution 95% confidence intervals over per-evaluation means
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracyabsolutepercentmean over 146 evaluation-level means after three repeatsdeterministic task-specific grader
Stepsabsolutesteps per evaluationmean over evaluation-level three-run meansNot reported
Latencyabsoluteseconds per evaluationmean over evaluation-level three-run meansNot reported
CostabsoluteUSD per evaluationmean over evaluation-level three-run means with missing cost logs excludedNot reported

Results

ModelMetricValuen
Claude Opus 4.5Accuracy38.36 percent
Table 1
146
Claude Opus 4.5Steps2.84 steps per evaluation
Table 1
146
Claude Opus 4.5Latency123.8 seconds per evaluation
Table 1
146
Claude Opus 4.5Cost0.143 USD per evaluation
Table 1
146
Claude Sonnet 4.5Accuracy28.31 percent
Table 1
146
Claude Sonnet 4.5Steps2.43 steps per evaluation
Table 1
146
Claude Sonnet 4.5Latency115.6 seconds per evaluation
Table 1
146
Claude Sonnet 4.5Cost0.081 USD per evaluation
Table 1
146
GPT-5.2Accuracy34.02 percent
Table 1
146
GPT-5.2Steps2.1 steps per evaluation
Table 1
146
GPT-5.2Latency89.2 seconds per evaluation
Table 1
146
GPT-5.2Cost0.037 USD per evaluation
Table 1
146
GPT-5.1Accuracy27.4 percent
Table 1
146
GPT-5.1Steps2.38 steps per evaluation
Table 1
146
GPT-5.1Latency55.8 seconds per evaluation
Table 1
146
GPT-5.1Cost0.02 USD per evaluation
Table 1
146
Gemini 2.5 ProAccuracy20.09 percent
Table 1
146
Gemini 2.5 ProSteps3.61 steps per evaluation
Table 1
146
Gemini 2.5 ProLatency193.5 seconds per evaluation
Table 1
146
Gemini 2.5 ProCost0.188 USD per evaluation
Table 1
146
Grok-4Accuracy22.83 percent
Table 1
146
Grok-4Steps9.9 steps per evaluation
Table 1
146
Grok-4Latency173.2 seconds per evaluation
Table 1
146
Grok-4Cost0.048 USD per evaluation
Table 1
146
Grok-4.1 (SpatialBench paper label)Accuracy24.66 percent
Table 1
146
Grok-4.1 (SpatialBench paper label)Steps9.93 steps per evaluation
Table 1
146
Grok-4.1 (SpatialBench paper label)Latency196.4 seconds per evaluation
Table 1
146
Grok-4.1 (SpatialBench paper label)Cost0.077 USD per evaluation
Table 1
146

Evidence

  • page: arXiv v2 §§3.2–3.7 and Appendix A.4 (Problem anatomy, deterministic graders, interactive compute, isolation, three repeats, 100-step limit, and two-stage statistics.) — supports /benchmark_version, /scope, /protocol, /metrics
  • table: arXiv v2 Table 1 (Exact base-harness accuracy, steps, latency, cost, and 95% confidence intervals for seven source labels.) — supports /results

Evaluation run

spatialbench-paper-v2-claude-code

From SpatialBench: Can Agents Analyze Real-World Spatial Biology Data?

spatialbench-paper-v2-claude-codevpaper-v2

Evaluated models / systems: Claude Opus 4.5, Claude Sonnet 4.5

Scopefull · n=146
ShotsNot reported
Turnsmulti-turn Claude Code agent
System prompt publicNo
Reasoning / effortClaude Code harness
BrowserNot reported
InternetNot reported
DatabasesNot reported
Code executionYes
ContainerYes
External toolsClaude Code harness tools
Token budgetNot reported
Time / cost budgetfixed harness-specific budget; numeric value not reported
TemperatureNot reported
Seedindependent random seeds
Repeats3
Graderfive deterministic grader families · human review: no
Statisticstwo-stage evaluation-weighted mean with t-distribution 95% confidence intervals
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracyabsolutepercentmean over 146 evaluation-level means after three repeatsdeterministic task-specific grader

Results

ModelMetricValuen
Claude Opus 4.5Accuracy48.1 percent
Table 4
146
Claude Sonnet 4.5Accuracy45.1 percent
Table 4
146

Evidence

  • page: arXiv v2 §§2.5 and 3.5–3.7; Appendix A.4 (Harness separation, full scope, deterministic graders, three repeats, and statistics.) — supports /benchmark_version, /scope, /protocol, /metrics
  • table: arXiv v2 Table 4 (Exact labeled Claude Code accuracy and 95% confidence intervals.) — supports /results

Evaluation run

spatialbench-paper-v2-latch

From SpatialBench: Can Agents Analyze Real-World Spatial Biology Data?

spatialbench-paper-v2-latchvpaper-v2

Evaluated models / systems: Claude Opus 4.5

Scopefull · n=146
ShotsNot reported
Turnsmulti-turn Latch agent
System prompt publicNo
Reasoning / effortLatch agent harness
BrowserNot reported
InternetNot reported
DatabasesNot reported
Code executionYes
ContainerYes
External toolsLatch agent harness tools
Token budgetNot reported
Time / cost budgetfixed harness-specific budget; numeric value not reported
TemperatureNot reported
Seedindependent random seeds
Repeats3
Graderfive deterministic grader families · human review: no
Statisticstwo-stage evaluation-weighted mean with t-distribution 95% confidence intervals
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracyabsolutepercentmean over 146 evaluation-level means after three repeatsdeterministic task-specific grader

Results

ModelMetricValuen
Claude Opus 4.5Accuracy61.7 percent
Table 4
146

Evidence

  • page: arXiv v2 §§2.5 and 3.5–3.7; Appendix A.4 (Harness separation, full scope, deterministic graders, three repeats, and statistics.) — supports /benchmark_version, /scope, /protocol, /metrics
  • table: arXiv v2 Table 4 (Exact labeled Latch accuracy and 95% confidence interval.) — supports /results

Evaluation run

spatialbench-repo-159-claude-code

From SpatialBench 159-evaluation repository snapshot

spatialbench-repo-159-claude-codevrepo-159-5042c4f

Evaluated models / systems: Claude Opus 4.7, Claude Opus 4.8

Scopefull · n=159
ShotsNot reported
Turnsmulti-turn Claude Code
System prompt publicYes
Reasoning / effortmaximum reasoning effort where applicable
BrowserNot reported
InternetYes
DatabasesNot reported
Code executionYes
ContainerYes
External toolsClaude Code with common scientific libraries and network-enabled shell
Token budgetNot reported
Time / cost budget21600 seconds per task with no step limit
TemperatureNot reported
SeedNot reported
Repeats3
Graderfive deterministic grader families · human review: no
Statisticst-distribution 95% confidence interval over per-evaluation means after averaging three runs
Contaminationfull suite withheld to reduce contamination
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracyabsolutepercentmean over 159 evaluation-level means after three repeatsdeterministic task-specific grader
CostabsoluteUSD per evaluationmean over the full result setNot reported
Durationabsoluteseconds per evaluationmean over the full result setNot reported

Results

ModelMetricValuen
Claude Opus 4.8Accuracy55.35 percent
model_results.csv
159
Claude Opus 4.8Cost0.8776 USD per evaluation
model_results.csv
159
Claude Opus 4.8Duration489.16 seconds per evaluation
model_results.csv
159
Claude Opus 4.7Accuracy51.36 percent
model_results.csv
159
Claude Opus 4.7Cost0.8023 USD per evaluation
model_results.csv
159
Claude Opus 4.7Duration532.85 seconds per evaluation
model_results.csv
159

Evidence

  • repository-path: METHODS.md at commit 5042c4f3ee597da1590650c7b894d068ae968e26 (Three runs, max reasoning, container resources, six-hour timeout, network access, binary placement, and aggregation.) — supports /benchmark_version, /scope, /protocol, /metrics
  • repository-path: results/model_results.csv at commit 5042c4f3ee597da1590650c7b894d068ae968e26; harness=claude-code (Exact model label, accuracy and CI, mean cost, mean duration, and N.) — supports /results

Evaluation run

spatialbench-repo-159-mini-swe-agent

From SpatialBench 159-evaluation repository snapshot

spatialbench-repo-159-mini-swe-agentvrepo-159-5042c4f

Evaluated models / systems: Claude Opus 4.5, Claude Opus 4.6, Claude Opus 4.7, Claude Opus 4.8, Claude Sonnet 4.5, Claude Sonnet 4.6, Gemini 2.5 Pro, gemini-3.1-pro-preview, Gemini 3.5 Flash, GPT-5.1, GPT-5.2, GPT-5.4, GPT-5.5, Grok-4, grok-4-1-fast-reasoning, grok-4.20-beta-0309-reasoning

Scopefull · n=159
ShotsNot reported
Turnsmulti-turn mini-swe-agent
System prompt publicYes
Reasoning / effortmaximum reasoning effort where applicable
BrowserNot reported
InternetYes
DatabasesNot reported
Code executionYes
ContainerYes
External toolsmini-swe-agent with common scientific libraries and network-enabled shell
Token budgetNot reported
Time / cost budget21600 seconds per task with no step limit
TemperatureNot reported
SeedNot reported
Repeats3
Graderfive deterministic grader families · human review: no
Statisticst-distribution 95% confidence interval over per-evaluation means after averaging three runs
Contaminationfull suite withheld to reduce contamination
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracyabsolutepercentmean over 159 evaluation-level means after three repeatsdeterministic task-specific grader
CostabsoluteUSD per evaluationmean over the full result setNot reported
Durationabsoluteseconds per evaluationmean over the full result setNot reported

Results

ModelMetricValuen
GPT-5.5Accuracy57.65 percent
model_results.csv
159
GPT-5.5Cost1.1207 USD per evaluation
model_results.csv
159
GPT-5.5Duration586.66 seconds per evaluation
model_results.csv
159
GPT-5.4Accuracy57.44 percent
model_results.csv
159
GPT-5.4Cost0.577 USD per evaluation
model_results.csv
159
GPT-5.4Duration1128.75 seconds per evaluation
model_results.csv
159
Claude Opus 4.6Accuracy52.83 percent
model_results.csv
159
Claude Opus 4.6Cost0.8456 USD per evaluation
model_results.csv
159
Claude Opus 4.6Duration609.15 seconds per evaluation
model_results.csv
159
Claude Opus 4.8Accuracy52.62 percent
model_results.csv
159
Claude Opus 4.8Cost1.1061 USD per evaluation
model_results.csv
159
Claude Opus 4.8Duration902.86 seconds per evaluation
model_results.csv
159
Claude Opus 4.7Accuracy52.41 percent
model_results.csv
159
Claude Opus 4.7Cost0.9817 USD per evaluation
model_results.csv
159
Claude Opus 4.7Duration626.84 seconds per evaluation
model_results.csv
159
gemini-3.1-pro-previewAccuracy51.57 percent
model_results.csv
159
gemini-3.1-pro-previewCost0.9362 USD per evaluation
model_results.csv
159
gemini-3.1-pro-previewDuration1061.63 seconds per evaluation
model_results.csv
159
GPT-5.2Accuracy50.1 percent
model_results.csv
159
GPT-5.2Cost0.6024 USD per evaluation
model_results.csv
159
GPT-5.2Duration931.76 seconds per evaluation
model_results.csv
159
Gemini 3.5 FlashAccuracy48.85 percent
model_results.csv
159
Gemini 3.5 FlashCost2.7608 USD per evaluation
model_results.csv
159
Gemini 3.5 FlashDuration1145.4 seconds per evaluation
model_results.csv
159
grok-4.20-beta-0309-reasoningAccuracy45.91 percent
model_results.csv
159
grok-4.20-beta-0309-reasoningCost0.1679 USD per evaluation
model_results.csv
159
grok-4.20-beta-0309-reasoningDuration342.63 seconds per evaluation
model_results.csv
159
Claude Sonnet 4.6Accuracy44.23 percent
model_results.csv
159
Claude Sonnet 4.6Cost0.273 USD per evaluation
model_results.csv
159
Claude Sonnet 4.6Duration405.3 seconds per evaluation
model_results.csv
159
Claude Opus 4.5Accuracy42.77 percent
model_results.csv
159
Claude Opus 4.5Cost0.4624 USD per evaluation
model_results.csv
159
Claude Opus 4.5Duration376.54 seconds per evaluation
model_results.csv
159
Claude Sonnet 4.5Accuracy41.51 percent
model_results.csv
159
Claude Sonnet 4.5Cost0.2247 USD per evaluation
model_results.csv
159
Claude Sonnet 4.5Duration294.44 seconds per evaluation
model_results.csv
159
GPT-5.1Accuracy39.83 percent
model_results.csv
159
GPT-5.1Cost0.1574 USD per evaluation
model_results.csv
159
GPT-5.1Duration309.79 seconds per evaluation
model_results.csv
159
grok-4-1-fast-reasoningAccuracy33.96 percent
model_results.csv
159
grok-4-1-fast-reasoningCost0.0164 USD per evaluation
model_results.csv
159
grok-4-1-fast-reasoningDuration357 seconds per evaluation
model_results.csv
159
Grok-4Accuracy31.87 percent
model_results.csv
159
Grok-4Cost0.4529 USD per evaluation
model_results.csv
159
Grok-4Duration732.47 seconds per evaluation
model_results.csv
159
Gemini 2.5 ProAccuracy28.93 percent
model_results.csv
159
Gemini 2.5 ProCost0.1086 USD per evaluation
model_results.csv
159
Gemini 2.5 ProDuration231.07 seconds per evaluation
model_results.csv
159

Evidence

  • repository-path: METHODS.md at commit 5042c4f3ee597da1590650c7b894d068ae968e26 (Three runs, max reasoning, container resources, six-hour timeout, network access, harness placement, and aggregation.) — supports /benchmark_version, /scope, /protocol, /metrics
  • repository-path: results/model_results.csv at commit 5042c4f3ee597da1590650c7b894d068ae968e26; harness=mini-swe-agent (Exact model label, accuracy and CI, mean cost, mean duration, and N.) — supports /results

Evaluation run

spatialbench-repo-159-openai-codex

From SpatialBench 159-evaluation repository snapshot

spatialbench-repo-159-openai-codexvrepo-159-5042c4f

Evaluated models / systems: GPT-5.5

Scopefull · n=159
ShotsNot reported
Turnsmulti-turn OpenAI Codex
System prompt publicYes
Reasoning / effortmaximum reasoning effort where applicable
BrowserNot reported
InternetYes
DatabasesNot reported
Code executionYes
ContainerYes
External toolsOpenAI Codex with common scientific libraries and network-enabled shell
Token budgetNot reported
Time / cost budget21600 seconds per task with no step limit
TemperatureNot reported
SeedNot reported
Repeats3
Graderfive deterministic grader families · human review: no
Statisticst-distribution 95% confidence interval over per-evaluation means after averaging three runs
Contaminationfull suite withheld to reduce contamination
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracyabsolutepercentmean over 159 evaluation-level means after three repeatsdeterministic task-specific grader
CostabsoluteUSD per evaluationmean over the full result setNot reported
Durationabsoluteseconds per evaluationmean over the full result setNot reported

Results

ModelMetricValuen
GPT-5.5Accuracy53.67 percent
model_results.csv
159
GPT-5.5Cost3.1616 USD per evaluation
model_results.csv
159
GPT-5.5Duration382.01 seconds per evaluation
model_results.csv
159

Evidence

  • repository-path: METHODS.md at commit 5042c4f3ee597da1590650c7b894d068ae968e26 (Three runs, max reasoning, container resources, six-hour timeout, network access, binary placement, and aggregation.) — supports /benchmark_version, /scope, /protocol, /metrics
  • repository-path: results/model_results.csv at commit 5042c4f3ee597da1590650c7b894d068ae968e26; harness=openai-codex (Exact model label, accuracy and CI, mean cost, mean duration, and N.) — supports /results

Evaluation run

spatialbench-repo-159-pi

From SpatialBench 159-evaluation repository snapshot

spatialbench-repo-159-pivrepo-159-5042c4f

Evaluated models / systems: Gemini 3.5 Flash

Scopefull · n=159
ShotsNot reported
Turnsmulti-turn Pi agent
System prompt publicYes
Reasoning / effortmaximum reasoning effort where applicable
BrowserNot reported
InternetYes
DatabasesNot reported
Code executionYes
ContainerYes
External toolsPi harness with common scientific libraries and network-enabled shell
Token budgetNot reported
Time / cost budget21600 seconds per task with no step limit
TemperatureNot reported
SeedNot reported
Repeats3
Graderfive deterministic grader families · human review: no
Statisticst-distribution 95% confidence interval over per-evaluation means after averaging three runs
Contaminationfull suite withheld to reduce contamination
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracyabsolutepercentmean over 159 evaluation-level means after three repeatsdeterministic task-specific grader
CostabsoluteUSD per evaluationmean over the full result setNot reported
Durationabsoluteseconds per evaluationmean over the full result setNot reported

Results

ModelMetricValuen
Gemini 3.5 FlashAccuracy55.56 percent
model_results.csv
159
Gemini 3.5 FlashCost1.4254 USD per evaluation
model_results.csv
159
Gemini 3.5 FlashDuration1739.29 seconds per evaluation
model_results.csv
159

Evidence

  • repository-path: METHODS.md at commit 5042c4f3ee597da1590650c7b894d068ae968e26 (Three runs, max reasoning, container resources, six-hour timeout, network access, and aggregation.) — supports /benchmark_version, /scope, /protocol, /metrics
  • repository-path: results/model_results.csv at commit 5042c4f3ee597da1590650c7b894d068ae968e26; harness=pi (Exact model label, accuracy and CI, mean cost, mean duration, and N.) — supports /results

Comparable result views

Accuracy

spatialbench-paper-v2-base · spatialbench-paper-v2-base

CSV ↓
Accessible data table
ModelValueComparability group
Claude Opus 4.538.36spatialbench-paper-v2-base
Claude Sonnet 4.528.31spatialbench-paper-v2-base
GPT-5.234.02spatialbench-paper-v2-base
GPT-5.127.4spatialbench-paper-v2-base
Gemini 2.5 Pro20.09spatialbench-paper-v2-base
Grok-422.83spatialbench-paper-v2-base
Grok-4.1 (SpatialBench paper label)24.66spatialbench-paper-v2-base

Steps

spatialbench-paper-v2-base · spatialbench-paper-v2-base

CSV ↓
Accessible data table
ModelValueComparability group
Claude Opus 4.52.84spatialbench-paper-v2-base
Claude Sonnet 4.52.43spatialbench-paper-v2-base
GPT-5.22.1spatialbench-paper-v2-base
GPT-5.12.38spatialbench-paper-v2-base
Gemini 2.5 Pro3.61spatialbench-paper-v2-base
Grok-49.9spatialbench-paper-v2-base
Grok-4.1 (SpatialBench paper label)9.93spatialbench-paper-v2-base

Latency

spatialbench-paper-v2-base · spatialbench-paper-v2-base

CSV ↓
Accessible data table
ModelValueComparability group
Claude Opus 4.5123.8spatialbench-paper-v2-base
Claude Sonnet 4.5115.6spatialbench-paper-v2-base
GPT-5.289.2spatialbench-paper-v2-base
GPT-5.155.8spatialbench-paper-v2-base
Gemini 2.5 Pro193.5spatialbench-paper-v2-base
Grok-4173.2spatialbench-paper-v2-base
Grok-4.1 (SpatialBench paper label)196.4spatialbench-paper-v2-base

Cost

spatialbench-paper-v2-base · spatialbench-paper-v2-base

CSV ↓
Accessible data table
ModelValueComparability group
Claude Opus 4.50.143spatialbench-paper-v2-base
Claude Sonnet 4.50.081spatialbench-paper-v2-base
GPT-5.20.037spatialbench-paper-v2-base
GPT-5.10.02spatialbench-paper-v2-base
Gemini 2.5 Pro0.188spatialbench-paper-v2-base
Grok-40.048spatialbench-paper-v2-base
Grok-4.1 (SpatialBench paper label)0.077spatialbench-paper-v2-base

Accuracy

spatialbench-paper-v2-claude-code · spatialbench-paper-v2-claude-code

CSV ↓
Accessible data table
ModelValueComparability group
Claude Opus 4.548.1spatialbench-paper-v2-claude-code
Claude Sonnet 4.545.1spatialbench-paper-v2-claude-code

Accuracy

spatialbench-paper-v2-latch · spatialbench-paper-v2-latch

CSV ↓
Accessible data table
ModelValueComparability group
Claude Opus 4.561.7spatialbench-paper-v2-latch

Accuracy

spatialbench-repo-159-claude-code · spatialbench-repo-159-claude-code

CSV ↓
Accessible data table
ModelValueComparability group
Claude Opus 4.855.35spatialbench-repo-159-claude-code
Claude Opus 4.751.36spatialbench-repo-159-claude-code

Cost

spatialbench-repo-159-claude-code · spatialbench-repo-159-claude-code

CSV ↓
Accessible data table
ModelValueComparability group
Claude Opus 4.80.8776spatialbench-repo-159-claude-code
Claude Opus 4.70.8023spatialbench-repo-159-claude-code

Duration

spatialbench-repo-159-claude-code · spatialbench-repo-159-claude-code

CSV ↓
Accessible data table
ModelValueComparability group
Claude Opus 4.8489.16spatialbench-repo-159-claude-code
Claude Opus 4.7532.85spatialbench-repo-159-claude-code

Accuracy

spatialbench-repo-159-mini-swe-agent · spatialbench-repo-159-mini-swe-agent

CSV ↓
Accessible data table
ModelValueComparability group
GPT-5.557.65spatialbench-repo-159-mini-swe-agent
GPT-5.457.44spatialbench-repo-159-mini-swe-agent
Claude Opus 4.652.83spatialbench-repo-159-mini-swe-agent
Claude Opus 4.852.62spatialbench-repo-159-mini-swe-agent
Claude Opus 4.752.41spatialbench-repo-159-mini-swe-agent
gemini-3.1-pro-preview51.57spatialbench-repo-159-mini-swe-agent
GPT-5.250.1spatialbench-repo-159-mini-swe-agent
Gemini 3.5 Flash48.85spatialbench-repo-159-mini-swe-agent
grok-4.20-beta-0309-reasoning45.91spatialbench-repo-159-mini-swe-agent
Claude Sonnet 4.644.23spatialbench-repo-159-mini-swe-agent
Claude Opus 4.542.77spatialbench-repo-159-mini-swe-agent
Claude Sonnet 4.541.51spatialbench-repo-159-mini-swe-agent
GPT-5.139.83spatialbench-repo-159-mini-swe-agent
grok-4-1-fast-reasoning33.96spatialbench-repo-159-mini-swe-agent
Grok-431.87spatialbench-repo-159-mini-swe-agent
Gemini 2.5 Pro28.93spatialbench-repo-159-mini-swe-agent

Cost

spatialbench-repo-159-mini-swe-agent · spatialbench-repo-159-mini-swe-agent

CSV ↓
Accessible data table
ModelValueComparability group
GPT-5.51.1207spatialbench-repo-159-mini-swe-agent
GPT-5.40.577spatialbench-repo-159-mini-swe-agent
Claude Opus 4.60.8456spatialbench-repo-159-mini-swe-agent
Claude Opus 4.81.1061spatialbench-repo-159-mini-swe-agent
Claude Opus 4.70.9817spatialbench-repo-159-mini-swe-agent
gemini-3.1-pro-preview0.9362spatialbench-repo-159-mini-swe-agent
GPT-5.20.6024spatialbench-repo-159-mini-swe-agent
Gemini 3.5 Flash2.7608spatialbench-repo-159-mini-swe-agent
grok-4.20-beta-0309-reasoning0.1679spatialbench-repo-159-mini-swe-agent
Claude Sonnet 4.60.273spatialbench-repo-159-mini-swe-agent
Claude Opus 4.50.4624spatialbench-repo-159-mini-swe-agent
Claude Sonnet 4.50.2247spatialbench-repo-159-mini-swe-agent
GPT-5.10.1574spatialbench-repo-159-mini-swe-agent
grok-4-1-fast-reasoning0.0164spatialbench-repo-159-mini-swe-agent
Grok-40.4529spatialbench-repo-159-mini-swe-agent
Gemini 2.5 Pro0.1086spatialbench-repo-159-mini-swe-agent

Duration

spatialbench-repo-159-mini-swe-agent · spatialbench-repo-159-mini-swe-agent

CSV ↓
Accessible data table
ModelValueComparability group
GPT-5.5586.66spatialbench-repo-159-mini-swe-agent
GPT-5.41128.75spatialbench-repo-159-mini-swe-agent
Claude Opus 4.6609.15spatialbench-repo-159-mini-swe-agent
Claude Opus 4.8902.86spatialbench-repo-159-mini-swe-agent
Claude Opus 4.7626.84spatialbench-repo-159-mini-swe-agent
gemini-3.1-pro-preview1061.63spatialbench-repo-159-mini-swe-agent
GPT-5.2931.76spatialbench-repo-159-mini-swe-agent
Gemini 3.5 Flash1145.4spatialbench-repo-159-mini-swe-agent
grok-4.20-beta-0309-reasoning342.63spatialbench-repo-159-mini-swe-agent
Claude Sonnet 4.6405.3spatialbench-repo-159-mini-swe-agent
Claude Opus 4.5376.54spatialbench-repo-159-mini-swe-agent
Claude Sonnet 4.5294.44spatialbench-repo-159-mini-swe-agent
GPT-5.1309.79spatialbench-repo-159-mini-swe-agent
grok-4-1-fast-reasoning357spatialbench-repo-159-mini-swe-agent
Grok-4732.47spatialbench-repo-159-mini-swe-agent
Gemini 2.5 Pro231.07spatialbench-repo-159-mini-swe-agent

Accuracy

spatialbench-repo-159-openai-codex · spatialbench-repo-159-openai-codex

CSV ↓
Accessible data table
ModelValueComparability group
GPT-5.553.67spatialbench-repo-159-openai-codex

Cost

spatialbench-repo-159-openai-codex · spatialbench-repo-159-openai-codex

CSV ↓
Accessible data table
ModelValueComparability group
GPT-5.53.1616spatialbench-repo-159-openai-codex

Duration

spatialbench-repo-159-openai-codex · spatialbench-repo-159-openai-codex

CSV ↓
Accessible data table
ModelValueComparability group
GPT-5.5382.01spatialbench-repo-159-openai-codex

Accuracy

spatialbench-repo-159-pi · spatialbench-repo-159-pi

CSV ↓
Accessible data table
ModelValueComparability group
Gemini 3.5 Flash55.56spatialbench-repo-159-pi

Cost

spatialbench-repo-159-pi · spatialbench-repo-159-pi

CSV ↓
Accessible data table
ModelValueComparability group
Gemini 3.5 Flash1.4254spatialbench-repo-159-pi

Duration

spatialbench-repo-159-pi · spatialbench-repo-159-pi

CSV ↓
Accessible data table
ModelValueComparability group
Gemini 3.5 Flash1739.29spatialbench-repo-159-pi

Evidence and change history

Source locators remain visible; expand an item to inspect the exact Registry fields it supports.

SpatialBench: Can Agents Analyze Real-World Spatial Biology Data? · page: arXiv v2 pp. 1–3, abstract and §§1–2.1 (Benchmark name, creators, objective, release, real-data setup, modalities, and deterministic evaluation design.) · Supports 11 fields

Open source →

  • /name
  • /aliases
  • /summary
  • /kind
  • /parent_id
  • /organizations
  • /release_date
  • /domains
  • /capabilities
  • /modalities
  • /task_formats
SpatialBench: Can Agents Analyze Real-World Spatial Biology Data? · table: arXiv v2 Appendix A.1, Table 9 and Table 10 (Complete 146-evaluation inventory and both category and platform partitions.) · Supports 5 fields

Open source →

  • /versions/0/release_date
  • /versions/0/as_of
  • /versions/0/task_counts/total
  • /versions/0/task_counts/basis
  • /versions/0/task_counts/subsets
spatialbench-repository-resource · repository-path: README.md, results/category_results.json, and results/platform_results.json at commit 5042c4f3ee597da1590650c7b894d068ae968e26 (159 total evaluations; seven category counts and five platform counts.) · Supports 14 fields

Open source →

  • /latest_version
  • /task_counts/total
  • /task_counts/basis
  • /task_counts/subsets
  • /versions/1/release_date
  • /versions/1/as_of
  • /versions/1/task_counts/total
  • /versions/1/task_counts/basis
  • /versions/1/task_counts/subsets
  • /coverage_notes
  • /scientific_task_classification/entries/0
  • /scientific_task_classification/entries/1
  • /scientific_task_classification/entries/2
  • /scientific_task_classification/entries/3
spatialbench-repository-resource · repository-path: README.md, METHODS.md, LICENSE, example_evals/, results/, and spatialbench/ at commit 5042c4f3ee597da1590650c7b894d068ae968e26 (Representative-only task release, Apache-2.0 code, deterministic grader families, runner, container methods, and public aggregate results.) · Supports 8 fields

Open source →

  • /access/level
  • /access/tasks
  • /access/artifacts
  • /access/grader
  • /access/license
  • /access/biosafety_notes
  • /resources
  • /implementations

Unresolved field claims

  • /versions/0/task_counts/subsetsConflicted · high — Paper v2 Table 9 states 146 total evaluations, while both its seven category rows and five platform rows sum to 147; no correction is published.
    Evidence: spatialbench-evidence-paper-counts

View source-level modification history on GitHub →