preprint · benchmark creator

SpatialBench: Can Agents Analyze Real-World Spatial Biology Data?

LatchBio · 2025-12-26

Relationship layer

Benchmark usage

This table records what the work did with each benchmark before attempting to normalize a run. Partial claims remain visible without being treated as comparable evaluations.

benchmark creation

spatialbench-preprint-creation

Non-evaluationunknown

Benchmark: SpatialBench · version paper-v2

Selection
not applicable
Models
Not reported / not applicable
Metrics
Not reported / not applicable
Linked runs
None

Creator preprint defining the 146-problem paper-v2 benchmark snapshot.

Evidence
  • page: arXiv v2 abstract and §§2.1, 3.1–3.4 (Introduces and constructs SpatialBench.)
    Supports: /benchmark_version, /relation_type, /status, /scope, /notes

evaluation

spatialbench-preprint-evaluation

Normalizedfull · n=146

Benchmark: SpatialBench · version paper-v2

Selection
not applicable · all 146 paper-v2 evaluations
Metrics
Accuracy, Steps, Latency, Cost

Base, Claude Code, and Latch harnesses are normalized as separate runs and comparability groups.

Evidence
  • table: arXiv v2 Tables 1 and 4; §§2.2, 2.5, 3.5–3.7 (Model and harness results over all 146 evaluations.)
    Supports: /benchmark_version, /relation_type, /status, /model_ids, /scope, /metric_labels, /evaluation_run_ids, /notes

Normalized evaluation runs

SpatialBench3 runs

Open benchmark record →

Evaluation run

spatialbench-paper-v2-base

From SpatialBench: Can Agents Analyze Real-World Spatial Biology Data?

spatialbench-paper-v2-basevpaper-v2

Evaluated models / systems: Claude Opus 4.5, Claude Sonnet 4.5, Gemini 2.5 Pro, GPT-5.1, GPT-5.2, Grok-4, Grok-4.1 (SpatialBench paper label)

Scopefull · n=146
ShotsNot reported
Turnsmulti-turn agent
System prompt publicNo
Reasoning / effortslightly modified Mini-SWE-Bench base harness
BrowserNot reported
InternetNot reported
DatabasesNot reported
Code executionYes
ContainerYes
External toolsinteractive compute and local workspace
Token budgetNot reported
Time / cost budgetmaximum 100 agent steps; exact timeout not reported
TemperatureNot reported
Seedindependent random seeds
Repeats3
Graderfive deterministic grader families · human review: no
Statisticstwo-stage evaluation-weighted mean with t-distribution 95% confidence intervals over per-evaluation means
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracyabsolutepercentmean over 146 evaluation-level means after three repeatsdeterministic task-specific grader
Stepsabsolutesteps per evaluationmean over evaluation-level three-run meansNot reported
Latencyabsoluteseconds per evaluationmean over evaluation-level three-run meansNot reported
CostabsoluteUSD per evaluationmean over evaluation-level three-run means with missing cost logs excludedNot reported

Results

ModelMetricValuen
Claude Opus 4.5Accuracy38.36 percent
Table 1
146
Claude Opus 4.5Steps2.84 steps per evaluation
Table 1
146
Claude Opus 4.5Latency123.8 seconds per evaluation
Table 1
146
Claude Opus 4.5Cost0.143 USD per evaluation
Table 1
146
Claude Sonnet 4.5Accuracy28.31 percent
Table 1
146
Claude Sonnet 4.5Steps2.43 steps per evaluation
Table 1
146
Claude Sonnet 4.5Latency115.6 seconds per evaluation
Table 1
146
Claude Sonnet 4.5Cost0.081 USD per evaluation
Table 1
146
GPT-5.2Accuracy34.02 percent
Table 1
146
GPT-5.2Steps2.1 steps per evaluation
Table 1
146
GPT-5.2Latency89.2 seconds per evaluation
Table 1
146
GPT-5.2Cost0.037 USD per evaluation
Table 1
146
GPT-5.1Accuracy27.4 percent
Table 1
146
GPT-5.1Steps2.38 steps per evaluation
Table 1
146
GPT-5.1Latency55.8 seconds per evaluation
Table 1
146
GPT-5.1Cost0.02 USD per evaluation
Table 1
146
Gemini 2.5 ProAccuracy20.09 percent
Table 1
146
Gemini 2.5 ProSteps3.61 steps per evaluation
Table 1
146
Gemini 2.5 ProLatency193.5 seconds per evaluation
Table 1
146
Gemini 2.5 ProCost0.188 USD per evaluation
Table 1
146
Grok-4Accuracy22.83 percent
Table 1
146
Grok-4Steps9.9 steps per evaluation
Table 1
146
Grok-4Latency173.2 seconds per evaluation
Table 1
146
Grok-4Cost0.048 USD per evaluation
Table 1
146
Grok-4.1 (SpatialBench paper label)Accuracy24.66 percent
Table 1
146
Grok-4.1 (SpatialBench paper label)Steps9.93 steps per evaluation
Table 1
146
Grok-4.1 (SpatialBench paper label)Latency196.4 seconds per evaluation
Table 1
146
Grok-4.1 (SpatialBench paper label)Cost0.077 USD per evaluation
Table 1
146

Evidence

  • page: arXiv v2 §§3.2–3.7 and Appendix A.4 (Problem anatomy, deterministic graders, interactive compute, isolation, three repeats, 100-step limit, and two-stage statistics.) — supports /benchmark_version, /scope, /protocol, /metrics
  • table: arXiv v2 Table 1 (Exact base-harness accuracy, steps, latency, cost, and 95% confidence intervals for seven source labels.) — supports /results

Evaluation run

spatialbench-paper-v2-claude-code

From SpatialBench: Can Agents Analyze Real-World Spatial Biology Data?

spatialbench-paper-v2-claude-codevpaper-v2

Evaluated models / systems: Claude Opus 4.5, Claude Sonnet 4.5

Scopefull · n=146
ShotsNot reported
Turnsmulti-turn Claude Code agent
System prompt publicNo
Reasoning / effortClaude Code harness
BrowserNot reported
InternetNot reported
DatabasesNot reported
Code executionYes
ContainerYes
External toolsClaude Code harness tools
Token budgetNot reported
Time / cost budgetfixed harness-specific budget; numeric value not reported
TemperatureNot reported
Seedindependent random seeds
Repeats3
Graderfive deterministic grader families · human review: no
Statisticstwo-stage evaluation-weighted mean with t-distribution 95% confidence intervals
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracyabsolutepercentmean over 146 evaluation-level means after three repeatsdeterministic task-specific grader

Results

ModelMetricValuen
Claude Opus 4.5Accuracy48.1 percent
Table 4
146
Claude Sonnet 4.5Accuracy45.1 percent
Table 4
146

Evidence

  • page: arXiv v2 §§2.5 and 3.5–3.7; Appendix A.4 (Harness separation, full scope, deterministic graders, three repeats, and statistics.) — supports /benchmark_version, /scope, /protocol, /metrics
  • table: arXiv v2 Table 4 (Exact labeled Claude Code accuracy and 95% confidence intervals.) — supports /results

Evaluation run

spatialbench-paper-v2-latch

From SpatialBench: Can Agents Analyze Real-World Spatial Biology Data?

spatialbench-paper-v2-latchvpaper-v2

Evaluated models / systems: Claude Opus 4.5

Scopefull · n=146
ShotsNot reported
Turnsmulti-turn Latch agent
System prompt publicNo
Reasoning / effortLatch agent harness
BrowserNot reported
InternetNot reported
DatabasesNot reported
Code executionYes
ContainerYes
External toolsLatch agent harness tools
Token budgetNot reported
Time / cost budgetfixed harness-specific budget; numeric value not reported
TemperatureNot reported
Seedindependent random seeds
Repeats3
Graderfive deterministic grader families · human review: no
Statisticstwo-stage evaluation-weighted mean with t-distribution 95% confidence intervals
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracyabsolutepercentmean over 146 evaluation-level means after three repeatsdeterministic task-specific grader

Results

ModelMetricValuen
Claude Opus 4.5Accuracy61.7 percent
Table 4
146

Evidence

  • page: arXiv v2 §§2.5 and 3.5–3.7; Appendix A.4 (Harness separation, full scope, deterministic graders, three repeats, and statistics.) — supports /benchmark_version, /scope, /protocol, /metrics
  • table: arXiv v2 Table 4 (Exact labeled Latch accuracy and 95% confidence interval.) — supports /results