paper · benchmark creator

Measuring Scientific Capabilities of Language Models with a Systems Biology Dry Lab

University of Toronto · SickKids · Axiom · Mila · Vector Institute · 2025-11-30

Relationship layer

Benchmark usage

This table records what the work did with each benchmark before attempting to normalize a run. Partial claims remain visible without being treated as comparable evaluations.

No BenchmarkUse relation is normalized for this legacy work yet. Existing EvaluationRuns remain available below.

Normalized evaluation runs

SCIGYM Small2 runs

Open benchmark record →

scigym-2025-small-react-initial-concentration-20step-3repeatv2025 release

Evaluated models / systems: Claude 3.5 Haiku 20241022 (SCIGYM), Claude 3.7 Sonnet 20250219 (SCIGYM), Gemini 2.5 Flash Preview 04-17 (SCIGYM), Gemini 2.5 Pro Preview 03-25 (SCIGYM), GPT-4.1 2025-04-14 (SCIGYM), GPT-4.1 Mini 2025-04-14 (SCIGYM)

Scopefull · n=137
ShotsNot reported
Turnsmulti-turn
System prompt publicYes
Reasoning / effortReAct-style Thoughts–Actions–Observations agent; no provider reasoning-effort control is reported.
BrowserNo
InternetNo
DatabasesNo
Code executionYes
ContainerNot reported
External toolsTellurium, libRoadRunner, libSBML, pandas, numpy, SCIGYM experiment API
Token budgetmaximum 8,192 output tokens per model response
Time / cost budgetmaximum 20 action iterations plus up to 3 invalid-submission debugging iterations
TemperatureNot reported
SeedNot reported
Repeats3
Graderdeterministic SBML structural and simulation evaluator · human review: no
StatisticsTable 1 reports arithmetic means across the small benchmark instances. The public artifact supplies three episodes for every model-system pair; Table 1 does not print confidence intervals.
ContaminationReference models are de-identified by stripping metadata, shuffling components, and replacing component IDs; species names are retained. No model-training decontamination test is reported.
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Network Topology Score (NTS) F1absolutescorepairwise species-interaction F1 with duplicate relationships counted once, then averaged across systems/episodesrelationship type includes reactant-product, reactant-modifier, and modifier-product
Simulation Trajectory Error (STE)absoluteSMAPE errorSMAPE averaged across species and benchmark systems/episodesevaluated under original and perturbed initial conditions
RMS with modifiers — Precisionabsolutescoreexact reactant/product/modifier reaction matching averaged across systems/episodesreactants, products, and modifiers must match
RMS with modifiers — Recallabsolutescoreexact reactant/product/modifier reaction matching averaged across systems/episodesreactants, products, and modifiers must match
RMS with modifiers — F1absolutescoreharmonic mean of with-modifier reaction precision and recall averaged across systems/episodesreactants, products, and modifiers must match
RMS without modifiers — Precisionabsolutescoreexact reactant/product reaction matching averaged across systems/episodesmodifiers are ignored
RMS without modifiers — Recallabsolutescoreexact reactant/product reaction matching averaged across systems/episodesmodifiers are ignored
RMS without modifiers — F1absolutescoreharmonic mean of without-modifier reaction precision and recall averaged across systems/episodesmodifiers are ignored

Results

ModelMetricValuen
Gemini 2.5 Flash Preview 04-17 (SCIGYM)Simulation Trajectory Error (STE)0.4181 SMAPE error
Table 1; three public episodes per system.
137
Gemini 2.5 Flash Preview 04-17 (SCIGYM)RMS with modifiers — Precision0.1527 score
Table 1; three public episodes per system.
137
Gemini 2.5 Flash Preview 04-17 (SCIGYM)RMS with modifiers — Recall0.1071 score
Table 1; three public episodes per system.
137
Gemini 2.5 Flash Preview 04-17 (SCIGYM)RMS with modifiers — F10.1217 score
Table 1; three public episodes per system.
137
Gemini 2.5 Flash Preview 04-17 (SCIGYM)RMS without modifiers — Precision0.2399 score
Table 1; three public episodes per system.
137
Gemini 2.5 Flash Preview 04-17 (SCIGYM)RMS without modifiers — Recall0.1839 score
Table 1; three public episodes per system.
137
Gemini 2.5 Flash Preview 04-17 (SCIGYM)RMS without modifiers — F10.2005 score
Table 1; three public episodes per system.
137
GPT-4.1 Mini 2025-04-14 (SCIGYM)Simulation Trajectory Error (STE)0.6007 SMAPE error
Table 1; three public episodes per system.
137
GPT-4.1 Mini 2025-04-14 (SCIGYM)RMS with modifiers — Precision0.1516 score
Table 1; three public episodes per system.
137
GPT-4.1 Mini 2025-04-14 (SCIGYM)RMS with modifiers — Recall0.1253 score
Table 1; three public episodes per system.
137
GPT-4.1 Mini 2025-04-14 (SCIGYM)RMS with modifiers — F10.132 score
Table 1; three public episodes per system.
137
GPT-4.1 Mini 2025-04-14 (SCIGYM)RMS without modifiers — Precision0.253 score
Table 1; three public episodes per system.
137
GPT-4.1 Mini 2025-04-14 (SCIGYM)RMS without modifiers — Recall0.2313 score
Table 1; three public episodes per system.
137
GPT-4.1 Mini 2025-04-14 (SCIGYM)RMS without modifiers — F10.2322 score
Table 1; three public episodes per system.
137
Claude 3.5 Haiku 20241022 (SCIGYM)Simulation Trajectory Error (STE)0.6281 SMAPE error
Table 1; three public episodes per system.
137
Claude 3.5 Haiku 20241022 (SCIGYM)RMS with modifiers — Precision0.0858 score
Table 1; three public episodes per system.
137
Claude 3.5 Haiku 20241022 (SCIGYM)RMS with modifiers — Recall0.0421 score
Table 1; three public episodes per system.
137
Claude 3.5 Haiku 20241022 (SCIGYM)RMS with modifiers — F10.053 score
Table 1; three public episodes per system.
137
Claude 3.5 Haiku 20241022 (SCIGYM)RMS without modifiers — Precision0.1454 score
Table 1; three public episodes per system.
137
Claude 3.5 Haiku 20241022 (SCIGYM)RMS without modifiers — Recall0.0805 score
Table 1; three public episodes per system.
137
Claude 3.5 Haiku 20241022 (SCIGYM)RMS without modifiers — F10.0987 score
Table 1; three public episodes per system.
137
Gemini 2.5 Pro Preview 03-25 (SCIGYM)Simulation Trajectory Error (STE)0.3212 SMAPE error
Table 1; three public episodes per system.
137
Gemini 2.5 Pro Preview 03-25 (SCIGYM)RMS with modifiers — Precision0.2138 score
Table 1; three public episodes per system.
137
Gemini 2.5 Pro Preview 03-25 (SCIGYM)RMS with modifiers — Recall0.1664 score
Table 1; three public episodes per system.
137
Gemini 2.5 Pro Preview 03-25 (SCIGYM)RMS with modifiers — F10.1817 score
Table 1; three public episodes per system.
137
Gemini 2.5 Pro Preview 03-25 (SCIGYM)RMS without modifiers — Precision0.3781 score
Table 1; three public episodes per system.
137
Gemini 2.5 Pro Preview 03-25 (SCIGYM)RMS without modifiers — Recall0.3219 score
Table 1; three public episodes per system.
137
Gemini 2.5 Pro Preview 03-25 (SCIGYM)RMS without modifiers — F10.3383 score
Table 1; three public episodes per system.
137
GPT-4.1 2025-04-14 (SCIGYM)Simulation Trajectory Error (STE)0.4611 SMAPE error
Table 1; three public episodes per system.
137
GPT-4.1 2025-04-14 (SCIGYM)RMS with modifiers — Precision0.2067 score
Table 1; three public episodes per system.
137
GPT-4.1 2025-04-14 (SCIGYM)RMS with modifiers — Recall0.1597 score
Table 1; three public episodes per system.
137
GPT-4.1 2025-04-14 (SCIGYM)RMS with modifiers — F10.174 score
Table 1; three public episodes per system.
137
GPT-4.1 2025-04-14 (SCIGYM)RMS without modifiers — Precision0.3517 score
Table 1; three public episodes per system.
137
GPT-4.1 2025-04-14 (SCIGYM)RMS without modifiers — Recall0.2888 score
Table 1; three public episodes per system.
137
GPT-4.1 2025-04-14 (SCIGYM)RMS without modifiers — F10.3038 score
Table 1; three public episodes per system.
137
Claude 3.7 Sonnet 20250219 (SCIGYM)Simulation Trajectory Error (STE)0.3615 SMAPE error
Table 1; three public episodes per system.
137
Claude 3.7 Sonnet 20250219 (SCIGYM)RMS with modifiers — Precision0.178 score
Table 1; three public episodes per system.
137
Claude 3.7 Sonnet 20250219 (SCIGYM)RMS with modifiers — Recall0.1698 score
Table 1; three public episodes per system.
137
Claude 3.7 Sonnet 20250219 (SCIGYM)RMS with modifiers — F10.1688 score
Table 1; three public episodes per system.
137
Claude 3.7 Sonnet 20250219 (SCIGYM)RMS without modifiers — Precision0.316 score
Table 1; three public episodes per system.
137
Claude 3.7 Sonnet 20250219 (SCIGYM)RMS without modifiers — Recall0.317 score
Table 1; three public episodes per system.
137
Claude 3.7 Sonnet 20250219 (SCIGYM)RMS without modifiers — F10.3047 score
Table 1; three public episodes per system.
137

Evidence

  • section: NeurIPS paper §§3.2–3.3 and 5; Appendix A (Defines the ReAct loop, three action types, public prompt, initial-concentration experiment, 20 action iterations, three debugging iterations, six exact model versions, full 137-system small scope, and NTS/RMS/STE.) — supports /benchmark_version, /scope, /model_ids, /protocol/shots, /protocol/turns, /protocol/system_prompt_public, /protocol/reasoning, /protocol/tools, /protocol/time_budget, /protocol/grader, /protocol/statistical, /protocol/contamination, /metrics, /comparability_group
  • dataset-card: data/small-00000-of-00001.parquet at commit 7d472c12855d46702c4915892578290355894c1a (2,466 rows; 137 unique systems, six exact model strings, and exactly three rows for every model-system pair.) — supports /model_ids, /protocol/repeats, /protocol/statistical
  • repository-path: README.md, scigym/llm.py, scigym/controller.py, scigym/system_prompts/, and scigym/evaluator.py at commit d290bb04bf54aad1c473c4701e0d0d88013c4f91 (Confirms exact API strings, maximum 8,192 response tokens, no browser/network/database tool, public Python/SBML tool loop, and deterministic evaluator.) — supports /model_ids, /protocol/system_prompt_public, /protocol/tools, /protocol/token_budget, /protocol/grader, /metrics
  • table: NeurIPS paper Table 1 (Prints STE and RMS precision/recall/F1 with and without modifiers for all six models; no confidence bounds are printed.) — supports /results
scigym-2025-small-zero-shot-no-toolsv2025 release

Evaluated models / systems: Claude 3.5 Haiku 20241022 (SCIGYM), Claude 3.7 Sonnet 20250219 (SCIGYM), Gemini 2.5 Flash Preview 04-17 (SCIGYM), Gemini 2.5 Pro Preview 03-25 (SCIGYM), GPT-4.1 2025-04-14 (SCIGYM), GPT-4.1 Mini 2025-04-14 (SCIGYM)

Scopefull · n=137
Shotszero-shot
Turnsone direct submission, with up to three debugging responses after invalid submissions
System prompt publicYes
Reasoning / effortdirect prompting without experimental feedback
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNo
External toolsnone
Token budgetmaximum 8,192 output tokens per model response
Time / cost budgetone direct submission plus up to 3 invalid-submission debugging iterations
TemperatureNot reported
SeedNot reported
RepeatsNot reported
Graderdeterministic SBML structural and simulation evaluator · human review: no
StatisticsFigure 5 plots per-system agent and zero-shot scores; exact aggregate baseline values and confidence intervals are not tabulated.
ContaminationThe same de-identified small systems are used; no model-training decontamination test is reported.
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Network Topology Score (NTS) F1absolutescoreper-system structural F1duplicate relationships counted once
Simulation Trajectory Error (STE)absoluteSMAPE errorspecies-level SMAPE averaged within systemoriginal and perturbed initial conditions
RMS with modifiers — F1absolutescoreexact-reaction F1reactants products and modifiers must match
RMS without modifiers — F1absolutescoreexact-reaction F1modifiers ignored

No numeric result rows are published yet; the verified protocol remains useful.

Evidence

  • figure: NeurIPS paper Figure 5 and §5.1 (Compares each of the six agents against a zero-shot/direct-prompt baseline across small systems, with experiment and tools removed and three debugging rounds; exact aggregates are not tabulated.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics, /results, /comparability_group