track · audited · verified 2026-07-21

SCIGYM Large

The formally released SCIGYM track containing the 213 systems not included in the creator paper's model evaluation, with systems reaching up to 400 reactions.

Benchmark definition

What is counted

Version
2025 release
Total
213 (unique SBML systems in the official large Parquet split, containing the remaining systems with up to 400 reactions)
Task formats
interactive simulated experiment; SBML reaction-network reconstruction
Capabilities
Experiment planningData analysisCodingTool useScientific reasoning
Modalities
TextTableCode

Version history

VersionStatusRelease / as-ofTotalFormal tracks
2025 release
scigym-large-2025-release
current2025-05-16213 (unique SBML systems in the official large Parquet split, containing the remaining systems with up to 400 reactions)None registered

Scientific Task Atlas

Scientific task classification

complete for 2025 release. Official large-system split.

Scientific taskCoverageCountMappingEvidence
Reaction-network reconstructionexplicitly-in-scope213 systems
unique SBML systems in the official large Parquet split, containing the remaining systems with up to 400 reactions
official-track
high confidence
scigym-large-evidence-count
Split-specific system count.
Simulation-based experimentexplicitly-in-scope213 systems
unique SBML systems in the official large Parquet split, containing the remaining systems with up to 400 reactions
official-track
high confidence
scigym-large-evidence-count
Same systems; overlapping task claim.

Scientific coverage notes

DomainCoverageCountInterpretation
GenomicsunknownNot reportedThe official split does not publish a task-level genomics classification.
Transcriptomicsnot-in-scope0Tasks consume SBML and simulated trajectories rather than transcriptomic measurements.
Protein designnot-in-scope0The task reconstructs reaction networks rather than proteins.
Protein-protein bindingnot-in-scope0No task evaluates protein-protein binding.
Protein-ligand bindingnot-in-scope0No task evaluates protein-ligand binding.

Evaluation registry

Works and run settings

A setting change—scope, prompt, tools, budget, grader, or repeats—creates a separate run. Charts never cross a comparability group.

No normalized evaluation run is published yet. Creator evidence is still attached below.

Evidence and change history

Source locators remain visible; expand an item to inspect the exact Registry fields it supports.

scigym-large-dataset-resource · dataset-card: README.md and data/large-00000-of-00001.parquet at commit dbb10c12a33427e8e05ebbac66de690c65b6acae (213 unique large-split systems; the paper describes them as the remaining released systems with up to 400 reactions and says they were not evaluated.) · Supports 17 fields

Open source →

  • /name
  • /summary
  • /kind
  • /parent_id
  • /organizations
  • /release_date
  • /latest_version
  • /domains
  • /capabilities
  • /modalities
  • /task_formats
  • /task_counts/total
  • /task_counts/basis
  • /task_counts/subsets
  • /versions/0/task_counts
  • /scientific_task_classification/entries/0
  • /scientific_task_classification/entries/1
scigym-large-repository-resource · repository-path: README.md, data/download.py, scigym/, and absence of LICENSE at commit 88a7b93609e35b6ecb4eb343d816d6ff09256c6a (Public large-split download path, environment, and grader with no stated software or data license.) · Supports 9 fields

Open source →

  • /access/level
  • /access/tasks
  • /access/artifacts
  • /access/grader
  • /access/license
  • /access/biosafety_notes
  • /resources
  • /implementations
  • /coverage_notes

View source-level modification history on GitHub →