suite · audited · verified 2026-07-21

SCIGYM

An agentic systems-biology suite in which language models iteratively perturb simulated SBML systems, analyze time-series observations in Python, and reconstruct hidden biological reactions.

Release, split, and protocol audit: SCIGYM releases 350 SBML systems: 137 small systems with fewer than ten reactions and 213 large systems with up to 400 reactions. Only the small track was evaluated in the creator paper. Its public evaluation artifact proves three episodes for every model–system pair; the code and dataset are public, but upstream states no code or data license.

Benchmark definition

What is counted

Version
2025 release
Total
350 (distinct curated BioModels systems released as SBML benchmark instances)
Task formats
interactive simulated experiment; SBML reaction-network reconstruction
Capabilities
Experiment planningData analysisCodingTool useScientific reasoning
Modalities
TextTableCode

Version history

VersionStatusRelease / as-ofTotalFormal tracks
2025 release
scigym-2025-release
current2025-05-16350 (distinct curated BioModels systems released as SBML benchmark instances)scigym-small, scigym-large

Tracks and subsets

IDCountBasisPartition?Notes
Small systems
scigym-small-systems
137systems with fewer than 10 reactions in the evaluated small splitExclusive & exhaustiveAll six creator-paper models were evaluated on this split.
Large systems
scigym-large-systems
213remaining released systems with up to 400 reactions in the large splitExclusive & exhaustiveReleased by the creators but not evaluated in the paper.

Registered child tracks

SCIGYM Large

The formally released SCIGYM track containing the 213 systems not included in the creator paper's model evaluation, with systems reaching up to 400 reactions.

0 evaluation run(s)

SCIGYM Small

The formally released and creator-evaluated SCIGYM track containing biological systems with fewer than ten reactions.

2 evaluation run(s)

Scientific Task Atlas

Scientific task classification

complete for 2025 release. All released systems use the same simulator-based reaction-network reconstruction task.

Scientific taskCoverageCountMappingEvidence
Reaction-network reconstructionexplicitly-in-scope350 systems
distinct curated BioModels systems released as SBML benchmark instances
official-taxonomy
high confidence
scigym-evidence-release-counts
scigym-evidence-taxonomy
Hidden biological reactions are reconstructed from interventions and trajectories.
Simulation-based experimentexplicitly-in-scope350 systems
distinct curated BioModels systems released as SBML benchmark instances
official-taxonomy
high confidence
scigym-evidence-release-counts
scigym-evidence-taxonomy
The same systems support iterative in-silico perturbation experiments; overlapping claims are not summed.

Scientific coverage notes

DomainCoverageCountInterpretation
Molecular and cell biologyexplicitly-in-scopeNot reportedThe source systems include signaling, metabolic, regulatory, and other biological processes, but the creators do not publish a mutually exclusive molecular/cell-biology count.
GenomicsunknownNot reportedGene-regulatory networks are named as an example system class, but no official task-level genomics count is published and the benchmark does not use genomic sequence data.
Transcriptomicsnot-in-scope0The paper motivates the task by analogy to Perturb-seq and spatial transcriptomics, but SCIGYM inputs are SBML systems and simulated concentration trajectories rather than transcriptomic measurements.
Single-cellnot-in-scope0No single-cell measurement is an official benchmark input or target.
Protein designnot-in-scope0Agents reconstruct reaction networks; they do not design or optimize protein sequences or structures.
Protein-protein bindingnot-in-scope0Species can represent proteins, but no task evaluates protein-protein binding prediction or affinity.
Protein-ligand bindingnot-in-scope0No task evaluates protein-ligand binding prediction or affinity.

Evaluation registry

Works and run settings

A setting change—scope, prompt, tools, budget, grader, or repeats—creates a separate run. Charts never cross a comparability group.

scigym-2025-small-react-initial-concentration-20step-3repeatv2025 release

Evaluated models / systems: Claude 3.5 Haiku 20241022 (SCIGYM), Claude 3.7 Sonnet 20250219 (SCIGYM), Gemini 2.5 Flash Preview 04-17 (SCIGYM), Gemini 2.5 Pro Preview 03-25 (SCIGYM), GPT-4.1 2025-04-14 (SCIGYM), GPT-4.1 Mini 2025-04-14 (SCIGYM)

Scopefull · n=137
ShotsNot reported
Turnsmulti-turn
System prompt publicYes
Reasoning / effortReAct-style Thoughts–Actions–Observations agent; no provider reasoning-effort control is reported.
BrowserNo
InternetNo
DatabasesNo
Code executionYes
ContainerNot reported
External toolsTellurium, libRoadRunner, libSBML, pandas, numpy, SCIGYM experiment API
Token budgetmaximum 8,192 output tokens per model response
Time / cost budgetmaximum 20 action iterations plus up to 3 invalid-submission debugging iterations
TemperatureNot reported
SeedNot reported
Repeats3
Graderdeterministic SBML structural and simulation evaluator · human review: no
StatisticsTable 1 reports arithmetic means across the small benchmark instances. The public artifact supplies three episodes for every model-system pair; Table 1 does not print confidence intervals.
ContaminationReference models are de-identified by stripping metadata, shuffling components, and replacing component IDs; species names are retained. No model-training decontamination test is reported.
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Network Topology Score (NTS) F1absolutescorepairwise species-interaction F1 with duplicate relationships counted once, then averaged across systems/episodesrelationship type includes reactant-product, reactant-modifier, and modifier-product
Simulation Trajectory Error (STE)absoluteSMAPE errorSMAPE averaged across species and benchmark systems/episodesevaluated under original and perturbed initial conditions
RMS with modifiers — Precisionabsolutescoreexact reactant/product/modifier reaction matching averaged across systems/episodesreactants, products, and modifiers must match
RMS with modifiers — Recallabsolutescoreexact reactant/product/modifier reaction matching averaged across systems/episodesreactants, products, and modifiers must match
RMS with modifiers — F1absolutescoreharmonic mean of with-modifier reaction precision and recall averaged across systems/episodesreactants, products, and modifiers must match
RMS without modifiers — Precisionabsolutescoreexact reactant/product reaction matching averaged across systems/episodesmodifiers are ignored
RMS without modifiers — Recallabsolutescoreexact reactant/product reaction matching averaged across systems/episodesmodifiers are ignored
RMS without modifiers — F1absolutescoreharmonic mean of without-modifier reaction precision and recall averaged across systems/episodesmodifiers are ignored

Results

ModelMetricValuen
Gemini 2.5 Flash Preview 04-17 (SCIGYM)Simulation Trajectory Error (STE)0.4181 SMAPE error
Table 1; three public episodes per system.
137
Gemini 2.5 Flash Preview 04-17 (SCIGYM)RMS with modifiers — Precision0.1527 score
Table 1; three public episodes per system.
137
Gemini 2.5 Flash Preview 04-17 (SCIGYM)RMS with modifiers — Recall0.1071 score
Table 1; three public episodes per system.
137
Gemini 2.5 Flash Preview 04-17 (SCIGYM)RMS with modifiers — F10.1217 score
Table 1; three public episodes per system.
137
Gemini 2.5 Flash Preview 04-17 (SCIGYM)RMS without modifiers — Precision0.2399 score
Table 1; three public episodes per system.
137
Gemini 2.5 Flash Preview 04-17 (SCIGYM)RMS without modifiers — Recall0.1839 score
Table 1; three public episodes per system.
137
Gemini 2.5 Flash Preview 04-17 (SCIGYM)RMS without modifiers — F10.2005 score
Table 1; three public episodes per system.
137
GPT-4.1 Mini 2025-04-14 (SCIGYM)Simulation Trajectory Error (STE)0.6007 SMAPE error
Table 1; three public episodes per system.
137
GPT-4.1 Mini 2025-04-14 (SCIGYM)RMS with modifiers — Precision0.1516 score
Table 1; three public episodes per system.
137
GPT-4.1 Mini 2025-04-14 (SCIGYM)RMS with modifiers — Recall0.1253 score
Table 1; three public episodes per system.
137
GPT-4.1 Mini 2025-04-14 (SCIGYM)RMS with modifiers — F10.132 score
Table 1; three public episodes per system.
137
GPT-4.1 Mini 2025-04-14 (SCIGYM)RMS without modifiers — Precision0.253 score
Table 1; three public episodes per system.
137
GPT-4.1 Mini 2025-04-14 (SCIGYM)RMS without modifiers — Recall0.2313 score
Table 1; three public episodes per system.
137
GPT-4.1 Mini 2025-04-14 (SCIGYM)RMS without modifiers — F10.2322 score
Table 1; three public episodes per system.
137
Claude 3.5 Haiku 20241022 (SCIGYM)Simulation Trajectory Error (STE)0.6281 SMAPE error
Table 1; three public episodes per system.
137
Claude 3.5 Haiku 20241022 (SCIGYM)RMS with modifiers — Precision0.0858 score
Table 1; three public episodes per system.
137
Claude 3.5 Haiku 20241022 (SCIGYM)RMS with modifiers — Recall0.0421 score
Table 1; three public episodes per system.
137
Claude 3.5 Haiku 20241022 (SCIGYM)RMS with modifiers — F10.053 score
Table 1; three public episodes per system.
137
Claude 3.5 Haiku 20241022 (SCIGYM)RMS without modifiers — Precision0.1454 score
Table 1; three public episodes per system.
137
Claude 3.5 Haiku 20241022 (SCIGYM)RMS without modifiers — Recall0.0805 score
Table 1; three public episodes per system.
137
Claude 3.5 Haiku 20241022 (SCIGYM)RMS without modifiers — F10.0987 score
Table 1; three public episodes per system.
137
Gemini 2.5 Pro Preview 03-25 (SCIGYM)Simulation Trajectory Error (STE)0.3212 SMAPE error
Table 1; three public episodes per system.
137
Gemini 2.5 Pro Preview 03-25 (SCIGYM)RMS with modifiers — Precision0.2138 score
Table 1; three public episodes per system.
137
Gemini 2.5 Pro Preview 03-25 (SCIGYM)RMS with modifiers — Recall0.1664 score
Table 1; three public episodes per system.
137
Gemini 2.5 Pro Preview 03-25 (SCIGYM)RMS with modifiers — F10.1817 score
Table 1; three public episodes per system.
137
Gemini 2.5 Pro Preview 03-25 (SCIGYM)RMS without modifiers — Precision0.3781 score
Table 1; three public episodes per system.
137
Gemini 2.5 Pro Preview 03-25 (SCIGYM)RMS without modifiers — Recall0.3219 score
Table 1; three public episodes per system.
137
Gemini 2.5 Pro Preview 03-25 (SCIGYM)RMS without modifiers — F10.3383 score
Table 1; three public episodes per system.
137
GPT-4.1 2025-04-14 (SCIGYM)Simulation Trajectory Error (STE)0.4611 SMAPE error
Table 1; three public episodes per system.
137
GPT-4.1 2025-04-14 (SCIGYM)RMS with modifiers — Precision0.2067 score
Table 1; three public episodes per system.
137
GPT-4.1 2025-04-14 (SCIGYM)RMS with modifiers — Recall0.1597 score
Table 1; three public episodes per system.
137
GPT-4.1 2025-04-14 (SCIGYM)RMS with modifiers — F10.174 score
Table 1; three public episodes per system.
137
GPT-4.1 2025-04-14 (SCIGYM)RMS without modifiers — Precision0.3517 score
Table 1; three public episodes per system.
137
GPT-4.1 2025-04-14 (SCIGYM)RMS without modifiers — Recall0.2888 score
Table 1; three public episodes per system.
137
GPT-4.1 2025-04-14 (SCIGYM)RMS without modifiers — F10.3038 score
Table 1; three public episodes per system.
137
Claude 3.7 Sonnet 20250219 (SCIGYM)Simulation Trajectory Error (STE)0.3615 SMAPE error
Table 1; three public episodes per system.
137
Claude 3.7 Sonnet 20250219 (SCIGYM)RMS with modifiers — Precision0.178 score
Table 1; three public episodes per system.
137
Claude 3.7 Sonnet 20250219 (SCIGYM)RMS with modifiers — Recall0.1698 score
Table 1; three public episodes per system.
137
Claude 3.7 Sonnet 20250219 (SCIGYM)RMS with modifiers — F10.1688 score
Table 1; three public episodes per system.
137
Claude 3.7 Sonnet 20250219 (SCIGYM)RMS without modifiers — Precision0.316 score
Table 1; three public episodes per system.
137
Claude 3.7 Sonnet 20250219 (SCIGYM)RMS without modifiers — Recall0.317 score
Table 1; three public episodes per system.
137
Claude 3.7 Sonnet 20250219 (SCIGYM)RMS without modifiers — F10.3047 score
Table 1; three public episodes per system.
137

Evidence

  • section: NeurIPS paper §§3.2–3.3 and 5; Appendix A (Defines the ReAct loop, three action types, public prompt, initial-concentration experiment, 20 action iterations, three debugging iterations, six exact model versions, full 137-system small scope, and NTS/RMS/STE.) — supports /benchmark_version, /scope, /model_ids, /protocol/shots, /protocol/turns, /protocol/system_prompt_public, /protocol/reasoning, /protocol/tools, /protocol/time_budget, /protocol/grader, /protocol/statistical, /protocol/contamination, /metrics, /comparability_group
  • dataset-card: data/small-00000-of-00001.parquet at commit 7d472c12855d46702c4915892578290355894c1a (2,466 rows; 137 unique systems, six exact model strings, and exactly three rows for every model-system pair.) — supports /model_ids, /protocol/repeats, /protocol/statistical
  • repository-path: README.md, scigym/llm.py, scigym/controller.py, scigym/system_prompts/, and scigym/evaluator.py at commit d290bb04bf54aad1c473c4701e0d0d88013c4f91 (Confirms exact API strings, maximum 8,192 response tokens, no browser/network/database tool, public Python/SBML tool loop, and deterministic evaluator.) — supports /model_ids, /protocol/system_prompt_public, /protocol/tools, /protocol/token_budget, /protocol/grader, /metrics
  • table: NeurIPS paper Table 1 (Prints STE and RMS precision/recall/F1 with and without modifiers for all six models; no confidence bounds are printed.) — supports /results
scigym-2025-small-zero-shot-no-toolsv2025 release

Evaluated models / systems: Claude 3.5 Haiku 20241022 (SCIGYM), Claude 3.7 Sonnet 20250219 (SCIGYM), Gemini 2.5 Flash Preview 04-17 (SCIGYM), Gemini 2.5 Pro Preview 03-25 (SCIGYM), GPT-4.1 2025-04-14 (SCIGYM), GPT-4.1 Mini 2025-04-14 (SCIGYM)

Scopefull · n=137
Shotszero-shot
Turnsone direct submission, with up to three debugging responses after invalid submissions
System prompt publicYes
Reasoning / effortdirect prompting without experimental feedback
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNo
External toolsnone
Token budgetmaximum 8,192 output tokens per model response
Time / cost budgetone direct submission plus up to 3 invalid-submission debugging iterations
TemperatureNot reported
SeedNot reported
RepeatsNot reported
Graderdeterministic SBML structural and simulation evaluator · human review: no
StatisticsFigure 5 plots per-system agent and zero-shot scores; exact aggregate baseline values and confidence intervals are not tabulated.
ContaminationThe same de-identified small systems are used; no model-training decontamination test is reported.
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Network Topology Score (NTS) F1absolutescoreper-system structural F1duplicate relationships counted once
Simulation Trajectory Error (STE)absoluteSMAPE errorspecies-level SMAPE averaged within systemoriginal and perturbed initial conditions
RMS with modifiers — F1absolutescoreexact-reaction F1reactants products and modifiers must match
RMS without modifiers — F1absolutescoreexact-reaction F1modifiers ignored

No numeric result rows are published yet; the verified protocol remains useful.

Evidence

  • figure: NeurIPS paper Figure 5 and §5.1 (Compares each of the six agents against a zero-shot/direct-prompt baseline across small systems, with experiment and tools removed and three debugging rounds; exact aggregates are not tabulated.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics, /results, /comparability_group

Comparable result views

Simulation Trajectory Error (STE)

scigym-small-creator-paper · scigym-2025-small-react-initial-concentration-20step-3repeat

CSV ↓
Accessible data table
ModelValueComparability group
Gemini 2.5 Flash Preview 04-17 (SCIGYM)0.4181scigym-2025-small-react-initial-concentration-20step-3repeat
GPT-4.1 Mini 2025-04-14 (SCIGYM)0.6007scigym-2025-small-react-initial-concentration-20step-3repeat
Claude 3.5 Haiku 20241022 (SCIGYM)0.6281scigym-2025-small-react-initial-concentration-20step-3repeat
Gemini 2.5 Pro Preview 03-25 (SCIGYM)0.3212scigym-2025-small-react-initial-concentration-20step-3repeat
GPT-4.1 2025-04-14 (SCIGYM)0.4611scigym-2025-small-react-initial-concentration-20step-3repeat
Claude 3.7 Sonnet 20250219 (SCIGYM)0.3615scigym-2025-small-react-initial-concentration-20step-3repeat

RMS with modifiers — Precision

scigym-small-creator-paper · scigym-2025-small-react-initial-concentration-20step-3repeat

CSV ↓
Accessible data table
ModelValueComparability group
Gemini 2.5 Flash Preview 04-17 (SCIGYM)0.1527scigym-2025-small-react-initial-concentration-20step-3repeat
GPT-4.1 Mini 2025-04-14 (SCIGYM)0.1516scigym-2025-small-react-initial-concentration-20step-3repeat
Claude 3.5 Haiku 20241022 (SCIGYM)0.0858scigym-2025-small-react-initial-concentration-20step-3repeat
Gemini 2.5 Pro Preview 03-25 (SCIGYM)0.2138scigym-2025-small-react-initial-concentration-20step-3repeat
GPT-4.1 2025-04-14 (SCIGYM)0.2067scigym-2025-small-react-initial-concentration-20step-3repeat
Claude 3.7 Sonnet 20250219 (SCIGYM)0.178scigym-2025-small-react-initial-concentration-20step-3repeat

RMS with modifiers — Recall

scigym-small-creator-paper · scigym-2025-small-react-initial-concentration-20step-3repeat

CSV ↓
Accessible data table
ModelValueComparability group
Gemini 2.5 Flash Preview 04-17 (SCIGYM)0.1071scigym-2025-small-react-initial-concentration-20step-3repeat
GPT-4.1 Mini 2025-04-14 (SCIGYM)0.1253scigym-2025-small-react-initial-concentration-20step-3repeat
Claude 3.5 Haiku 20241022 (SCIGYM)0.0421scigym-2025-small-react-initial-concentration-20step-3repeat
Gemini 2.5 Pro Preview 03-25 (SCIGYM)0.1664scigym-2025-small-react-initial-concentration-20step-3repeat
GPT-4.1 2025-04-14 (SCIGYM)0.1597scigym-2025-small-react-initial-concentration-20step-3repeat
Claude 3.7 Sonnet 20250219 (SCIGYM)0.1698scigym-2025-small-react-initial-concentration-20step-3repeat

RMS with modifiers — F1

scigym-small-creator-paper · scigym-2025-small-react-initial-concentration-20step-3repeat

CSV ↓
Accessible data table
ModelValueComparability group
Gemini 2.5 Flash Preview 04-17 (SCIGYM)0.1217scigym-2025-small-react-initial-concentration-20step-3repeat
GPT-4.1 Mini 2025-04-14 (SCIGYM)0.132scigym-2025-small-react-initial-concentration-20step-3repeat
Claude 3.5 Haiku 20241022 (SCIGYM)0.053scigym-2025-small-react-initial-concentration-20step-3repeat
Gemini 2.5 Pro Preview 03-25 (SCIGYM)0.1817scigym-2025-small-react-initial-concentration-20step-3repeat
GPT-4.1 2025-04-14 (SCIGYM)0.174scigym-2025-small-react-initial-concentration-20step-3repeat
Claude 3.7 Sonnet 20250219 (SCIGYM)0.1688scigym-2025-small-react-initial-concentration-20step-3repeat

RMS without modifiers — Precision

scigym-small-creator-paper · scigym-2025-small-react-initial-concentration-20step-3repeat

CSV ↓
Accessible data table
ModelValueComparability group
Gemini 2.5 Flash Preview 04-17 (SCIGYM)0.2399scigym-2025-small-react-initial-concentration-20step-3repeat
GPT-4.1 Mini 2025-04-14 (SCIGYM)0.253scigym-2025-small-react-initial-concentration-20step-3repeat
Claude 3.5 Haiku 20241022 (SCIGYM)0.1454scigym-2025-small-react-initial-concentration-20step-3repeat
Gemini 2.5 Pro Preview 03-25 (SCIGYM)0.3781scigym-2025-small-react-initial-concentration-20step-3repeat
GPT-4.1 2025-04-14 (SCIGYM)0.3517scigym-2025-small-react-initial-concentration-20step-3repeat
Claude 3.7 Sonnet 20250219 (SCIGYM)0.316scigym-2025-small-react-initial-concentration-20step-3repeat

RMS without modifiers — Recall

scigym-small-creator-paper · scigym-2025-small-react-initial-concentration-20step-3repeat

CSV ↓
Accessible data table
ModelValueComparability group
Gemini 2.5 Flash Preview 04-17 (SCIGYM)0.1839scigym-2025-small-react-initial-concentration-20step-3repeat
GPT-4.1 Mini 2025-04-14 (SCIGYM)0.2313scigym-2025-small-react-initial-concentration-20step-3repeat
Claude 3.5 Haiku 20241022 (SCIGYM)0.0805scigym-2025-small-react-initial-concentration-20step-3repeat
Gemini 2.5 Pro Preview 03-25 (SCIGYM)0.3219scigym-2025-small-react-initial-concentration-20step-3repeat
GPT-4.1 2025-04-14 (SCIGYM)0.2888scigym-2025-small-react-initial-concentration-20step-3repeat
Claude 3.7 Sonnet 20250219 (SCIGYM)0.317scigym-2025-small-react-initial-concentration-20step-3repeat

RMS without modifiers — F1

scigym-small-creator-paper · scigym-2025-small-react-initial-concentration-20step-3repeat

CSV ↓
Accessible data table
ModelValueComparability group
Gemini 2.5 Flash Preview 04-17 (SCIGYM)0.2005scigym-2025-small-react-initial-concentration-20step-3repeat
GPT-4.1 Mini 2025-04-14 (SCIGYM)0.2322scigym-2025-small-react-initial-concentration-20step-3repeat
Claude 3.5 Haiku 20241022 (SCIGYM)0.0987scigym-2025-small-react-initial-concentration-20step-3repeat
Gemini 2.5 Pro Preview 03-25 (SCIGYM)0.3383scigym-2025-small-react-initial-concentration-20step-3repeat
GPT-4.1 2025-04-14 (SCIGYM)0.3038scigym-2025-small-react-initial-concentration-20step-3repeat
Claude 3.7 Sonnet 20250219 (SCIGYM)0.3047scigym-2025-small-react-initial-concentration-20step-3repeat

Evidence and change history

Source locators remain visible; expand an item to inspect the exact Registry fields it supports.

Measuring Scientific Capabilities of Language Models with a Systems Biology Dry Lab · page: NeurIPS paper pp. 1–5, title, author affiliations, abstract, and §§1–3.2 (Defines SCIGYM, all five represented organizations, end-to-end simulated discovery, SBML inputs, ReAct agent actions, and the reconstruction task.) · Supports 9 fields

Open source →

  • /name
  • /aliases
  • /summary
  • /organizations
  • /kind
  • /domains
  • /capabilities
  • /modalities
  • /task_formats
scigym-dataset-resource · dataset-card: README.md and data/{small,large}-00000-of-00001.parquet at commit dbb10c12a33427e8e05ebbac66de690c65b6acae (The first public release commit is dated 2025-05-16; the files contain 137 and 213 unique system rows, totaling 350.) · Supports 9 fields

Open source →

  • /release_date
  • /latest_version
  • /task_counts/total
  • /task_counts/basis
  • /task_counts/subsets
  • /versions/0/task_counts
  • /versions/0/formal_tracks
  • /scientific_task_classification/entries/0
  • /scientific_task_classification/entries/1
Measuring Scientific Capabilities of Language Models with a Systems Biology Dry Lab · section: NeurIPS paper §§1–3, 5–6 and Appendix C.3 (Describes signaling, metabolic, regulatory, and epidemiological system classes; the released tasks use SBML and simulated time series, not sequence, omics, binding, or protein-design targets.) · Supports 5 fields

Open source →

  • /domains
  • /coverage_notes
  • /access/biosafety_notes
  • /scientific_task_classification/entries/0
  • /scientific_task_classification/entries/1
scigym-repository-resource · repository-path: Complete tree, README.md, pyproject.toml, scigym/prompts/, scigym/eval/, and absence of LICENSE at commit 88a7b93609e35b6ecb4eb343d816d6ff09256c6a (The environment, prompts, and grader are public, but no software license is stated. The two pinned Hugging Face cards likewise have no license field.) · Supports 7 fields

Open source →

  • /access/level
  • /access/tasks
  • /access/artifacts
  • /access/grader
  • /access/license
  • /resources
  • /implementations
scigym-results-resource · dataset-card: data/small-00000-of-00001.parquet at commit 7d472c12855d46702c4915892578290355894c1a (2,466 public rows comprise 137 systems × 6 exact models × 3 episodes; each row includes chat_history and final_model.) · Supports 1 field

Open source →

  • /access/artifacts

View source-level modification history on GitHub →