track · audited · verified 2026-07-21

Biology-Instructions Mean Ribosome Loading Prediction

Regression of mean ribosome loading from an RNA sequence, reported as squared Pearson correlation under the paper's R2 label.

Benchmark definition

What is counted

Version
emnlp-2025
Total
91519 (distinct examples across the published train, validation, and test splits)
Task formats
RNA-sequence regression
Capabilities
PredictionRegression
Modalities
TextDNA or RNA sequence

Version history

VersionStatusRelease / as-ofTotalFormal tracks
emnlp-2025
bioinstruction-mrl-emnlp-2025
current2025-11-0491519 (distinct examples across the published train, validation, and test splits)None registered

Tracks and subsets

IDCountBasisPartition?Notes
Training split
bioinstruction-mrl-train
76319examplesExclusive & exhaustivePublished training split.
Validation split
bioinstruction-mrl-validation
7600examplesExclusive & exhaustivePublished validation split.
Test split
bioinstruction-mrl-test
7600examplesExclusive & exhaustiveHeld-out split used for creator-paper Tables 4-7.

Scientific Task Atlas

Scientific task classification

complete for emnlp-2025. Single-purpose formal Biology-Instructions evaluation track.

Scientific taskCoverageCountMappingEvidence
RNA processing and translation predictionexplicitly-in-scope91519 examples
distinct examples across the published train, validation, and test splits
official-track
high confidence
bioinstruction-mrl-evidence-paper
Mean ribosome-loading prediction.

Evaluation registry

Works and run settings

A setting change—scope, prompt, tools, budget, grader, or repeats—creates a separate run. Charts never cross a comparability group.

bioinstruction-mrl-closed-baselines-emnlp-2025vemnlp-2025

Evaluated models / systems: GPT-4o (Biology-Instructions snapshot not reported), GPT-4o-mini (Biology-Instructions snapshot not reported)

Scopesubset · n=7600
Shots0
Turnssingle-turn
System prompt publicYes
Reasoning / effortThe prompt requests a direct JSON answer and no chain-of-thought.
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNot reported
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
RepeatsNot reported
Graderfirst-number extraction followed by squared Pearson correlation labeled R2 · human review: no
StatisticsPoint metric over the complete published test split, scaled by 100 and rounded to two decimals; no confidence interval is reported.
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
R2absolutepercentSquared Pearson correlation across held-out test examples, scaled by 100Not reported

Results

ModelMetricValuen
GPT-4o-mini (Biology-Instructions snapshot not reported)R20.01 percent
MRL creator-paper result; metric scaled by 100.
7600
GPT-4o (Biology-Instructions snapshot not reported)R20.01 percent
MRL creator-paper result; metric scaled by 100.
7600

Evidence

  • table: Table 2 (MRL test split); Appendix A.3; Table 8; Table 9 closed-source prompt (Scope, prompt, output parser, grader, scaling, and aggregation.) — supports /scope, /benchmark_version, /model_ids, /protocol, /metrics
  • table: Table 5 (MRL column) (All registered model values; literature-SOTA row omitted.) — supports /results
bioinstruction-mrl-creator-systems-emnlp-2025vemnlp-2025

Evaluated models / systems: ChatMultiOmics stage 1 + balanced stage 2, ChatMultiOmics stage 1 + stage 2, ChatMultiOmics, ChatMultiOmics stage 2 only

Scopesubset · n=7600
Shots0
Turnssingle-turn
System prompt publicYes
Reasoning / effortPsc requests clear, concise task answers and numeric output for regression tasks.
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNot reported
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
RepeatsNot reported
Graderfirst-number extraction followed by squared Pearson correlation labeled R2 · human review: no
StatisticsPoint metric over the complete published test split, scaled by 100 and rounded to two decimals; no confidence interval is reported.
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
R2absolutepercentSquared Pearson correlation across held-out test examples, scaled by 100Not reported

Results

ModelMetricValuen
ChatMultiOmics stage 1 + balanced stage 2R20 percent
MRL creator-paper result; metric scaled by 100.
7600
ChatMultiOmics stage 2 onlyR20 percent
MRL creator-paper result; metric scaled by 100.
7600
ChatMultiOmics stage 1 + stage 2R229.12 percent
MRL creator-paper result; metric scaled by 100.
7600
ChatMultiOmicsR247.64 percent
MRL creator-paper result; metric scaled by 100.
7600

Evidence

  • table: Table 2 (MRL test split); Appendix A.3; Table 8; Section 4.2 Psc prompt (Scope, prompt, output parser, grader, scaling, and aggregation.) — supports /scope, /benchmark_version, /model_ids, /protocol, /metrics
  • table: Table 5 (MRL column) (All registered model values; literature-SOTA row omitted.) — supports /results
bioinstruction-mrl-open-baselines-emnlp-2025vemnlp-2025

Evaluated models / systems: Alpaca-7B (Biology-Instructions label), BioMedGPT-LM-7B (Biology-Instructions label), Galactica-1.3B (Biology-Instructions label), GLM-4-9B-Chat (Biology-Instructions label), InstructProtein-1.3B (Biology-Instructions label), Llama-molinst-protein-7B (Mol-Ins), Llama2-7B-Chat (Biology-Instructions label), LLaMA3.1-8B-Instruct (Biology-Instructions label), Qwen2-7B (Biology-Instructions label), Vicuna-v1.5-7B (Biology-Instructions label)

Scopesubset · n=7600
Shots0
Turnssingle-turn
System prompt publicNo
Reasoning / effortThe prompt requests the task-formatted answer and says not to explain or repeat.
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNot reported
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
RepeatsNot reported
Graderfirst-number extraction followed by squared Pearson correlation labeled R2 · human review: no
StatisticsPoint metric over the complete published test split, scaled by 100 and rounded to two decimals; no confidence interval is reported.
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
R2absolutepercentSquared Pearson correlation across held-out test examples, scaled by 100Not reported

Results

ModelMetricValuen
LLaMA3.1-8B-Instruct (Biology-Instructions label)R20.01 percent
MRL creator-paper result; metric scaled by 100.
7600
Qwen2-7B (Biology-Instructions label)R20 percent
MRL creator-paper result; metric scaled by 100.
7600
Llama2-7B-Chat (Biology-Instructions label)R20 percent
MRL creator-paper result; metric scaled by 100.
7600
Alpaca-7B (Biology-Instructions label)R20.03 percent
MRL creator-paper result; metric scaled by 100.
7600
GLM-4-9B-Chat (Biology-Instructions label)R20 percent
MRL creator-paper result; metric scaled by 100.
7600
Vicuna-v1.5-7B (Biology-Instructions label)R20.01 percent
MRL creator-paper result; metric scaled by 100.
7600
Galactica-1.3B (Biology-Instructions label)R20 percent
MRL creator-paper result; metric scaled by 100.
7600
InstructProtein-1.3B (Biology-Instructions label)R20.02 percent
MRL creator-paper result; metric scaled by 100.
7600
Llama-molinst-protein-7B (Mol-Ins)R20 percent
MRL creator-paper result; metric scaled by 100.
7600
BioMedGPT-LM-7B (Biology-Instructions label)R20.01 percent
MRL creator-paper result; metric scaled by 100.
7600

Evidence

  • table: Table 2 (MRL test split); Appendix A.3; Table 8; Table 9 open-source prompt (Scope, prompt, output parser, grader, scaling, and aggregation.) — supports /scope, /benchmark_version, /model_ids, /protocol, /metrics
  • table: Table 5 (MRL column) (All registered model values; literature-SOTA row omitted.) — supports /results

Comparable result views

R2

bioinstruction-mrl-closed-baselines · bioinstruction-mrl-closed-baselines-emnlp-2025

CSV ↓
Accessible data table
ModelValueComparability group
GPT-4o-mini (Biology-Instructions snapshot not reported)0.01bioinstruction-mrl-closed-baselines-emnlp-2025
GPT-4o (Biology-Instructions snapshot not reported)0.01bioinstruction-mrl-closed-baselines-emnlp-2025

R2

bioinstruction-mrl-creator-systems · bioinstruction-mrl-creator-systems-emnlp-2025

CSV ↓
Accessible data table
ModelValueComparability group
ChatMultiOmics stage 1 + balanced stage 20bioinstruction-mrl-creator-systems-emnlp-2025
ChatMultiOmics stage 2 only0bioinstruction-mrl-creator-systems-emnlp-2025
ChatMultiOmics stage 1 + stage 229.12bioinstruction-mrl-creator-systems-emnlp-2025
ChatMultiOmics47.64bioinstruction-mrl-creator-systems-emnlp-2025

R2

bioinstruction-mrl-open-baselines · bioinstruction-mrl-open-baselines-emnlp-2025

CSV ↓
Accessible data table
ModelValueComparability group
LLaMA3.1-8B-Instruct (Biology-Instructions label)0.01bioinstruction-mrl-open-baselines-emnlp-2025
Qwen2-7B (Biology-Instructions label)0bioinstruction-mrl-open-baselines-emnlp-2025
Llama2-7B-Chat (Biology-Instructions label)0bioinstruction-mrl-open-baselines-emnlp-2025
Alpaca-7B (Biology-Instructions label)0.03bioinstruction-mrl-open-baselines-emnlp-2025
GLM-4-9B-Chat (Biology-Instructions label)0bioinstruction-mrl-open-baselines-emnlp-2025
Vicuna-v1.5-7B (Biology-Instructions label)0.01bioinstruction-mrl-open-baselines-emnlp-2025
Galactica-1.3B (Biology-Instructions label)0bioinstruction-mrl-open-baselines-emnlp-2025
InstructProtein-1.3B (Biology-Instructions label)0.02bioinstruction-mrl-open-baselines-emnlp-2025
Llama-molinst-protein-7B (Mol-Ins)0bioinstruction-mrl-open-baselines-emnlp-2025
BioMedGPT-LM-7B (Biology-Instructions label)0.01bioinstruction-mrl-open-baselines-emnlp-2025

Evidence and change history

Source locators remain visible; expand an item to inspect the exact Registry fields it supports.

Biology-Instructions: A Dataset and Benchmark for Multi-Omics Sequence Understanding Capability of Large Language Models · table: Table 2 (MRL row); Appendix A.2-A.3; Table 8; Tables 4-7 (Task definition, split counts, input/output format, metric, and creator evaluation.) · Supports 29 fields

Open source →

  • /name
  • /aliases
  • /summary
  • /kind
  • /parent_id
  • /organizations
  • /release_date
  • /latest_version
  • /domains
  • /capabilities
  • /modalities
  • /task_formats
  • /task_counts/total
  • /task_counts/basis
  • /task_counts/subsets
  • /access/level
  • /access/tasks
  • /access/artifacts
  • /access/grader
  • /access/license
  • /access/biosafety_notes
  • /resources
  • /implementations
  • /versions/0/release_date
  • /versions/0/as_of
  • /versions/0/task_counts/total
  • /versions/0/task_counts/basis
  • /versions/0/task_counts/subsets
  • /scientific_task_classification/entries/0
bioinstruction-mrl-repository-resource · repository-path: evaluation/evaluate.py and evaluation/register_tasks.json at 600acaa08c0302e8f5ce86de0fe041f21c13b53e (Public grader implementation, partial artifact release, and absent repository license.) · Supports 7 fields

Open source →

  • /access/level
  • /access/tasks
  • /access/artifacts
  • /access/grader
  • /access/license
  • /resources
  • /implementations

View source-level modification history on GitHub →