paper · benchmark creator

Biology-Instructions: A Dataset and Benchmark for Multi-Omics Sequence Understanding Capability of Large Language Models

Shanghai Artificial Intelligence Laboratory · University of Science and Technology of China · University of Sydney · University of Toronto · Chinese University of Hong Kong · Shanghai Jiao Tong University · Fudan University · Shanghai Innovation Institute · 2025-11-04

Relationship layer

Benchmark usage

This table records what the work did with each benchmark before attempting to normalize a run. Partial claims remain visible without being treated as comparable evaluations.

No BenchmarkUse relation is normalized for this legacy work yet. Existing EvaluationRuns remain available below.

Normalized evaluation runs

Biology-Instructions Antibody-Antigen Neutralization3 runs

Open benchmark record →

bioinstruction-aan-closed-baselines-emnlp-2025vemnlp-2025

Evaluated models / systems: GPT-4o (Biology-Instructions snapshot not reported), GPT-4o-mini (Biology-Instructions snapshot not reported)

Scopesubset · n=3301
Shots0
Turnssingle-turn
System prompt publicYes
Reasoning / effortThe prompt requests a direct JSON answer and no chain-of-thought.
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNot reported
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
RepeatsNot reported
Graderkeyword-first binary extraction with sentiment-model fallback, followed by MCC · model: cardiffnlp/twitter-roberta-base-sentiment-latest fallback · human review: no
StatisticsPoint metric over the complete published test split, scaled by 100 and rounded to two decimals; no confidence interval is reported.
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
MCCabsolutepercentMatthews correlation coefficient across held-out test examples, scaled by 100Not reported

Results

ModelMetricValuen
GPT-4o-mini (Biology-Instructions snapshot not reported)MCC1.59 percent
AAN creator-paper result; metric scaled by 100.
3301
GPT-4o (Biology-Instructions snapshot not reported)MCC-3.29 percent
AAN creator-paper result; metric scaled by 100.
3301

Evidence

  • table: Table 2 (AAN test split); Appendix A.3; Table 8; Table 9 closed-source prompt (Scope, prompt, output parser, grader, scaling, and aggregation.) — supports /scope, /benchmark_version, /model_ids, /protocol, /metrics
  • table: Table 7 (AAN column) (All registered model values; literature-SOTA row omitted.) — supports /results
bioinstruction-aan-creator-systems-emnlp-2025vemnlp-2025

Evaluated models / systems: ChatMultiOmics stage 1 + balanced stage 2, ChatMultiOmics stage 1 + stage 2, ChatMultiOmics, ChatMultiOmics stage 2 only

Scopesubset · n=3301
Shots0
Turnssingle-turn
System prompt publicYes
Reasoning / effortPsc requests clear, concise task answers and numeric output for regression tasks.
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNot reported
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
RepeatsNot reported
Graderkeyword-first binary extraction with sentiment-model fallback, followed by MCC · model: cardiffnlp/twitter-roberta-base-sentiment-latest fallback · human review: no
StatisticsPoint metric over the complete published test split, scaled by 100 and rounded to two decimals; no confidence interval is reported.
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
MCCabsolutepercentMatthews correlation coefficient across held-out test examples, scaled by 100Not reported

Results

ModelMetricValuen
ChatMultiOmics stage 1 + balanced stage 2MCC-1.48 percent
AAN creator-paper result; metric scaled by 100.
3301
ChatMultiOmics stage 2 onlyMCC0.72 percent
AAN creator-paper result; metric scaled by 100.
3301
ChatMultiOmics stage 1 + stage 2MCC10.26 percent
AAN creator-paper result; metric scaled by 100.
3301
ChatMultiOmicsMCC1.06 percent
AAN creator-paper result; metric scaled by 100.
3301

Evidence

  • table: Table 2 (AAN test split); Appendix A.3; Table 8; Section 4.2 Psc prompt (Scope, prompt, output parser, grader, scaling, and aggregation.) — supports /scope, /benchmark_version, /model_ids, /protocol, /metrics
  • table: Table 7 (AAN column) (All registered model values; literature-SOTA row omitted.) — supports /results
bioinstruction-aan-open-baselines-emnlp-2025vemnlp-2025

Evaluated models / systems: Alpaca-7B (Biology-Instructions label), BioMedGPT-LM-7B (Biology-Instructions label), Galactica-1.3B (Biology-Instructions label), GLM-4-9B-Chat (Biology-Instructions label), InstructProtein-1.3B (Biology-Instructions label), Llama-molinst-protein-7B (Mol-Ins), Llama2-7B-Chat (Biology-Instructions label), LLaMA3.1-8B-Instruct (Biology-Instructions label), Qwen2-7B (Biology-Instructions label), Vicuna-v1.5-7B (Biology-Instructions label)

Scopesubset · n=3301
Shots0
Turnssingle-turn
System prompt publicNo
Reasoning / effortThe prompt requests the task-formatted answer and says not to explain or repeat.
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNot reported
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
RepeatsNot reported
Graderkeyword-first binary extraction with sentiment-model fallback, followed by MCC · model: cardiffnlp/twitter-roberta-base-sentiment-latest fallback · human review: no
StatisticsPoint metric over the complete published test split, scaled by 100 and rounded to two decimals; no confidence interval is reported.
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
MCCabsolutepercentMatthews correlation coefficient across held-out test examples, scaled by 100Not reported

Results

ModelMetricValuen
LLaMA3.1-8B-Instruct (Biology-Instructions label)MCC-1.05 percent
AAN creator-paper result; metric scaled by 100.
3301
Qwen2-7B (Biology-Instructions label)MCC2.98 percent
AAN creator-paper result; metric scaled by 100.
3301
Llama2-7B-Chat (Biology-Instructions label)MCC-0.63 percent
AAN creator-paper result; metric scaled by 100.
3301
Alpaca-7B (Biology-Instructions label)MCC-0.81 percent
AAN creator-paper result; metric scaled by 100.
3301
GLM-4-9B-Chat (Biology-Instructions label)MCC1.32 percent
AAN creator-paper result; metric scaled by 100.
3301
Vicuna-v1.5-7B (Biology-Instructions label)MCC2 percent
AAN creator-paper result; metric scaled by 100.
3301
Galactica-1.3B (Biology-Instructions label)MCC0.01 percent
AAN creator-paper result; metric scaled by 100.
3301
InstructProtein-1.3B (Biology-Instructions label)MCC1.53 percent
AAN creator-paper result; metric scaled by 100.
3301
Llama-molinst-protein-7B (Mol-Ins)MCC-1.38 percent
AAN creator-paper result; metric scaled by 100.
3301
BioMedGPT-LM-7B (Biology-Instructions label)MCC0.92 percent
AAN creator-paper result; metric scaled by 100.
3301

Evidence

  • table: Table 2 (AAN test split); Appendix A.3; Table 8; Table 9 open-source prompt (Scope, prompt, output parser, grader, scaling, and aggregation.) — supports /scope, /benchmark_version, /model_ids, /protocol, /metrics
  • table: Table 7 (AAN column) (All registered model values; literature-SOTA row omitted.) — supports /results
Biology-Instructions APA Isoform Prediction3 runs

Open benchmark record →

bioinstruction-apa-closed-baselines-emnlp-2025vemnlp-2025

Evaluated models / systems: GPT-4o (Biology-Instructions snapshot not reported), GPT-4o-mini (Biology-Instructions snapshot not reported)

Scopesubset · n=49755
Shots0
Turnssingle-turn
System prompt publicYes
Reasoning / effortThe prompt requests a direct JSON answer and no chain-of-thought.
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNot reported
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
RepeatsNot reported
Graderfirst-number extraction followed by squared Pearson correlation labeled R2 · human review: no
StatisticsPoint metric over the complete published test split, scaled by 100 and rounded to two decimals; no confidence interval is reported.
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
R2absolutepercentSquared Pearson correlation across held-out test examples, scaled by 100Not reported

Results

ModelMetricValuen
GPT-4o-mini (Biology-Instructions snapshot not reported)R20.05 percent
APA creator-paper result; metric scaled by 100.
49755
GPT-4o (Biology-Instructions snapshot not reported)R20 percent
APA creator-paper result; metric scaled by 100.
49755

Evidence

  • table: Table 2 (APA test split); Appendix A.3; Table 8; Table 9 closed-source prompt (Scope, prompt, output parser, grader, scaling, and aggregation.) — supports /scope, /benchmark_version, /model_ids, /protocol, /metrics
  • table: Table 5 (APA column) (All registered model values; literature-SOTA row omitted.) — supports /results
bioinstruction-apa-creator-systems-emnlp-2025vemnlp-2025

Evaluated models / systems: ChatMultiOmics stage 1 + balanced stage 2, ChatMultiOmics stage 1 + stage 2, ChatMultiOmics, ChatMultiOmics stage 2 only

Scopesubset · n=49755
Shots0
Turnssingle-turn
System prompt publicYes
Reasoning / effortPsc requests clear, concise task answers and numeric output for regression tasks.
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNot reported
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
RepeatsNot reported
Graderfirst-number extraction followed by squared Pearson correlation labeled R2 · human review: no
StatisticsPoint metric over the complete published test split, scaled by 100 and rounded to two decimals; no confidence interval is reported.
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
R2absolutepercentSquared Pearson correlation across held-out test examples, scaled by 100Not reported

Results

ModelMetricValuen
ChatMultiOmics stage 1 + balanced stage 2R20.01 percent
APA creator-paper result; metric scaled by 100.
49755
ChatMultiOmics stage 2 onlyR20 percent
APA creator-paper result; metric scaled by 100.
49755
ChatMultiOmics stage 1 + stage 2R250.68 percent
APA creator-paper result; metric scaled by 100.
49755
ChatMultiOmicsR259.01 percent
APA creator-paper result; metric scaled by 100.
49755

Evidence

  • table: Table 2 (APA test split); Appendix A.3; Table 8; Section 4.2 Psc prompt (Scope, prompt, output parser, grader, scaling, and aggregation.) — supports /scope, /benchmark_version, /model_ids, /protocol, /metrics
  • table: Table 5 (APA column) (All registered model values; literature-SOTA row omitted.) — supports /results
bioinstruction-apa-open-baselines-emnlp-2025vemnlp-2025

Evaluated models / systems: Alpaca-7B (Biology-Instructions label), BioMedGPT-LM-7B (Biology-Instructions label), Galactica-1.3B (Biology-Instructions label), GLM-4-9B-Chat (Biology-Instructions label), InstructProtein-1.3B (Biology-Instructions label), Llama-molinst-protein-7B (Mol-Ins), Llama2-7B-Chat (Biology-Instructions label), LLaMA3.1-8B-Instruct (Biology-Instructions label), Qwen2-7B (Biology-Instructions label), Vicuna-v1.5-7B (Biology-Instructions label)

Scopesubset · n=49755
Shots0
Turnssingle-turn
System prompt publicNo
Reasoning / effortThe prompt requests the task-formatted answer and says not to explain or repeat.
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNot reported
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
RepeatsNot reported
Graderfirst-number extraction followed by squared Pearson correlation labeled R2 · human review: no
StatisticsPoint metric over the complete published test split, scaled by 100 and rounded to two decimals; no confidence interval is reported.
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
R2absolutepercentSquared Pearson correlation across held-out test examples, scaled by 100Not reported

Results

ModelMetricValuen
LLaMA3.1-8B-Instruct (Biology-Instructions label)R20.01 percent
APA creator-paper result; metric scaled by 100.
49755
Qwen2-7B (Biology-Instructions label)R20 percent
APA creator-paper result; metric scaled by 100.
49755
Llama2-7B-Chat (Biology-Instructions label)R20 percent
APA creator-paper result; metric scaled by 100.
49755
Alpaca-7B (Biology-Instructions label)R20 percent
APA creator-paper result; metric scaled by 100.
49755
GLM-4-9B-Chat (Biology-Instructions label)R20 percent
APA creator-paper result; metric scaled by 100.
49755
Vicuna-v1.5-7B (Biology-Instructions label)R20.01 percent
APA creator-paper result; metric scaled by 100.
49755
Galactica-1.3B (Biology-Instructions label)R20 percent
APA creator-paper result; metric scaled by 100.
49755
InstructProtein-1.3B (Biology-Instructions label)R20 percent
APA creator-paper result; metric scaled by 100.
49755
Llama-molinst-protein-7B (Mol-Ins)R20.02 percent
APA creator-paper result; metric scaled by 100.
49755
BioMedGPT-LM-7B (Biology-Instructions label)R20 percent
APA creator-paper result; metric scaled by 100.
49755

Evidence

  • table: Table 2 (APA test split); Appendix A.3; Table 8; Table 9 open-source prompt (Scope, prompt, output parser, grader, scaling, and aggregation.) — supports /scope, /benchmark_version, /model_ids, /protocol, /metrics
  • table: Table 5 (APA column) (All registered model values; literature-SOTA row omitted.) — supports /results
Biology-Instructions Core Promoter Detection3 runs

Open benchmark record →

bioinstruction-cpd-closed-baselines-emnlp-2025vemnlp-2025

Evaluated models / systems: GPT-4o (Biology-Instructions snapshot not reported), GPT-4o-mini (Biology-Instructions snapshot not reported)

Scopesubset · n=11840
Shots0
Turnssingle-turn
System prompt publicYes
Reasoning / effortThe prompt requests a direct JSON answer and no chain-of-thought.
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNot reported
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
RepeatsNot reported
Graderkeyword-first binary extraction with sentiment-model fallback, followed by MCC · model: cardiffnlp/twitter-roberta-base-sentiment-latest fallback · human review: no
StatisticsPoint metric over the complete published test split, scaled by 100 and rounded to two decimals; no confidence interval is reported.
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
MCCabsolutepercentMatthews correlation coefficient across held-out test examples, scaled by 100Not reported

Results

ModelMetricValuen
GPT-4o-mini (Biology-Instructions snapshot not reported)MCC-2.95 percent
CPD creator-paper result; metric scaled by 100.
11840
GPT-4o (Biology-Instructions snapshot not reported)MCC-0.84 percent
CPD creator-paper result; metric scaled by 100.
11840

Evidence

  • table: Table 2 (CPD test split); Appendix A.3; Table 8; Table 9 closed-source prompt (Scope, prompt, output parser, grader, scaling, and aggregation.) — supports /scope, /benchmark_version, /model_ids, /protocol, /metrics
  • table: Table 4 (CPD column) (All registered model values; literature-SOTA row omitted.) — supports /results
bioinstruction-cpd-creator-systems-emnlp-2025vemnlp-2025

Evaluated models / systems: ChatMultiOmics stage 1 + balanced stage 2, ChatMultiOmics stage 1 + stage 2, ChatMultiOmics, ChatMultiOmics stage 2 only

Scopesubset · n=11840
Shots0
Turnssingle-turn
System prompt publicYes
Reasoning / effortPsc requests clear, concise task answers and numeric output for regression tasks.
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNot reported
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
RepeatsNot reported
Graderkeyword-first binary extraction with sentiment-model fallback, followed by MCC · model: cardiffnlp/twitter-roberta-base-sentiment-latest fallback · human review: no
StatisticsPoint metric over the complete published test split, scaled by 100 and rounded to two decimals; no confidence interval is reported.
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
MCCabsolutepercentMatthews correlation coefficient across held-out test examples, scaled by 100Not reported

Results

ModelMetricValuen
ChatMultiOmics stage 1 + balanced stage 2MCC5.57 percent
CPD creator-paper result; metric scaled by 100.
11840
ChatMultiOmics stage 2 onlyMCC1.8 percent
CPD creator-paper result; metric scaled by 100.
11840
ChatMultiOmics stage 1 + stage 2MCC41.18 percent
CPD creator-paper result; metric scaled by 100.
11840
ChatMultiOmicsMCC44.54 percent
CPD creator-paper result; metric scaled by 100.
11840

Evidence

  • table: Table 2 (CPD test split); Appendix A.3; Table 8; Section 4.2 Psc prompt (Scope, prompt, output parser, grader, scaling, and aggregation.) — supports /scope, /benchmark_version, /model_ids, /protocol, /metrics
  • table: Table 4 (CPD column) (All registered model values; literature-SOTA row omitted.) — supports /results
bioinstruction-cpd-open-baselines-emnlp-2025vemnlp-2025

Evaluated models / systems: Alpaca-7B (Biology-Instructions label), Galactica-1.3B (Biology-Instructions label), GLM-4-9B-Chat (Biology-Instructions label), InstructProtein-1.3B (Biology-Instructions label), Llama-molinst-protein-7B (Mol-Ins), Llama2-7B-Chat (Biology-Instructions label), LLaMA3.1-8B-Instruct (Biology-Instructions label), Qwen2-7B (Biology-Instructions label), Vicuna-v1.5-7B (Biology-Instructions label)

Scopesubset · n=11840
Shots0
Turnssingle-turn
System prompt publicNo
Reasoning / effortThe prompt requests the task-formatted answer and says not to explain or repeat.
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNot reported
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
RepeatsNot reported
Graderkeyword-first binary extraction with sentiment-model fallback, followed by MCC · model: cardiffnlp/twitter-roberta-base-sentiment-latest fallback · human review: no
StatisticsPoint metric over the complete published test split, scaled by 100 and rounded to two decimals; no confidence interval is reported.
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
MCCabsolutepercentMatthews correlation coefficient across held-out test examples, scaled by 100Not reported

Results

ModelMetricValuen
LLaMA3.1-8B-Instruct (Biology-Instructions label)MCC0 percent
CPD creator-paper result; metric scaled by 100.
11840
Qwen2-7B (Biology-Instructions label)MCC1.35 percent
CPD creator-paper result; metric scaled by 100.
11840
Llama2-7B-Chat (Biology-Instructions label)MCC-0.55 percent
CPD creator-paper result; metric scaled by 100.
11840
Alpaca-7B (Biology-Instructions label)MCC-1.3 percent
CPD creator-paper result; metric scaled by 100.
11840
GLM-4-9B-Chat (Biology-Instructions label)MCC-2.53 percent
CPD creator-paper result; metric scaled by 100.
11840
Vicuna-v1.5-7B (Biology-Instructions label)MCC0 percent
CPD creator-paper result; metric scaled by 100.
11840
Galactica-1.3B (Biology-Instructions label)MCC-1.01 percent
CPD creator-paper result; metric scaled by 100.
11840
InstructProtein-1.3B (Biology-Instructions label)MCC-0.33 percent
CPD creator-paper result; metric scaled by 100.
11840
Llama-molinst-protein-7B (Mol-Ins)MCC1.98 percent
CPD creator-paper result; metric scaled by 100.
11840

Evidence

  • table: Table 2 (CPD test split); Appendix A.3; Table 8; Table 9 open-source prompt (Scope, prompt, output parser, grader, scaling, and aggregation.) — supports /scope, /benchmark_version, /model_ids, /protocol, /metrics
  • table: Table 4 (CPD column) (All registered model values; literature-SOTA row omitted.) — supports /results
Biology-Instructions CRISPR On-Target Prediction3 runs

Open benchmark record →

bioinstruction-crispr-on-target-closed-baselines-emnlp-2025vemnlp-2025

Evaluated models / systems: GPT-4o (Biology-Instructions snapshot not reported), GPT-4o-mini (Biology-Instructions snapshot not reported)

Scopesubset · n=416
Shots0
Turnssingle-turn
System prompt publicYes
Reasoning / effortThe prompt requests a direct JSON answer and no chain-of-thought.
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNot reported
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
RepeatsNot reported
Graderfirst-number extraction followed by Spearman rank correlation · human review: no
StatisticsPoint metric over the complete published test split, scaled by 100 and rounded to two decimals; no confidence interval is reported.
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Spearman's ρabsolutepercentSpearman rank correlation across held-out test examples, scaled by 100Not reported

Results

ModelMetricValuen
GPT-4o-mini (Biology-Instructions snapshot not reported)Spearman's ρ3.77 percent
CRI-On creator-paper result; metric scaled by 100.
416
GPT-4o (Biology-Instructions snapshot not reported)Spearman's ρ-3.31 percent
CRI-On creator-paper result; metric scaled by 100.
416

Evidence

  • table: Table 2 (CRI-On test split); Appendix A.3; Table 8; Table 9 closed-source prompt (Scope, prompt, output parser, grader, scaling, and aggregation.) — supports /scope, /benchmark_version, /model_ids, /protocol, /metrics
  • table: Table 5 (CRI-On column) (All registered model values; literature-SOTA row omitted.) — supports /results
bioinstruction-crispr-on-target-creator-systems-emnlp-2025vemnlp-2025

Evaluated models / systems: ChatMultiOmics stage 1 + balanced stage 2, ChatMultiOmics stage 1 + stage 2, ChatMultiOmics, ChatMultiOmics stage 2 only

Scopesubset · n=416
Shots0
Turnssingle-turn
System prompt publicYes
Reasoning / effortPsc requests clear, concise task answers and numeric output for regression tasks.
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNot reported
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
RepeatsNot reported
Graderfirst-number extraction followed by Spearman rank correlation · human review: no
StatisticsPoint metric over the complete published test split, scaled by 100 and rounded to two decimals; no confidence interval is reported.
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Spearman's ρabsolutepercentSpearman rank correlation across held-out test examples, scaled by 100Not reported

Results

ModelMetricValuen
ChatMultiOmics stage 1 + balanced stage 2Spearman's ρ-0.31 percent
CRI-On creator-paper result; metric scaled by 100.
416
ChatMultiOmics stage 2 onlySpearman's ρ2.87 percent
CRI-On creator-paper result; metric scaled by 100.
416
ChatMultiOmics stage 1 + stage 2Spearman's ρ-2.99 percent
CRI-On creator-paper result; metric scaled by 100.
416
ChatMultiOmicsSpearman's ρ-0.02 percent
CRI-On creator-paper result; metric scaled by 100.
416

Evidence

  • table: Table 2 (CRI-On test split); Appendix A.3; Table 8; Section 4.2 Psc prompt (Scope, prompt, output parser, grader, scaling, and aggregation.) — supports /scope, /benchmark_version, /model_ids, /protocol, /metrics
  • table: Table 5 (CRI-On column) (All registered model values; literature-SOTA row omitted.) — supports /results
bioinstruction-crispr-on-target-open-baselines-emnlp-2025vemnlp-2025

Evaluated models / systems: Alpaca-7B (Biology-Instructions label), BioMedGPT-LM-7B (Biology-Instructions label), Galactica-1.3B (Biology-Instructions label), GLM-4-9B-Chat (Biology-Instructions label), InstructProtein-1.3B (Biology-Instructions label), Llama-molinst-protein-7B (Mol-Ins), Llama2-7B-Chat (Biology-Instructions label), LLaMA3.1-8B-Instruct (Biology-Instructions label), Qwen2-7B (Biology-Instructions label), Vicuna-v1.5-7B (Biology-Instructions label)

Scopesubset · n=416
Shots0
Turnssingle-turn
System prompt publicNo
Reasoning / effortThe prompt requests the task-formatted answer and says not to explain or repeat.
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNot reported
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
RepeatsNot reported
Graderfirst-number extraction followed by Spearman rank correlation · human review: no
StatisticsPoint metric over the complete published test split, scaled by 100 and rounded to two decimals; no confidence interval is reported.
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Spearman's ρabsolutepercentSpearman rank correlation across held-out test examples, scaled by 100Not reported

Results

ModelMetricValuen
LLaMA3.1-8B-Instruct (Biology-Instructions label)Spearman's ρ-0.09 percent
CRI-On creator-paper result; metric scaled by 100.
416
Qwen2-7B (Biology-Instructions label)Spearman's ρ-6.21 percent
CRI-On creator-paper result; metric scaled by 100.
416
Llama2-7B-Chat (Biology-Instructions label)Spearman's ρ0.92 percent
CRI-On creator-paper result; metric scaled by 100.
416
Alpaca-7B (Biology-Instructions label)Spearman's ρ-3.55 percent
CRI-On creator-paper result; metric scaled by 100.
416
GLM-4-9B-Chat (Biology-Instructions label)Spearman's ρ-0.02 percent
CRI-On creator-paper result; metric scaled by 100.
416
Vicuna-v1.5-7B (Biology-Instructions label)Spearman's ρ1.88 percent
CRI-On creator-paper result; metric scaled by 100.
416
Galactica-1.3B (Biology-Instructions label)Spearman's ρ-5.56 percent
CRI-On creator-paper result; metric scaled by 100.
416
InstructProtein-1.3B (Biology-Instructions label)Spearman's ρ0 percent
CRI-On creator-paper result; metric scaled by 100.
416
Llama-molinst-protein-7B (Mol-Ins)Spearman's ρ-0.1 percent
CRI-On creator-paper result; metric scaled by 100.
416
BioMedGPT-LM-7B (Biology-Instructions label)Spearman's ρ0.12 percent
CRI-On creator-paper result; metric scaled by 100.
416

Evidence

  • table: Table 2 (CRI-On test split); Appendix A.3; Table 8; Table 9 open-source prompt (Scope, prompt, output parser, grader, scaling, and aggregation.) — supports /scope, /benchmark_version, /model_ids, /protocol, /metrics
  • table: Table 5 (CRI-On column) (All registered model values; literature-SOTA row omitted.) — supports /results
Biology-Instructions Enhancer Activity Prediction3 runs

Open benchmark record →

bioinstruction-ea-closed-baselines-emnlp-2025vemnlp-2025

Evaluated models / systems: GPT-4o (Biology-Instructions snapshot not reported), GPT-4o-mini (Biology-Instructions snapshot not reported)

Scopesubset · n=41186
Shots0
Turnssingle-turn
System prompt publicYes
Reasoning / effortThe prompt requests a direct JSON answer and no chain-of-thought.
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNot reported
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
RepeatsNot reported
Gradertwo-number extraction and separate Pearson correlations · human review: no
StatisticsPoint metric over the complete published test split, scaled by 100 and rounded to two decimals; no confidence interval is reported.
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
PCC (housekeeping enhancer activity)absolutepercentPearson correlation across held-out test examples, scaled by 100Not reported
PCC (developmental enhancer activity)absolutepercentPearson correlation across held-out test examples, scaled by 100Not reported

Results

ModelMetricValuen
GPT-4o-mini (Biology-Instructions snapshot not reported)PCC (housekeeping enhancer activity)-0.76 percent
EA creator-paper result; metric scaled by 100.
41186
GPT-4o-mini (Biology-Instructions snapshot not reported)PCC (developmental enhancer activity)0.09 percent
EA creator-paper result; metric scaled by 100.
41186
GPT-4o (Biology-Instructions snapshot not reported)PCC (housekeeping enhancer activity)-1.17 percent
EA creator-paper result; metric scaled by 100.
41186
GPT-4o (Biology-Instructions snapshot not reported)PCC (developmental enhancer activity)-1.49 percent
EA creator-paper result; metric scaled by 100.
41186

Evidence

  • table: Table 2 (EA test split); Appendix A.3; Table 8; Table 9 closed-source prompt (Scope, prompt, output parser, grader, scaling, and aggregation.) — supports /scope, /benchmark_version, /model_ids, /protocol, /metrics
  • table: Table 4 (EA column) (All registered model values; literature-SOTA row omitted.) — supports /results
bioinstruction-ea-creator-systems-emnlp-2025vemnlp-2025

Evaluated models / systems: ChatMultiOmics stage 1 + balanced stage 2, ChatMultiOmics stage 1 + stage 2, ChatMultiOmics, ChatMultiOmics stage 2 only

Scopesubset · n=41186
Shots0
Turnssingle-turn
System prompt publicYes
Reasoning / effortPsc requests clear, concise task answers and numeric output for regression tasks.
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNot reported
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
RepeatsNot reported
Gradertwo-number extraction and separate Pearson correlations · human review: no
StatisticsPoint metric over the complete published test split, scaled by 100 and rounded to two decimals; no confidence interval is reported.
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
PCC (housekeeping enhancer activity)absolutepercentPearson correlation across held-out test examples, scaled by 100Not reported
PCC (developmental enhancer activity)absolutepercentPearson correlation across held-out test examples, scaled by 100Not reported

Results

ModelMetricValuen
ChatMultiOmics stage 1 + balanced stage 2PCC (housekeeping enhancer activity)0.92 percent
EA creator-paper result; metric scaled by 100.
41186
ChatMultiOmics stage 1 + balanced stage 2PCC (developmental enhancer activity)0.06 percent
EA creator-paper result; metric scaled by 100.
41186
ChatMultiOmics stage 2 onlyPCC (housekeeping enhancer activity)-0.16 percent
EA creator-paper result; metric scaled by 100.
41186
ChatMultiOmics stage 2 onlyPCC (developmental enhancer activity)0.08 percent
EA creator-paper result; metric scaled by 100.
41186
ChatMultiOmics stage 1 + stage 2PCC (housekeeping enhancer activity)59.74 percent
EA creator-paper result; metric scaled by 100.
41186
ChatMultiOmics stage 1 + stage 2PCC (developmental enhancer activity)46.82 percent
EA creator-paper result; metric scaled by 100.
41186
ChatMultiOmicsPCC (housekeeping enhancer activity)57.24 percent
EA creator-paper result; metric scaled by 100.
41186
ChatMultiOmicsPCC (developmental enhancer activity)45.92 percent
EA creator-paper result; metric scaled by 100.
41186

Evidence

  • table: Table 2 (EA test split); Appendix A.3; Table 8; Section 4.2 Psc prompt (Scope, prompt, output parser, grader, scaling, and aggregation.) — supports /scope, /benchmark_version, /model_ids, /protocol, /metrics
  • table: Table 4 (EA column) (All registered model values; literature-SOTA row omitted.) — supports /results
bioinstruction-ea-open-baselines-emnlp-2025vemnlp-2025

Evaluated models / systems: Alpaca-7B (Biology-Instructions label), Galactica-1.3B (Biology-Instructions label), GLM-4-9B-Chat (Biology-Instructions label), InstructProtein-1.3B (Biology-Instructions label), Llama-molinst-protein-7B (Mol-Ins), Llama2-7B-Chat (Biology-Instructions label), LLaMA3.1-8B-Instruct (Biology-Instructions label), Qwen2-7B (Biology-Instructions label), Vicuna-v1.5-7B (Biology-Instructions label)

Scopesubset · n=41186
Shots0
Turnssingle-turn
System prompt publicNo
Reasoning / effortThe prompt requests the task-formatted answer and says not to explain or repeat.
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNot reported
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
RepeatsNot reported
Gradertwo-number extraction and separate Pearson correlations · human review: no
StatisticsPoint metric over the complete published test split, scaled by 100 and rounded to two decimals; no confidence interval is reported.
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
PCC (housekeeping enhancer activity)absolutepercentPearson correlation across held-out test examples, scaled by 100Not reported
PCC (developmental enhancer activity)absolutepercentPearson correlation across held-out test examples, scaled by 100Not reported

Results

ModelMetricValuen
LLaMA3.1-8B-Instruct (Biology-Instructions label)PCC (housekeeping enhancer activity)0.61 percent
EA creator-paper result; metric scaled by 100.
41186
LLaMA3.1-8B-Instruct (Biology-Instructions label)PCC (developmental enhancer activity)0.27 percent
EA creator-paper result; metric scaled by 100.
41186
Qwen2-7B (Biology-Instructions label)PCC (housekeeping enhancer activity)0.4 percent
EA creator-paper result; metric scaled by 100.
41186
Qwen2-7B (Biology-Instructions label)PCC (developmental enhancer activity)0.35 percent
EA creator-paper result; metric scaled by 100.
41186
Llama2-7B-Chat (Biology-Instructions label)PCC (housekeeping enhancer activity)0.55 percent
EA creator-paper result; metric scaled by 100.
41186
Llama2-7B-Chat (Biology-Instructions label)PCC (developmental enhancer activity)0.13 percent
EA creator-paper result; metric scaled by 100.
41186
Alpaca-7B (Biology-Instructions label)PCC (housekeeping enhancer activity)-0.11 percent
EA creator-paper result; metric scaled by 100.
41186
Alpaca-7B (Biology-Instructions label)PCC (developmental enhancer activity)0.31 percent
EA creator-paper result; metric scaled by 100.
41186
GLM-4-9B-Chat (Biology-Instructions label)PCC (housekeeping enhancer activity)0.87 percent
EA creator-paper result; metric scaled by 100.
41186
GLM-4-9B-Chat (Biology-Instructions label)PCC (developmental enhancer activity)0.17 percent
EA creator-paper result; metric scaled by 100.
41186
Vicuna-v1.5-7B (Biology-Instructions label)PCC (housekeeping enhancer activity)0.18 percent
EA creator-paper result; metric scaled by 100.
41186
Vicuna-v1.5-7B (Biology-Instructions label)PCC (developmental enhancer activity)0.69 percent
EA creator-paper result; metric scaled by 100.
41186
Galactica-1.3B (Biology-Instructions label)PCC (housekeeping enhancer activity)0.13 percent
EA creator-paper result; metric scaled by 100.
41186
Galactica-1.3B (Biology-Instructions label)PCC (developmental enhancer activity)0.09 percent
EA creator-paper result; metric scaled by 100.
41186
InstructProtein-1.3B (Biology-Instructions label)PCC (housekeeping enhancer activity)0 percent
EA creator-paper result; metric scaled by 100.
41186
InstructProtein-1.3B (Biology-Instructions label)PCC (developmental enhancer activity)0.39 percent
EA creator-paper result; metric scaled by 100.
41186
Llama-molinst-protein-7B (Mol-Ins)PCC (housekeeping enhancer activity)0.02 percent
EA creator-paper result; metric scaled by 100.
41186
Llama-molinst-protein-7B (Mol-Ins)PCC (developmental enhancer activity)0.1 percent
EA creator-paper result; metric scaled by 100.
41186

Evidence

  • table: Table 2 (EA test split); Appendix A.3; Table 8; Table 9 open-source prompt (Scope, prompt, output parser, grader, scaling, and aggregation.) — supports /scope, /benchmark_version, /model_ids, /protocol, /metrics
  • table: Table 4 (EA column) (All registered model values; literature-SOTA row omitted.) — supports /results
Biology-Instructions Enzyme Commission Number Prediction3 runs

Open benchmark record →

bioinstruction-ec-closed-baselines-emnlp-2025vemnlp-2025

Evaluated models / systems: GPT-4o (Biology-Instructions snapshot not reported), GPT-4o-mini (Biology-Instructions snapshot not reported)

Scopesubset · n=1919
Shots0
Turnssingle-turn
System prompt publicYes
Reasoning / effortThe prompt requests a direct JSON answer and no chain-of-thought.
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNot reported
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
RepeatsNot reported
GraderEC-number regex and multi-hot conversion followed by the creator Fmax implementation · human review: no
StatisticsPoint metric over the complete published test split, scaled by 100 and rounded to two decimals; no confidence interval is reported.
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
FmaxabsolutepercentCreator Fmax implementation across held-out test examples, scaled by 100Not reported

Results

ModelMetricValuen
GPT-4o-mini (Biology-Instructions snapshot not reported)Fmax1.73 percent
EC creator-paper result; metric scaled by 100.
1919
GPT-4o (Biology-Instructions snapshot not reported)Fmax5.89 percent
EC creator-paper result; metric scaled by 100.
1919

Evidence

  • table: Table 2 (EC test split); Appendix A.3; Table 8; Table 9 closed-source prompt (Scope, prompt, output parser, grader, scaling, and aggregation.) — supports /scope, /benchmark_version, /model_ids, /protocol, /metrics
  • table: Table 6 (EC column) (All registered model values; literature-SOTA row omitted.) — supports /results
bioinstruction-ec-creator-systems-emnlp-2025vemnlp-2025

Evaluated models / systems: ChatMultiOmics stage 1 + balanced stage 2, ChatMultiOmics stage 1 + stage 2, ChatMultiOmics, ChatMultiOmics stage 2 only

Scopesubset · n=1919
Shots0
Turnssingle-turn
System prompt publicYes
Reasoning / effortPsc requests clear, concise task answers and numeric output for regression tasks.
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNot reported
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
RepeatsNot reported
GraderEC-number regex and multi-hot conversion followed by the creator Fmax implementation · human review: no
StatisticsPoint metric over the complete published test split, scaled by 100 and rounded to two decimals; no confidence interval is reported.
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
FmaxabsolutepercentCreator Fmax implementation across held-out test examples, scaled by 100Not reported

Results

ModelMetricValuen
ChatMultiOmics stage 1 + balanced stage 2Fmax10.76 percent
EC creator-paper result; metric scaled by 100.
1919
ChatMultiOmics stage 2 onlyFmax1.85 percent
EC creator-paper result; metric scaled by 100.
1919
ChatMultiOmics stage 1 + stage 2Fmax19.35 percent
EC creator-paper result; metric scaled by 100.
1919
ChatMultiOmicsFmax19.79 percent
EC creator-paper result; metric scaled by 100.
1919

Evidence

  • table: Table 2 (EC test split); Appendix A.3; Table 8; Section 4.2 Psc prompt (Scope, prompt, output parser, grader, scaling, and aggregation.) — supports /scope, /benchmark_version, /model_ids, /protocol, /metrics
  • table: Table 6 (EC column) (All registered model values; literature-SOTA row omitted.) — supports /results
bioinstruction-ec-open-baselines-emnlp-2025vemnlp-2025

Evaluated models / systems: Alpaca-7B (Biology-Instructions label), BioMedGPT-LM-7B (Biology-Instructions label), Galactica-1.3B (Biology-Instructions label), GLM-4-9B-Chat (Biology-Instructions label), InstructProtein-1.3B (Biology-Instructions label), Llama-molinst-protein-7B (Mol-Ins), Llama2-7B-Chat (Biology-Instructions label), LLaMA3.1-8B-Instruct (Biology-Instructions label), Qwen2-7B (Biology-Instructions label), Vicuna-v1.5-7B (Biology-Instructions label)

Scopesubset · n=1919
Shots0
Turnssingle-turn
System prompt publicNo
Reasoning / effortThe prompt requests the task-formatted answer and says not to explain or repeat.
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNot reported
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
RepeatsNot reported
GraderEC-number regex and multi-hot conversion followed by the creator Fmax implementation · human review: no
StatisticsPoint metric over the complete published test split, scaled by 100 and rounded to two decimals; no confidence interval is reported.
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
FmaxabsolutepercentCreator Fmax implementation across held-out test examples, scaled by 100Not reported

Results

ModelMetricValuen
LLaMA3.1-8B-Instruct (Biology-Instructions label)Fmax1.42 percent
EC creator-paper result; metric scaled by 100.
1919
Qwen2-7B (Biology-Instructions label)Fmax0.9 percent
EC creator-paper result; metric scaled by 100.
1919
Llama2-7B-Chat (Biology-Instructions label)Fmax0.97 percent
EC creator-paper result; metric scaled by 100.
1919
Alpaca-7B (Biology-Instructions label)Fmax0.88 percent
EC creator-paper result; metric scaled by 100.
1919
GLM-4-9B-Chat (Biology-Instructions label)Fmax0.91 percent
EC creator-paper result; metric scaled by 100.
1919
Vicuna-v1.5-7B (Biology-Instructions label)Fmax0.88 percent
EC creator-paper result; metric scaled by 100.
1919
Galactica-1.3B (Biology-Instructions label)Fmax0.91 percent
EC creator-paper result; metric scaled by 100.
1919
InstructProtein-1.3B (Biology-Instructions label)Fmax1.85 percent
EC creator-paper result; metric scaled by 100.
1919
Llama-molinst-protein-7B (Mol-Ins)Fmax1.85 percent
EC creator-paper result; metric scaled by 100.
1919
BioMedGPT-LM-7B (Biology-Instructions label)Fmax1.07 percent
EC creator-paper result; metric scaled by 100.
1919

Evidence

  • table: Table 2 (EC test split); Appendix A.3; Table 8; Table 9 open-source prompt (Scope, prompt, output parser, grader, scaling, and aggregation.) — supports /scope, /benchmark_version, /model_ids, /protocol, /metrics
  • table: Table 6 (EC column) (All registered model values; literature-SOTA row omitted.) — supports /results
Biology-Instructions Epigenetic Marks Prediction3 runs

Open benchmark record →

bioinstruction-emp-closed-baselines-emnlp-2025vemnlp-2025

Evaluated models / systems: GPT-4o (Biology-Instructions snapshot not reported), GPT-4o-mini (Biology-Instructions snapshot not reported)

Scopesubset · n=28741
Shots0
Turnssingle-turn
System prompt publicYes
Reasoning / effortThe prompt requests a direct JSON answer and no chain-of-thought.
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNot reported
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
RepeatsNot reported
Graderkeyword-first binary extraction with sentiment-model fallback, followed by MCC · model: cardiffnlp/twitter-roberta-base-sentiment-latest fallback · human review: no
StatisticsPoint metric over the complete published test split, scaled by 100 and rounded to two decimals; no confidence interval is reported.
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
MCCabsolutepercentMatthews correlation coefficient across held-out test examples, scaled by 100Not reported

Results

ModelMetricValuen
GPT-4o-mini (Biology-Instructions snapshot not reported)MCC-0.91 percent
EMP creator-paper result; metric scaled by 100.
28741
GPT-4o (Biology-Instructions snapshot not reported)MCC-0.49 percent
EMP creator-paper result; metric scaled by 100.
28741

Evidence

  • table: Table 2 (EMP test split); Appendix A.3; Table 8; Table 9 closed-source prompt (Scope, prompt, output parser, grader, scaling, and aggregation.) — supports /scope, /benchmark_version, /model_ids, /protocol, /metrics
  • table: Table 4 (EMP column) (All registered model values; literature-SOTA row omitted.) — supports /results
bioinstruction-emp-creator-systems-emnlp-2025vemnlp-2025

Evaluated models / systems: ChatMultiOmics stage 1 + balanced stage 2, ChatMultiOmics stage 1 + stage 2, ChatMultiOmics, ChatMultiOmics stage 2 only

Scopesubset · n=28741
Shots0
Turnssingle-turn
System prompt publicYes
Reasoning / effortPsc requests clear, concise task answers and numeric output for regression tasks.
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNot reported
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
RepeatsNot reported
Graderkeyword-first binary extraction with sentiment-model fallback, followed by MCC · model: cardiffnlp/twitter-roberta-base-sentiment-latest fallback · human review: no
StatisticsPoint metric over the complete published test split, scaled by 100 and rounded to two decimals; no confidence interval is reported.
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
MCCabsolutepercentMatthews correlation coefficient across held-out test examples, scaled by 100Not reported

Results

ModelMetricValuen
ChatMultiOmics stage 1 + balanced stage 2MCC1.4 percent
EMP creator-paper result; metric scaled by 100.
28741
ChatMultiOmics stage 2 onlyMCC0.31 percent
EMP creator-paper result; metric scaled by 100.
28741
ChatMultiOmics stage 1 + stage 2MCC8.1 percent
EMP creator-paper result; metric scaled by 100.
28741
ChatMultiOmicsMCC3.64 percent
EMP creator-paper result; metric scaled by 100.
28741

Evidence

  • table: Table 2 (EMP test split); Appendix A.3; Table 8; Section 4.2 Psc prompt (Scope, prompt, output parser, grader, scaling, and aggregation.) — supports /scope, /benchmark_version, /model_ids, /protocol, /metrics
  • table: Table 4 (EMP column) (All registered model values; literature-SOTA row omitted.) — supports /results
bioinstruction-emp-open-baselines-emnlp-2025vemnlp-2025

Evaluated models / systems: Alpaca-7B (Biology-Instructions label), Galactica-1.3B (Biology-Instructions label), GLM-4-9B-Chat (Biology-Instructions label), InstructProtein-1.3B (Biology-Instructions label), Llama-molinst-protein-7B (Mol-Ins), Llama2-7B-Chat (Biology-Instructions label), LLaMA3.1-8B-Instruct (Biology-Instructions label), Qwen2-7B (Biology-Instructions label), Vicuna-v1.5-7B (Biology-Instructions label)

Scopesubset · n=28741
Shots0
Turnssingle-turn
System prompt publicNo
Reasoning / effortThe prompt requests the task-formatted answer and says not to explain or repeat.
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNot reported
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
RepeatsNot reported
Graderkeyword-first binary extraction with sentiment-model fallback, followed by MCC · model: cardiffnlp/twitter-roberta-base-sentiment-latest fallback · human review: no
StatisticsPoint metric over the complete published test split, scaled by 100 and rounded to two decimals; no confidence interval is reported.
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
MCCabsolutepercentMatthews correlation coefficient across held-out test examples, scaled by 100Not reported

Results

ModelMetricValuen
LLaMA3.1-8B-Instruct (Biology-Instructions label)MCC-0.37 percent
EMP creator-paper result; metric scaled by 100.
28741
Qwen2-7B (Biology-Instructions label)MCC-0.66 percent
EMP creator-paper result; metric scaled by 100.
28741
Llama2-7B-Chat (Biology-Instructions label)MCC0.94 percent
EMP creator-paper result; metric scaled by 100.
28741
Alpaca-7B (Biology-Instructions label)MCC-0.36 percent
EMP creator-paper result; metric scaled by 100.
28741
GLM-4-9B-Chat (Biology-Instructions label)MCC-0.22 percent
EMP creator-paper result; metric scaled by 100.
28741
Vicuna-v1.5-7B (Biology-Instructions label)MCC0 percent
EMP creator-paper result; metric scaled by 100.
28741
Galactica-1.3B (Biology-Instructions label)MCC0.07 percent
EMP creator-paper result; metric scaled by 100.
28741
InstructProtein-1.3B (Biology-Instructions label)MCC0.22 percent
EMP creator-paper result; metric scaled by 100.
28741
Llama-molinst-protein-7B (Mol-Ins)MCC-0.29 percent
EMP creator-paper result; metric scaled by 100.
28741

Evidence

  • table: Table 2 (EMP test split); Appendix A.3; Table 8; Table 9 open-source prompt (Scope, prompt, output parser, grader, scaling, and aggregation.) — supports /scope, /benchmark_version, /model_ids, /protocol, /metrics
  • table: Table 4 (EMP column) (All registered model values; literature-SOTA row omitted.) — supports /results
Biology-Instructions Enhancer-Promoter Interaction Prediction3 runs

Open benchmark record →

bioinstruction-epi-closed-baselines-emnlp-2025vemnlp-2025

Evaluated models / systems: GPT-4o (Biology-Instructions snapshot not reported), GPT-4o-mini (Biology-Instructions snapshot not reported)

Scopesubset · n=308
Shots0
Turnssingle-turn
System prompt publicYes
Reasoning / effortThe prompt requests a direct JSON answer and no chain-of-thought.
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNot reported
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
RepeatsNot reported
Graderkeyword-first binary extraction with sentiment-model fallback, followed by MCC · model: cardiffnlp/twitter-roberta-base-sentiment-latest fallback · human review: no
StatisticsPoint metric over the complete published test split, scaled by 100 and rounded to two decimals; no confidence interval is reported.
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
MCCabsolutepercentMatthews correlation coefficient across held-out test examples, scaled by 100Not reported

Results

ModelMetricValuen
GPT-4o-mini (Biology-Instructions snapshot not reported)MCC-0.39 percent
EPI creator-paper result; metric scaled by 100.
308
GPT-4o (Biology-Instructions snapshot not reported)MCC0 percent
EPI creator-paper result; metric scaled by 100.
308

Evidence

  • table: Table 2 (EPI test split); Appendix A.3; Table 8; Table 9 closed-source prompt (Scope, prompt, output parser, grader, scaling, and aggregation.) — supports /scope, /benchmark_version, /model_ids, /protocol, /metrics
  • table: Table 7 (EPI column) (All registered model values; literature-SOTA row omitted.) — supports /results
bioinstruction-epi-creator-systems-emnlp-2025vemnlp-2025

Evaluated models / systems: ChatMultiOmics stage 1 + balanced stage 2, ChatMultiOmics stage 1 + stage 2, ChatMultiOmics, ChatMultiOmics stage 2 only

Scopesubset · n=308
Shots0
Turnssingle-turn
System prompt publicYes
Reasoning / effortPsc requests clear, concise task answers and numeric output for regression tasks.
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNot reported
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
RepeatsNot reported
Graderkeyword-first binary extraction with sentiment-model fallback, followed by MCC · model: cardiffnlp/twitter-roberta-base-sentiment-latest fallback · human review: no
StatisticsPoint metric over the complete published test split, scaled by 100 and rounded to two decimals; no confidence interval is reported.
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
MCCabsolutepercentMatthews correlation coefficient across held-out test examples, scaled by 100Not reported

Results

ModelMetricValuen
ChatMultiOmics stage 1 + balanced stage 2MCC4.13 percent
EPI creator-paper result; metric scaled by 100.
308
ChatMultiOmics stage 2 onlyMCC4.77 percent
EPI creator-paper result; metric scaled by 100.
308
ChatMultiOmics stage 1 + stage 2MCC1.68 percent
EPI creator-paper result; metric scaled by 100.
308
ChatMultiOmicsMCC3.37 percent
EPI creator-paper result; metric scaled by 100.
308

Evidence

  • table: Table 2 (EPI test split); Appendix A.3; Table 8; Section 4.2 Psc prompt (Scope, prompt, output parser, grader, scaling, and aggregation.) — supports /scope, /benchmark_version, /model_ids, /protocol, /metrics
  • table: Table 7 (EPI column) (All registered model values; literature-SOTA row omitted.) — supports /results
bioinstruction-epi-open-baselines-emnlp-2025vemnlp-2025

Evaluated models / systems: Alpaca-7B (Biology-Instructions label), BioMedGPT-LM-7B (Biology-Instructions label), Galactica-1.3B (Biology-Instructions label), GLM-4-9B-Chat (Biology-Instructions label), InstructProtein-1.3B (Biology-Instructions label), Llama-molinst-protein-7B (Mol-Ins), Llama2-7B-Chat (Biology-Instructions label), LLaMA3.1-8B-Instruct (Biology-Instructions label), Qwen2-7B (Biology-Instructions label), Vicuna-v1.5-7B (Biology-Instructions label)

Scopesubset · n=308
Shots0
Turnssingle-turn
System prompt publicNo
Reasoning / effortThe prompt requests the task-formatted answer and says not to explain or repeat.
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNot reported
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
RepeatsNot reported
Graderkeyword-first binary extraction with sentiment-model fallback, followed by MCC · model: cardiffnlp/twitter-roberta-base-sentiment-latest fallback · human review: no
StatisticsPoint metric over the complete published test split, scaled by 100 and rounded to two decimals; no confidence interval is reported.
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
MCCabsolutepercentMatthews correlation coefficient across held-out test examples, scaled by 100Not reported

Results

ModelMetricValuen
LLaMA3.1-8B-Instruct (Biology-Instructions label)MCC0 percent
EPI creator-paper result; metric scaled by 100.
308
Qwen2-7B (Biology-Instructions label)MCC0 percent
EPI creator-paper result; metric scaled by 100.
308
Llama2-7B-Chat (Biology-Instructions label)MCC0 percent
EPI creator-paper result; metric scaled by 100.
308
Alpaca-7B (Biology-Instructions label)MCC0 percent
EPI creator-paper result; metric scaled by 100.
308
GLM-4-9B-Chat (Biology-Instructions label)MCC0 percent
EPI creator-paper result; metric scaled by 100.
308
Vicuna-v1.5-7B (Biology-Instructions label)MCC0 percent
EPI creator-paper result; metric scaled by 100.
308
Galactica-1.3B (Biology-Instructions label)MCC0 percent
EPI creator-paper result; metric scaled by 100.
308
InstructProtein-1.3B (Biology-Instructions label)MCC0 percent
EPI creator-paper result; metric scaled by 100.
308
Llama-molinst-protein-7B (Mol-Ins)MCC0 percent
EPI creator-paper result; metric scaled by 100.
308
BioMedGPT-LM-7B (Biology-Instructions label)MCC0 percent
EPI creator-paper result; metric scaled by 100.
308

Evidence

  • table: Table 2 (EPI test split); Appendix A.3; Table 8; Table 9 open-source prompt (Scope, prompt, output parser, grader, scaling, and aggregation.) — supports /scope, /benchmark_version, /model_ids, /protocol, /metrics
  • table: Table 7 (EPI column) (All registered model values; literature-SOTA row omitted.) — supports /results
Biology-Instructions Protein Fluorescence Prediction3 runs

Open benchmark record →

bioinstruction-fluorescence-closed-baselines-emnlp-2025vemnlp-2025

Evaluated models / systems: GPT-4o (Biology-Instructions snapshot not reported), GPT-4o-mini (Biology-Instructions snapshot not reported)

Scopesubset · n=27217
Shots0
Turnssingle-turn
System prompt publicYes
Reasoning / effortThe prompt requests a direct JSON answer and no chain-of-thought.
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNot reported
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
RepeatsNot reported
Graderfirst-number extraction followed by Spearman rank correlation · human review: no
StatisticsPoint metric over the complete published test split, scaled by 100 and rounded to two decimals; no confidence interval is reported.
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Spearman's ρabsolutepercentSpearman rank correlation across held-out test examples, scaled by 100Not reported

Results

ModelMetricValuen
GPT-4o-mini (Biology-Instructions snapshot not reported)Spearman's ρ-0.47 percent
Flu creator-paper result; metric scaled by 100.
27217
GPT-4o (Biology-Instructions snapshot not reported)Spearman's ρ0.69 percent
Flu creator-paper result; metric scaled by 100.
27217

Evidence

  • table: Table 2 (Flu test split); Appendix A.3; Table 8; Table 9 closed-source prompt (Scope, prompt, output parser, grader, scaling, and aggregation.) — supports /scope, /benchmark_version, /model_ids, /protocol, /metrics
  • table: Table 6 (Flu column) (All registered model values; literature-SOTA row omitted.) — supports /results
bioinstruction-fluorescence-creator-systems-emnlp-2025vemnlp-2025

Evaluated models / systems: ChatMultiOmics stage 1 + balanced stage 2, ChatMultiOmics stage 1 + stage 2, ChatMultiOmics, ChatMultiOmics stage 2 only

Scopesubset · n=27217
Shots0
Turnssingle-turn
System prompt publicYes
Reasoning / effortPsc requests clear, concise task answers and numeric output for regression tasks.
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNot reported
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
RepeatsNot reported
Graderfirst-number extraction followed by Spearman rank correlation · human review: no
StatisticsPoint metric over the complete published test split, scaled by 100 and rounded to two decimals; no confidence interval is reported.
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Spearman's ρabsolutepercentSpearman rank correlation across held-out test examples, scaled by 100Not reported

Results

ModelMetricValuen
ChatMultiOmics stage 1 + balanced stage 2Spearman's ρ0.55 percent
Flu creator-paper result; metric scaled by 100.
27217
ChatMultiOmics stage 2 onlySpearman's ρ0.37 percent
Flu creator-paper result; metric scaled by 100.
27217
ChatMultiOmics stage 1 + stage 2Spearman's ρ1.49 percent
Flu creator-paper result; metric scaled by 100.
27217
ChatMultiOmicsSpearman's ρ2.57 percent
Flu creator-paper result; metric scaled by 100.
27217

Evidence

  • table: Table 2 (Flu test split); Appendix A.3; Table 8; Section 4.2 Psc prompt (Scope, prompt, output parser, grader, scaling, and aggregation.) — supports /scope, /benchmark_version, /model_ids, /protocol, /metrics
  • table: Table 6 (Flu column) (All registered model values; literature-SOTA row omitted.) — supports /results
bioinstruction-fluorescence-open-baselines-emnlp-2025vemnlp-2025

Evaluated models / systems: Alpaca-7B (Biology-Instructions label), BioMedGPT-LM-7B (Biology-Instructions label), Galactica-1.3B (Biology-Instructions label), GLM-4-9B-Chat (Biology-Instructions label), InstructProtein-1.3B (Biology-Instructions label), Llama-molinst-protein-7B (Mol-Ins), Llama2-7B-Chat (Biology-Instructions label), LLaMA3.1-8B-Instruct (Biology-Instructions label), Qwen2-7B (Biology-Instructions label), Vicuna-v1.5-7B (Biology-Instructions label)

Scopesubset · n=27217
Shots0
Turnssingle-turn
System prompt publicNo
Reasoning / effortThe prompt requests the task-formatted answer and says not to explain or repeat.
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNot reported
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
RepeatsNot reported
Graderfirst-number extraction followed by Spearman rank correlation · human review: no
StatisticsPoint metric over the complete published test split, scaled by 100 and rounded to two decimals; no confidence interval is reported.
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Spearman's ρabsolutepercentSpearman rank correlation across held-out test examples, scaled by 100Not reported

Results

ModelMetricValuen
LLaMA3.1-8B-Instruct (Biology-Instructions label)Spearman's ρ0.91 percent
Flu creator-paper result; metric scaled by 100.
27217
Qwen2-7B (Biology-Instructions label)Spearman's ρ0.81 percent
Flu creator-paper result; metric scaled by 100.
27217
Llama2-7B-Chat (Biology-Instructions label)Spearman's ρ0.28 percent
Flu creator-paper result; metric scaled by 100.
27217
Alpaca-7B (Biology-Instructions label)Spearman's ρ-0.2 percent
Flu creator-paper result; metric scaled by 100.
27217
GLM-4-9B-Chat (Biology-Instructions label)Spearman's ρ0.63 percent
Flu creator-paper result; metric scaled by 100.
27217
Vicuna-v1.5-7B (Biology-Instructions label)Spearman's ρ-0.51 percent
Flu creator-paper result; metric scaled by 100.
27217
Galactica-1.3B (Biology-Instructions label)Spearman's ρ-0.73 percent
Flu creator-paper result; metric scaled by 100.
27217
InstructProtein-1.3B (Biology-Instructions label)Spearman's ρ-0.03 percent
Flu creator-paper result; metric scaled by 100.
27217
Llama-molinst-protein-7B (Mol-Ins)Spearman's ρ0.27 percent
Flu creator-paper result; metric scaled by 100.
27217
BioMedGPT-LM-7B (Biology-Instructions label)Spearman's ρ0.43 percent
Flu creator-paper result; metric scaled by 100.
27217

Evidence

  • table: Table 2 (Flu test split); Appendix A.3; Table 8; Table 9 open-source prompt (Scope, prompt, output parser, grader, scaling, and aggregation.) — supports /scope, /benchmark_version, /model_ids, /protocol, /metrics
  • table: Table 6 (Flu column) (All registered model values; literature-SOTA row omitted.) — supports /results
Biology-Instructions RNA Modification Prediction3 runs

Open benchmark record →

bioinstruction-modification-closed-baselines-emnlp-2025vemnlp-2025

Evaluated models / systems: GPT-4o (Biology-Instructions snapshot not reported), GPT-4o-mini (Biology-Instructions snapshot not reported)

Scopesubset · n=1200
Shots0
Turnssingle-turn
System prompt publicYes
Reasoning / effortThe prompt requests a direct JSON answer and no chain-of-thought.
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNot reported
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
RepeatsNot reported
Gradermodification-label extraction with sentiment fallback for none, followed by macro ROC AUC · model: cardiffnlp/twitter-roberta-base-sentiment-latest fallback · human review: no
StatisticsPoint metric over the complete published test split, scaled by 100 and rounded to two decimals; no confidence interval is reported.
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
AUCabsolutepercentMacro ROC AUC across modification labels on held-out test examples, scaled by 100Not reported

Results

ModelMetricValuen
GPT-4o-mini (Biology-Instructions snapshot not reported)AUC50.49 percent
Modif creator-paper result; metric scaled by 100.
1200
GPT-4o (Biology-Instructions snapshot not reported)AUC50.47 percent
Modif creator-paper result; metric scaled by 100.
1200

Evidence

  • table: Table 2 (Modif test split); Appendix A.3; Table 8; Table 9 closed-source prompt (Scope, prompt, output parser, grader, scaling, and aggregation.) — supports /scope, /benchmark_version, /model_ids, /protocol, /metrics
  • table: Table 5 (Modif column) (All registered model values; literature-SOTA row omitted.) — supports /results
bioinstruction-modification-creator-systems-emnlp-2025vemnlp-2025

Evaluated models / systems: ChatMultiOmics stage 1 + balanced stage 2, ChatMultiOmics stage 1 + stage 2, ChatMultiOmics, ChatMultiOmics stage 2 only

Scopesubset · n=1200
Shots0
Turnssingle-turn
System prompt publicYes
Reasoning / effortPsc requests clear, concise task answers and numeric output for regression tasks.
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNot reported
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
RepeatsNot reported
Gradermodification-label extraction with sentiment fallback for none, followed by macro ROC AUC · model: cardiffnlp/twitter-roberta-base-sentiment-latest fallback · human review: no
StatisticsPoint metric over the complete published test split, scaled by 100 and rounded to two decimals; no confidence interval is reported.
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
AUCabsolutepercentMacro ROC AUC across modification labels on held-out test examples, scaled by 100Not reported

Results

ModelMetricValuen
ChatMultiOmics stage 1 + balanced stage 2AUC53.76 percent
Modif creator-paper result; metric scaled by 100.
1200
ChatMultiOmics stage 2 onlyAUC51.21 percent
Modif creator-paper result; metric scaled by 100.
1200
ChatMultiOmics stage 1 + stage 2AUC57.45 percent
Modif creator-paper result; metric scaled by 100.
1200
ChatMultiOmicsAUC59.06 percent
Modif creator-paper result; metric scaled by 100.
1200

Evidence

  • table: Table 2 (Modif test split); Appendix A.3; Table 8; Section 4.2 Psc prompt (Scope, prompt, output parser, grader, scaling, and aggregation.) — supports /scope, /benchmark_version, /model_ids, /protocol, /metrics
  • table: Table 5 (Modif column) (All registered model values; literature-SOTA row omitted.) — supports /results
bioinstruction-modification-open-baselines-emnlp-2025vemnlp-2025

Evaluated models / systems: Alpaca-7B (Biology-Instructions label), BioMedGPT-LM-7B (Biology-Instructions label), Galactica-1.3B (Biology-Instructions label), GLM-4-9B-Chat (Biology-Instructions label), InstructProtein-1.3B (Biology-Instructions label), Llama-molinst-protein-7B (Mol-Ins), Llama2-7B-Chat (Biology-Instructions label), LLaMA3.1-8B-Instruct (Biology-Instructions label), Qwen2-7B (Biology-Instructions label), Vicuna-v1.5-7B (Biology-Instructions label)

Scopesubset · n=1200
Shots0
Turnssingle-turn
System prompt publicNo
Reasoning / effortThe prompt requests the task-formatted answer and says not to explain or repeat.
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNot reported
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
RepeatsNot reported
Gradermodification-label extraction with sentiment fallback for none, followed by macro ROC AUC · model: cardiffnlp/twitter-roberta-base-sentiment-latest fallback · human review: no
StatisticsPoint metric over the complete published test split, scaled by 100 and rounded to two decimals; no confidence interval is reported.
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
AUCabsolutepercentMacro ROC AUC across modification labels on held-out test examples, scaled by 100Not reported

Results

ModelMetricValuen
LLaMA3.1-8B-Instruct (Biology-Instructions label)AUC50.52 percent
Modif creator-paper result; metric scaled by 100.
1200
Qwen2-7B (Biology-Instructions label)AUC50.34 percent
Modif creator-paper result; metric scaled by 100.
1200
Llama2-7B-Chat (Biology-Instructions label)AUC50.4 percent
Modif creator-paper result; metric scaled by 100.
1200
Alpaca-7B (Biology-Instructions label)AUC50 percent
Modif creator-paper result; metric scaled by 100.
1200
GLM-4-9B-Chat (Biology-Instructions label)AUC50.05 percent
Modif creator-paper result; metric scaled by 100.
1200
Vicuna-v1.5-7B (Biology-Instructions label)AUC50.27 percent
Modif creator-paper result; metric scaled by 100.
1200
Galactica-1.3B (Biology-Instructions label)AUC53.78 percent
Modif creator-paper result; metric scaled by 100.
1200
InstructProtein-1.3B (Biology-Instructions label)AUC51.08 percent
Modif creator-paper result; metric scaled by 100.
1200
Llama-molinst-protein-7B (Mol-Ins)AUC52.51 percent
Modif creator-paper result; metric scaled by 100.
1200
BioMedGPT-LM-7B (Biology-Instructions label)AUC51.65 percent
Modif creator-paper result; metric scaled by 100.
1200

Evidence

  • table: Table 2 (Modif test split); Appendix A.3; Table 8; Table 9 open-source prompt (Scope, prompt, output parser, grader, scaling, and aggregation.) — supports /scope, /benchmark_version, /model_ids, /protocol, /metrics
  • table: Table 5 (Modif column) (All registered model values; literature-SOTA row omitted.) — supports /results
Biology-Instructions Mean Ribosome Loading Prediction3 runs

Open benchmark record →

bioinstruction-mrl-closed-baselines-emnlp-2025vemnlp-2025

Evaluated models / systems: GPT-4o (Biology-Instructions snapshot not reported), GPT-4o-mini (Biology-Instructions snapshot not reported)

Scopesubset · n=7600
Shots0
Turnssingle-turn
System prompt publicYes
Reasoning / effortThe prompt requests a direct JSON answer and no chain-of-thought.
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNot reported
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
RepeatsNot reported
Graderfirst-number extraction followed by squared Pearson correlation labeled R2 · human review: no
StatisticsPoint metric over the complete published test split, scaled by 100 and rounded to two decimals; no confidence interval is reported.
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
R2absolutepercentSquared Pearson correlation across held-out test examples, scaled by 100Not reported

Results

ModelMetricValuen
GPT-4o-mini (Biology-Instructions snapshot not reported)R20.01 percent
MRL creator-paper result; metric scaled by 100.
7600
GPT-4o (Biology-Instructions snapshot not reported)R20.01 percent
MRL creator-paper result; metric scaled by 100.
7600

Evidence

  • table: Table 2 (MRL test split); Appendix A.3; Table 8; Table 9 closed-source prompt (Scope, prompt, output parser, grader, scaling, and aggregation.) — supports /scope, /benchmark_version, /model_ids, /protocol, /metrics
  • table: Table 5 (MRL column) (All registered model values; literature-SOTA row omitted.) — supports /results
bioinstruction-mrl-creator-systems-emnlp-2025vemnlp-2025

Evaluated models / systems: ChatMultiOmics stage 1 + balanced stage 2, ChatMultiOmics stage 1 + stage 2, ChatMultiOmics, ChatMultiOmics stage 2 only

Scopesubset · n=7600
Shots0
Turnssingle-turn
System prompt publicYes
Reasoning / effortPsc requests clear, concise task answers and numeric output for regression tasks.
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNot reported
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
RepeatsNot reported
Graderfirst-number extraction followed by squared Pearson correlation labeled R2 · human review: no
StatisticsPoint metric over the complete published test split, scaled by 100 and rounded to two decimals; no confidence interval is reported.
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
R2absolutepercentSquared Pearson correlation across held-out test examples, scaled by 100Not reported

Results

ModelMetricValuen
ChatMultiOmics stage 1 + balanced stage 2R20 percent
MRL creator-paper result; metric scaled by 100.
7600
ChatMultiOmics stage 2 onlyR20 percent
MRL creator-paper result; metric scaled by 100.
7600
ChatMultiOmics stage 1 + stage 2R229.12 percent
MRL creator-paper result; metric scaled by 100.
7600
ChatMultiOmicsR247.64 percent
MRL creator-paper result; metric scaled by 100.
7600

Evidence

  • table: Table 2 (MRL test split); Appendix A.3; Table 8; Section 4.2 Psc prompt (Scope, prompt, output parser, grader, scaling, and aggregation.) — supports /scope, /benchmark_version, /model_ids, /protocol, /metrics
  • table: Table 5 (MRL column) (All registered model values; literature-SOTA row omitted.) — supports /results
bioinstruction-mrl-open-baselines-emnlp-2025vemnlp-2025

Evaluated models / systems: Alpaca-7B (Biology-Instructions label), BioMedGPT-LM-7B (Biology-Instructions label), Galactica-1.3B (Biology-Instructions label), GLM-4-9B-Chat (Biology-Instructions label), InstructProtein-1.3B (Biology-Instructions label), Llama-molinst-protein-7B (Mol-Ins), Llama2-7B-Chat (Biology-Instructions label), LLaMA3.1-8B-Instruct (Biology-Instructions label), Qwen2-7B (Biology-Instructions label), Vicuna-v1.5-7B (Biology-Instructions label)

Scopesubset · n=7600
Shots0
Turnssingle-turn
System prompt publicNo
Reasoning / effortThe prompt requests the task-formatted answer and says not to explain or repeat.
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNot reported
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
RepeatsNot reported
Graderfirst-number extraction followed by squared Pearson correlation labeled R2 · human review: no
StatisticsPoint metric over the complete published test split, scaled by 100 and rounded to two decimals; no confidence interval is reported.
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
R2absolutepercentSquared Pearson correlation across held-out test examples, scaled by 100Not reported

Results

ModelMetricValuen
LLaMA3.1-8B-Instruct (Biology-Instructions label)R20.01 percent
MRL creator-paper result; metric scaled by 100.
7600
Qwen2-7B (Biology-Instructions label)R20 percent
MRL creator-paper result; metric scaled by 100.
7600
Llama2-7B-Chat (Biology-Instructions label)R20 percent
MRL creator-paper result; metric scaled by 100.
7600
Alpaca-7B (Biology-Instructions label)R20.03 percent
MRL creator-paper result; metric scaled by 100.
7600
GLM-4-9B-Chat (Biology-Instructions label)R20 percent
MRL creator-paper result; metric scaled by 100.
7600
Vicuna-v1.5-7B (Biology-Instructions label)R20.01 percent
MRL creator-paper result; metric scaled by 100.
7600
Galactica-1.3B (Biology-Instructions label)R20 percent
MRL creator-paper result; metric scaled by 100.
7600
InstructProtein-1.3B (Biology-Instructions label)R20.02 percent
MRL creator-paper result; metric scaled by 100.
7600
Llama-molinst-protein-7B (Mol-Ins)R20 percent
MRL creator-paper result; metric scaled by 100.
7600
BioMedGPT-LM-7B (Biology-Instructions label)R20.01 percent
MRL creator-paper result; metric scaled by 100.
7600

Evidence

  • table: Table 2 (MRL test split); Appendix A.3; Table 8; Table 9 open-source prompt (Scope, prompt, output parser, grader, scaling, and aggregation.) — supports /scope, /benchmark_version, /model_ids, /protocol, /metrics
  • table: Table 5 (MRL column) (All registered model values; literature-SOTA row omitted.) — supports /results
Biology-Instructions Non-coding RNA Function Classification3 runs

Open benchmark record →

bioinstruction-ncrna-closed-baselines-emnlp-2025vemnlp-2025

Evaluated models / systems: GPT-4o (Biology-Instructions snapshot not reported), GPT-4o-mini (Biology-Instructions snapshot not reported)

Scopesubset · n=4840
Shots0
Turnssingle-turn
System prompt publicYes
Reasoning / effortThe prompt requests a direct JSON answer and no chain-of-thought.
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNot reported
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
RepeatsNot reported
Graderordered RNA-family name extraction followed by exact accuracy · human review: no
StatisticsPoint metric over the complete published test split, scaled by 100 and rounded to two decimals; no confidence interval is reported.
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
AccuracyabsolutepercentExact extracted-class accuracy across held-out test examples, scaled by 100Not reported

Results

ModelMetricValuen
GPT-4o-mini (Biology-Instructions snapshot not reported)Accuracy3 percent
ncRNA creator-paper result; metric scaled by 100.
4840
GPT-4o (Biology-Instructions snapshot not reported)Accuracy5.6 percent
ncRNA creator-paper result; metric scaled by 100.
4840

Evidence

  • table: Table 2 (ncRNA test split); Appendix A.3; Table 8; Table 9 closed-source prompt (Scope, prompt, output parser, grader, scaling, and aggregation.) — supports /scope, /benchmark_version, /model_ids, /protocol, /metrics
  • table: Table 5 (ncRNA column) (All registered model values; literature-SOTA row omitted.) — supports /results
bioinstruction-ncrna-creator-systems-emnlp-2025vemnlp-2025

Evaluated models / systems: ChatMultiOmics stage 1 + balanced stage 2, ChatMultiOmics stage 1 + stage 2, ChatMultiOmics, ChatMultiOmics stage 2 only

Scopesubset · n=4840
Shots0
Turnssingle-turn
System prompt publicYes
Reasoning / effortPsc requests clear, concise task answers and numeric output for regression tasks.
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNot reported
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
RepeatsNot reported
Graderordered RNA-family name extraction followed by exact accuracy · human review: no
StatisticsPoint metric over the complete published test split, scaled by 100 and rounded to two decimals; no confidence interval is reported.
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
AccuracyabsolutepercentExact extracted-class accuracy across held-out test examples, scaled by 100Not reported

Results

ModelMetricValuen
ChatMultiOmics stage 1 + balanced stage 2Accuracy35.68 percent
ncRNA creator-paper result; metric scaled by 100.
4840
ChatMultiOmics stage 2 onlyAccuracy0 percent
ncRNA creator-paper result; metric scaled by 100.
4840
ChatMultiOmics stage 1 + stage 2Accuracy62.77 percent
ncRNA creator-paper result; metric scaled by 100.
4840
ChatMultiOmicsAccuracy63.09 percent
ncRNA creator-paper result; metric scaled by 100.
4840

Evidence

  • table: Table 2 (ncRNA test split); Appendix A.3; Table 8; Section 4.2 Psc prompt (Scope, prompt, output parser, grader, scaling, and aggregation.) — supports /scope, /benchmark_version, /model_ids, /protocol, /metrics
  • table: Table 5 (ncRNA column) (All registered model values; literature-SOTA row omitted.) — supports /results
bioinstruction-ncrna-open-baselines-emnlp-2025vemnlp-2025

Evaluated models / systems: Alpaca-7B (Biology-Instructions label), BioMedGPT-LM-7B (Biology-Instructions label), Galactica-1.3B (Biology-Instructions label), GLM-4-9B-Chat (Biology-Instructions label), InstructProtein-1.3B (Biology-Instructions label), Llama-molinst-protein-7B (Mol-Ins), Llama2-7B-Chat (Biology-Instructions label), LLaMA3.1-8B-Instruct (Biology-Instructions label), Qwen2-7B (Biology-Instructions label), Vicuna-v1.5-7B (Biology-Instructions label)

Scopesubset · n=4840
Shots0
Turnssingle-turn
System prompt publicNo
Reasoning / effortThe prompt requests the task-formatted answer and says not to explain or repeat.
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNot reported
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
RepeatsNot reported
Graderordered RNA-family name extraction followed by exact accuracy · human review: no
StatisticsPoint metric over the complete published test split, scaled by 100 and rounded to two decimals; no confidence interval is reported.
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
AccuracyabsolutepercentExact extracted-class accuracy across held-out test examples, scaled by 100Not reported

Results

ModelMetricValuen
LLaMA3.1-8B-Instruct (Biology-Instructions label)Accuracy6.32 percent
ncRNA creator-paper result; metric scaled by 100.
4840
Qwen2-7B (Biology-Instructions label)Accuracy7.08 percent
ncRNA creator-paper result; metric scaled by 100.
4840
Llama2-7B-Chat (Biology-Instructions label)Accuracy4.88 percent
ncRNA creator-paper result; metric scaled by 100.
4840
Alpaca-7B (Biology-Instructions label)Accuracy7.42 percent
ncRNA creator-paper result; metric scaled by 100.
4840
GLM-4-9B-Chat (Biology-Instructions label)Accuracy8.23 percent
ncRNA creator-paper result; metric scaled by 100.
4840
Vicuna-v1.5-7B (Biology-Instructions label)Accuracy3.81 percent
ncRNA creator-paper result; metric scaled by 100.
4840
Galactica-1.3B (Biology-Instructions label)Accuracy6.73 percent
ncRNA creator-paper result; metric scaled by 100.
4840
InstructProtein-1.3B (Biology-Instructions label)Accuracy0 percent
ncRNA creator-paper result; metric scaled by 100.
4840
Llama-molinst-protein-7B (Mol-Ins)Accuracy0 percent
ncRNA creator-paper result; metric scaled by 100.
4840
BioMedGPT-LM-7B (Biology-Instructions label)Accuracy1.62 percent
ncRNA creator-paper result; metric scaled by 100.
4840

Evidence

  • table: Table 2 (ncRNA test split); Appendix A.3; Table 8; Table 9 open-source prompt (Scope, prompt, output parser, grader, scaling, and aggregation.) — supports /scope, /benchmark_version, /model_ids, /protocol, /metrics
  • table: Table 5 (ncRNA column) (All registered model values; literature-SOTA row omitted.) — supports /results
Biology-Instructions Promoter Detection 3003 runs

Open benchmark record →

bioinstruction-pd300-closed-baselines-emnlp-2025vemnlp-2025

Evaluated models / systems: GPT-4o (Biology-Instructions snapshot not reported), GPT-4o-mini (Biology-Instructions snapshot not reported)

Scopesubset · n=11840
Shots0
Turnssingle-turn
System prompt publicYes
Reasoning / effortThe prompt requests a direct JSON answer and no chain-of-thought.
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNot reported
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
RepeatsNot reported
Graderkeyword-first binary extraction with sentiment-model fallback, followed by MCC · model: cardiffnlp/twitter-roberta-base-sentiment-latest fallback · human review: no
StatisticsPoint metric over the complete published test split, scaled by 100 and rounded to two decimals; no confidence interval is reported.
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
MCCabsolutepercentMatthews correlation coefficient across held-out test examples, scaled by 100Not reported

Results

ModelMetricValuen
GPT-4o-mini (Biology-Instructions snapshot not reported)MCC-4.44 percent
PD300 creator-paper result; metric scaled by 100.
11840
GPT-4o (Biology-Instructions snapshot not reported)MCC8.67 percent
PD300 creator-paper result; metric scaled by 100.
11840

Evidence

  • table: Table 2 (PD300 test split); Appendix A.3; Table 8; Table 9 closed-source prompt (Scope, prompt, output parser, grader, scaling, and aggregation.) — supports /scope, /benchmark_version, /model_ids, /protocol, /metrics
  • table: Table 4 (PD300 column) (All registered model values; literature-SOTA row omitted.) — supports /results
bioinstruction-pd300-creator-systems-emnlp-2025vemnlp-2025

Evaluated models / systems: ChatMultiOmics stage 1 + balanced stage 2, ChatMultiOmics stage 1 + stage 2, ChatMultiOmics, ChatMultiOmics stage 2 only

Scopesubset · n=11840
Shots0
Turnssingle-turn
System prompt publicYes
Reasoning / effortPsc requests clear, concise task answers and numeric output for regression tasks.
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNot reported
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
RepeatsNot reported
Graderkeyword-first binary extraction with sentiment-model fallback, followed by MCC · model: cardiffnlp/twitter-roberta-base-sentiment-latest fallback · human review: no
StatisticsPoint metric over the complete published test split, scaled by 100 and rounded to two decimals; no confidence interval is reported.
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
MCCabsolutepercentMatthews correlation coefficient across held-out test examples, scaled by 100Not reported

Results

ModelMetricValuen
ChatMultiOmics stage 1 + balanced stage 2MCC5.19 percent
PD300 creator-paper result; metric scaled by 100.
11840
ChatMultiOmics stage 2 onlyMCC0.87 percent
PD300 creator-paper result; metric scaled by 100.
11840
ChatMultiOmics stage 1 + stage 2MCC49.01 percent
PD300 creator-paper result; metric scaled by 100.
11840
ChatMultiOmicsMCC58.18 percent
PD300 creator-paper result; metric scaled by 100.
11840

Evidence

  • table: Table 2 (PD300 test split); Appendix A.3; Table 8; Section 4.2 Psc prompt (Scope, prompt, output parser, grader, scaling, and aggregation.) — supports /scope, /benchmark_version, /model_ids, /protocol, /metrics
  • table: Table 4 (PD300 column) (All registered model values; literature-SOTA row omitted.) — supports /results
bioinstruction-pd300-open-baselines-emnlp-2025vemnlp-2025

Evaluated models / systems: Alpaca-7B (Biology-Instructions label), Galactica-1.3B (Biology-Instructions label), GLM-4-9B-Chat (Biology-Instructions label), InstructProtein-1.3B (Biology-Instructions label), Llama-molinst-protein-7B (Mol-Ins), Llama2-7B-Chat (Biology-Instructions label), LLaMA3.1-8B-Instruct (Biology-Instructions label), Qwen2-7B (Biology-Instructions label), Vicuna-v1.5-7B (Biology-Instructions label)

Scopesubset · n=11840
Shots0
Turnssingle-turn
System prompt publicNo
Reasoning / effortThe prompt requests the task-formatted answer and says not to explain or repeat.
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNot reported
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
RepeatsNot reported
Graderkeyword-first binary extraction with sentiment-model fallback, followed by MCC · model: cardiffnlp/twitter-roberta-base-sentiment-latest fallback · human review: no
StatisticsPoint metric over the complete published test split, scaled by 100 and rounded to two decimals; no confidence interval is reported.
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
MCCabsolutepercentMatthews correlation coefficient across held-out test examples, scaled by 100Not reported

Results

ModelMetricValuen
LLaMA3.1-8B-Instruct (Biology-Instructions label)MCC0.01 percent
PD300 creator-paper result; metric scaled by 100.
11840
Qwen2-7B (Biology-Instructions label)MCC-4.83 percent
PD300 creator-paper result; metric scaled by 100.
11840
Llama2-7B-Chat (Biology-Instructions label)MCC-0.29 percent
PD300 creator-paper result; metric scaled by 100.
11840
Alpaca-7B (Biology-Instructions label)MCC-0.15 percent
PD300 creator-paper result; metric scaled by 100.
11840
GLM-4-9B-Chat (Biology-Instructions label)MCC-0.25 percent
PD300 creator-paper result; metric scaled by 100.
11840
Vicuna-v1.5-7B (Biology-Instructions label)MCC0 percent
PD300 creator-paper result; metric scaled by 100.
11840
Galactica-1.3B (Biology-Instructions label)MCC0.41 percent
PD300 creator-paper result; metric scaled by 100.
11840
InstructProtein-1.3B (Biology-Instructions label)MCC2.75 percent
PD300 creator-paper result; metric scaled by 100.
11840
Llama-molinst-protein-7B (Mol-Ins)MCC-5.76 percent
PD300 creator-paper result; metric scaled by 100.
11840

Evidence

  • table: Table 2 (PD300 test split); Appendix A.3; Table 8; Table 9 open-source prompt (Scope, prompt, output parser, grader, scaling, and aggregation.) — supports /scope, /benchmark_version, /model_ids, /protocol, /metrics
  • table: Table 4 (PD300 column) (All registered model values; literature-SOTA row omitted.) — supports /results
Biology-Instructions Programmable RNA Switches3 runs

Open benchmark record →

bioinstruction-prs-closed-baselines-emnlp-2025vemnlp-2025

Evaluated models / systems: GPT-4o (Biology-Instructions snapshot not reported), GPT-4o-mini (Biology-Instructions snapshot not reported)

Scopesubset · n=11019
Shots0
Turnssingle-turn
System prompt publicYes
Reasoning / effortThe prompt requests a direct JSON answer and no chain-of-thought.
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNot reported
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
RepeatsNot reported
Graderthree-number extraction followed by the mean of three squared Pearson correlations · human review: no
StatisticsPoint metric over the complete published test split, scaled by 100 and rounded to two decimals; no confidence interval is reported.
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
R2absolutepercentMean of ON, OFF, and ON/OFF squared Pearson correlations, scaled by 100Not reported

Results

ModelMetricValuen
GPT-4o-mini (Biology-Instructions snapshot not reported)R20.03 percent
PRS creator-paper result; metric scaled by 100.
11019
GPT-4o (Biology-Instructions snapshot not reported)R20 percent
PRS creator-paper result; metric scaled by 100.
11019

Evidence

  • table: Table 2 (PRS test split); Appendix A.3; Table 8; Table 9 closed-source prompt (Scope, prompt, output parser, grader, scaling, and aggregation.) — supports /scope, /benchmark_version, /model_ids, /protocol, /metrics
  • table: Table 5 (PRS column) (All registered model values; literature-SOTA row omitted.) — supports /results
bioinstruction-prs-creator-systems-emnlp-2025vemnlp-2025

Evaluated models / systems: ChatMultiOmics stage 1 + balanced stage 2, ChatMultiOmics stage 1 + stage 2, ChatMultiOmics, ChatMultiOmics stage 2 only

Scopesubset · n=11019
Shots0
Turnssingle-turn
System prompt publicYes
Reasoning / effortPsc requests clear, concise task answers and numeric output for regression tasks.
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNot reported
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
RepeatsNot reported
Graderthree-number extraction followed by the mean of three squared Pearson correlations · human review: no
StatisticsPoint metric over the complete published test split, scaled by 100 and rounded to two decimals; no confidence interval is reported.
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
R2absolutepercentMean of ON, OFF, and ON/OFF squared Pearson correlations, scaled by 100Not reported

Results

ModelMetricValuen
ChatMultiOmics stage 1 + balanced stage 2R20.01 percent
PRS creator-paper result; metric scaled by 100.
11019
ChatMultiOmics stage 2 onlyR20 percent
PRS creator-paper result; metric scaled by 100.
11019
ChatMultiOmics stage 1 + stage 2R226.65 percent
PRS creator-paper result; metric scaled by 100.
11019
ChatMultiOmicsR226.57 percent
PRS creator-paper result; metric scaled by 100.
11019

Evidence

  • table: Table 2 (PRS test split); Appendix A.3; Table 8; Section 4.2 Psc prompt (Scope, prompt, output parser, grader, scaling, and aggregation.) — supports /scope, /benchmark_version, /model_ids, /protocol, /metrics
  • table: Table 5 (PRS column) (All registered model values; literature-SOTA row omitted.) — supports /results
bioinstruction-prs-open-baselines-emnlp-2025vemnlp-2025

Evaluated models / systems: Alpaca-7B (Biology-Instructions label), BioMedGPT-LM-7B (Biology-Instructions label), Galactica-1.3B (Biology-Instructions label), GLM-4-9B-Chat (Biology-Instructions label), InstructProtein-1.3B (Biology-Instructions label), Llama-molinst-protein-7B (Mol-Ins), Llama2-7B-Chat (Biology-Instructions label), LLaMA3.1-8B-Instruct (Biology-Instructions label), Qwen2-7B (Biology-Instructions label), Vicuna-v1.5-7B (Biology-Instructions label)

Scopesubset · n=11019
Shots0
Turnssingle-turn
System prompt publicNo
Reasoning / effortThe prompt requests the task-formatted answer and says not to explain or repeat.
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNot reported
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
RepeatsNot reported
Graderthree-number extraction followed by the mean of three squared Pearson correlations · human review: no
StatisticsPoint metric over the complete published test split, scaled by 100 and rounded to two decimals; no confidence interval is reported.
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
R2absolutepercentMean of ON, OFF, and ON/OFF squared Pearson correlations, scaled by 100Not reported

Results

ModelMetricValuen
LLaMA3.1-8B-Instruct (Biology-Instructions label)R20.02 percent
PRS creator-paper result; metric scaled by 100.
11019
Qwen2-7B (Biology-Instructions label)R20.01 percent
PRS creator-paper result; metric scaled by 100.
11019
Llama2-7B-Chat (Biology-Instructions label)R20.01 percent
PRS creator-paper result; metric scaled by 100.
11019
Alpaca-7B (Biology-Instructions label)R20.01 percent
PRS creator-paper result; metric scaled by 100.
11019
GLM-4-9B-Chat (Biology-Instructions label)R20.01 percent
PRS creator-paper result; metric scaled by 100.
11019
Vicuna-v1.5-7B (Biology-Instructions label)R20 percent
PRS creator-paper result; metric scaled by 100.
11019
Galactica-1.3B (Biology-Instructions label)R20.02 percent
PRS creator-paper result; metric scaled by 100.
11019
InstructProtein-1.3B (Biology-Instructions label)R20 percent
PRS creator-paper result; metric scaled by 100.
11019
Llama-molinst-protein-7B (Mol-Ins)R20.02 percent
PRS creator-paper result; metric scaled by 100.
11019
BioMedGPT-LM-7B (Biology-Instructions label)R20.03 percent
PRS creator-paper result; metric scaled by 100.
11019

Evidence

  • table: Table 2 (PRS test split); Appendix A.3; Table 8; Table 9 open-source prompt (Scope, prompt, output parser, grader, scaling, and aggregation.) — supports /scope, /benchmark_version, /model_ids, /protocol, /metrics
  • table: Table 5 (PRS column) (All registered model values; literature-SOTA row omitted.) — supports /results
Biology-Instructions RNA-Protein Interaction Prediction3 runs

Open benchmark record →

bioinstruction-rpi-closed-baselines-emnlp-2025vemnlp-2025

Evaluated models / systems: GPT-4o (Biology-Instructions snapshot not reported), GPT-4o-mini (Biology-Instructions snapshot not reported)

Scopesubset · n=4164
Shots0
Turnssingle-turn
System prompt publicYes
Reasoning / effortThe prompt requests a direct JSON answer and no chain-of-thought.
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNot reported
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
RepeatsNot reported
Graderkeyword-first binary extraction with sentiment-model fallback, followed by MCC · model: cardiffnlp/twitter-roberta-base-sentiment-latest fallback · human review: no
StatisticsPoint metric over the complete published test split, scaled by 100 and rounded to two decimals; no confidence interval is reported.
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
MCCabsolutepercentMatthews correlation coefficient across held-out test examples, scaled by 100Not reported

Results

ModelMetricValuen
GPT-4o-mini (Biology-Instructions snapshot not reported)MCC1.22 percent
RPI creator-paper result; metric scaled by 100.
4164
GPT-4o (Biology-Instructions snapshot not reported)MCC1.17 percent
RPI creator-paper result; metric scaled by 100.
4164

Evidence

  • table: Table 2 (RPI test split); Appendix A.3; Table 8; Table 9 closed-source prompt (Scope, prompt, output parser, grader, scaling, and aggregation.) — supports /scope, /benchmark_version, /model_ids, /protocol, /metrics
  • table: Table 7 (RPI column) (All registered model values; literature-SOTA row omitted.) — supports /results
bioinstruction-rpi-creator-systems-emnlp-2025vemnlp-2025

Evaluated models / systems: ChatMultiOmics stage 1 + balanced stage 2, ChatMultiOmics stage 1 + stage 2, ChatMultiOmics, ChatMultiOmics stage 2 only

Scopesubset · n=4164
Shots0
Turnssingle-turn
System prompt publicYes
Reasoning / effortPsc requests clear, concise task answers and numeric output for regression tasks.
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNot reported
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
RepeatsNot reported
Graderkeyword-first binary extraction with sentiment-model fallback, followed by MCC · model: cardiffnlp/twitter-roberta-base-sentiment-latest fallback · human review: no
StatisticsPoint metric over the complete published test split, scaled by 100 and rounded to two decimals; no confidence interval is reported.
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
MCCabsolutepercentMatthews correlation coefficient across held-out test examples, scaled by 100Not reported

Results

ModelMetricValuen
ChatMultiOmics stage 1 + balanced stage 2MCC8.29 percent
RPI creator-paper result; metric scaled by 100.
4164
ChatMultiOmics stage 2 onlyMCC1.61 percent
RPI creator-paper result; metric scaled by 100.
4164
ChatMultiOmics stage 1 + stage 2MCC70.8 percent
RPI creator-paper result; metric scaled by 100.
4164
ChatMultiOmicsMCC74.26 percent
RPI creator-paper result; metric scaled by 100.
4164

Evidence

  • table: Table 2 (RPI test split); Appendix A.3; Table 8; Section 4.2 Psc prompt (Scope, prompt, output parser, grader, scaling, and aggregation.) — supports /scope, /benchmark_version, /model_ids, /protocol, /metrics
  • table: Table 7 (RPI column) (All registered model values; literature-SOTA row omitted.) — supports /results
bioinstruction-rpi-open-baselines-emnlp-2025vemnlp-2025

Evaluated models / systems: Alpaca-7B (Biology-Instructions label), BioMedGPT-LM-7B (Biology-Instructions label), Galactica-1.3B (Biology-Instructions label), GLM-4-9B-Chat (Biology-Instructions label), InstructProtein-1.3B (Biology-Instructions label), Llama-molinst-protein-7B (Mol-Ins), Llama2-7B-Chat (Biology-Instructions label), LLaMA3.1-8B-Instruct (Biology-Instructions label), Qwen2-7B (Biology-Instructions label), Vicuna-v1.5-7B (Biology-Instructions label)

Scopesubset · n=4164
Shots0
Turnssingle-turn
System prompt publicNo
Reasoning / effortThe prompt requests the task-formatted answer and says not to explain or repeat.
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNot reported
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
RepeatsNot reported
Graderkeyword-first binary extraction with sentiment-model fallback, followed by MCC · model: cardiffnlp/twitter-roberta-base-sentiment-latest fallback · human review: no
StatisticsPoint metric over the complete published test split, scaled by 100 and rounded to two decimals; no confidence interval is reported.
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
MCCabsolutepercentMatthews correlation coefficient across held-out test examples, scaled by 100Not reported

Results

ModelMetricValuen
LLaMA3.1-8B-Instruct (Biology-Instructions label)MCC3.82 percent
RPI creator-paper result; metric scaled by 100.
4164
Qwen2-7B (Biology-Instructions label)MCC-2.15 percent
RPI creator-paper result; metric scaled by 100.
4164
Llama2-7B-Chat (Biology-Instructions label)MCC5.87 percent
RPI creator-paper result; metric scaled by 100.
4164
Alpaca-7B (Biology-Instructions label)MCC4.38 percent
RPI creator-paper result; metric scaled by 100.
4164
GLM-4-9B-Chat (Biology-Instructions label)MCC0.13 percent
RPI creator-paper result; metric scaled by 100.
4164
Vicuna-v1.5-7B (Biology-Instructions label)MCC0 percent
RPI creator-paper result; metric scaled by 100.
4164
Galactica-1.3B (Biology-Instructions label)MCC0.24 percent
RPI creator-paper result; metric scaled by 100.
4164
InstructProtein-1.3B (Biology-Instructions label)MCC-1.55 percent
RPI creator-paper result; metric scaled by 100.
4164
Llama-molinst-protein-7B (Mol-Ins)MCC3.71 percent
RPI creator-paper result; metric scaled by 100.
4164
BioMedGPT-LM-7B (Biology-Instructions label)MCC-2.39 percent
RPI creator-paper result; metric scaled by 100.
4164

Evidence

  • table: Table 2 (RPI test split); Appendix A.3; Table 8; Table 9 open-source prompt (Scope, prompt, output parser, grader, scaling, and aggregation.) — supports /scope, /benchmark_version, /model_ids, /protocol, /metrics
  • table: Table 7 (RPI column) (All registered model values; literature-SOTA row omitted.) — supports /results
Biology-Instructions siRNA Efficiency Prediction3 runs

Open benchmark record →

bioinstruction-sirna-closed-baselines-emnlp-2025vemnlp-2025

Evaluated models / systems: GPT-4o (Biology-Instructions snapshot not reported), GPT-4o-mini (Biology-Instructions snapshot not reported)

Scopesubset · n=6688
Shots0
Turnssingle-turn
System prompt publicYes
Reasoning / effortThe prompt requests a direct JSON answer and no chain-of-thought.
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNot reported
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
RepeatsNot reported
Graderfirst-number extraction followed by the creator mixed-score implementation · human review: no
StatisticsPoint metric over the complete published test split, scaled by 100 and rounded to two decimals; no confidence interval is reported.
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Mixed ScoreabsolutepercentCreator mixed score combining capped MAE and range-MAE-weighted binary F1, scaled by 100Not reported

Results

ModelMetricValuen
GPT-4o-mini (Biology-Instructions snapshot not reported)Mixed Score30.37 percent
siRNA creator-paper result; metric scaled by 100.
6688
GPT-4o (Biology-Instructions snapshot not reported)Mixed Score0 percent
siRNA creator-paper result; metric scaled by 100.
6688

Evidence

  • table: Table 2 (siRNA test split); Appendix A.3; Table 8; Table 9 closed-source prompt (Scope, prompt, output parser, grader, scaling, and aggregation.) — supports /scope, /benchmark_version, /model_ids, /protocol, /metrics
  • table: Table 7 (siRNA column) (All registered model values; literature-SOTA row omitted.) — supports /results
bioinstruction-sirna-creator-systems-emnlp-2025vemnlp-2025

Evaluated models / systems: ChatMultiOmics stage 1 + balanced stage 2, ChatMultiOmics stage 1 + stage 2, ChatMultiOmics, ChatMultiOmics stage 2 only

Scopesubset · n=6688
Shots0
Turnssingle-turn
System prompt publicYes
Reasoning / effortPsc requests clear, concise task answers and numeric output for regression tasks.
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNot reported
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
RepeatsNot reported
Graderfirst-number extraction followed by the creator mixed-score implementation · human review: no
StatisticsPoint metric over the complete published test split, scaled by 100 and rounded to two decimals; no confidence interval is reported.
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Mixed ScoreabsolutepercentCreator mixed score combining capped MAE and range-MAE-weighted binary F1, scaled by 100Not reported

Results

ModelMetricValuen
ChatMultiOmics stage 1 + balanced stage 2Mixed Score42.92 percent
siRNA creator-paper result; metric scaled by 100.
6688
ChatMultiOmics stage 2 onlyMixed Score4.25 percent
siRNA creator-paper result; metric scaled by 100.
6688
ChatMultiOmics stage 1 + stage 2Mixed Score56.31 percent
siRNA creator-paper result; metric scaled by 100.
6688
ChatMultiOmicsMixed Score56.25 percent
siRNA creator-paper result; metric scaled by 100.
6688

Evidence

  • table: Table 2 (siRNA test split); Appendix A.3; Table 8; Section 4.2 Psc prompt (Scope, prompt, output parser, grader, scaling, and aggregation.) — supports /scope, /benchmark_version, /model_ids, /protocol, /metrics
  • table: Table 7 (siRNA column) (All registered model values; literature-SOTA row omitted.) — supports /results
bioinstruction-sirna-open-baselines-emnlp-2025vemnlp-2025

Evaluated models / systems: Alpaca-7B (Biology-Instructions label), BioMedGPT-LM-7B (Biology-Instructions label), Galactica-1.3B (Biology-Instructions label), GLM-4-9B-Chat (Biology-Instructions label), InstructProtein-1.3B (Biology-Instructions label), Llama-molinst-protein-7B (Mol-Ins), Llama2-7B-Chat (Biology-Instructions label), LLaMA3.1-8B-Instruct (Biology-Instructions label), Qwen2-7B (Biology-Instructions label), Vicuna-v1.5-7B (Biology-Instructions label)

Scopesubset · n=6688
Shots0
Turnssingle-turn
System prompt publicNo
Reasoning / effortThe prompt requests the task-formatted answer and says not to explain or repeat.
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNot reported
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
RepeatsNot reported
Graderfirst-number extraction followed by the creator mixed-score implementation · human review: no
StatisticsPoint metric over the complete published test split, scaled by 100 and rounded to two decimals; no confidence interval is reported.
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Mixed ScoreabsolutepercentCreator mixed score combining capped MAE and range-MAE-weighted binary F1, scaled by 100Not reported

Results

ModelMetricValuen
LLaMA3.1-8B-Instruct (Biology-Instructions label)Mixed Score32.76 percent
siRNA creator-paper result; metric scaled by 100.
6688
Qwen2-7B (Biology-Instructions label)Mixed Score33.39 percent
siRNA creator-paper result; metric scaled by 100.
6688
Llama2-7B-Chat (Biology-Instructions label)Mixed Score17.43 percent
siRNA creator-paper result; metric scaled by 100.
6688
Alpaca-7B (Biology-Instructions label)Mixed Score19.12 percent
siRNA creator-paper result; metric scaled by 100.
6688
GLM-4-9B-Chat (Biology-Instructions label)Mixed Score23.33 percent
siRNA creator-paper result; metric scaled by 100.
6688
Vicuna-v1.5-7B (Biology-Instructions label)Mixed Score14.28 percent
siRNA creator-paper result; metric scaled by 100.
6688
Galactica-1.3B (Biology-Instructions label)Mixed Score33.55 percent
siRNA creator-paper result; metric scaled by 100.
6688
InstructProtein-1.3B (Biology-Instructions label)Mixed Score5.58 percent
siRNA creator-paper result; metric scaled by 100.
6688
Llama-molinst-protein-7B (Mol-Ins)Mixed Score13.85 percent
siRNA creator-paper result; metric scaled by 100.
6688
BioMedGPT-LM-7B (Biology-Instructions label)Mixed Score19.71 percent
siRNA creator-paper result; metric scaled by 100.
6688

Evidence

  • table: Table 2 (siRNA test split); Appendix A.3; Table 8; Table 9 open-source prompt (Scope, prompt, output parser, grader, scaling, and aggregation.) — supports /scope, /benchmark_version, /model_ids, /protocol, /metrics
  • table: Table 7 (siRNA column) (All registered model values; literature-SOTA row omitted.) — supports /results
Biology-Instructions Protein Solubility Prediction3 runs

Open benchmark record →

bioinstruction-solubility-closed-baselines-emnlp-2025vemnlp-2025

Evaluated models / systems: GPT-4o (Biology-Instructions snapshot not reported), GPT-4o-mini (Biology-Instructions snapshot not reported)

Scopesubset · n=2001
Shots0
Turnssingle-turn
System prompt publicYes
Reasoning / effortThe prompt requests a direct JSON answer and no chain-of-thought.
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNot reported
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
RepeatsNot reported
Graderkeyword-first binary extraction with sentiment-model fallback, followed by accuracy · model: cardiffnlp/twitter-roberta-base-sentiment-latest fallback · human review: no
StatisticsPoint metric over the complete published test split, scaled by 100 and rounded to two decimals; no confidence interval is reported.
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
AccuracyabsolutepercentExact binary accuracy across held-out test examples, scaled by 100Not reported

Results

ModelMetricValuen
GPT-4o-mini (Biology-Instructions snapshot not reported)Accuracy50.02 percent
Sol creator-paper result; metric scaled by 100.
2001
GPT-4o (Biology-Instructions snapshot not reported)Accuracy51.67 percent
Sol creator-paper result; metric scaled by 100.
2001

Evidence

  • table: Table 2 (Sol test split); Appendix A.3; Table 8; Table 9 closed-source prompt (Scope, prompt, output parser, grader, scaling, and aggregation.) — supports /scope, /benchmark_version, /model_ids, /protocol, /metrics
  • table: Table 6 (Sol column) (All registered model values; literature-SOTA row omitted.) — supports /results
bioinstruction-solubility-creator-systems-emnlp-2025vemnlp-2025

Evaluated models / systems: ChatMultiOmics stage 1 + balanced stage 2, ChatMultiOmics stage 1 + stage 2, ChatMultiOmics, ChatMultiOmics stage 2 only

Scopesubset · n=2001
Shots0
Turnssingle-turn
System prompt publicYes
Reasoning / effortPsc requests clear, concise task answers and numeric output for regression tasks.
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNot reported
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
RepeatsNot reported
Graderkeyword-first binary extraction with sentiment-model fallback, followed by accuracy · model: cardiffnlp/twitter-roberta-base-sentiment-latest fallback · human review: no
StatisticsPoint metric over the complete published test split, scaled by 100 and rounded to two decimals; no confidence interval is reported.
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
AccuracyabsolutepercentExact binary accuracy across held-out test examples, scaled by 100Not reported

Results

ModelMetricValuen
ChatMultiOmics stage 1 + balanced stage 2Accuracy52.37 percent
Sol creator-paper result; metric scaled by 100.
2001
ChatMultiOmics stage 2 onlyAccuracy49.28 percent
Sol creator-paper result; metric scaled by 100.
2001
ChatMultiOmics stage 1 + stage 2Accuracy62.07 percent
Sol creator-paper result; metric scaled by 100.
2001
ChatMultiOmicsAccuracy63.02 percent
Sol creator-paper result; metric scaled by 100.
2001

Evidence

  • table: Table 2 (Sol test split); Appendix A.3; Table 8; Section 4.2 Psc prompt (Scope, prompt, output parser, grader, scaling, and aggregation.) — supports /scope, /benchmark_version, /model_ids, /protocol, /metrics
  • table: Table 6 (Sol column) (All registered model values; literature-SOTA row omitted.) — supports /results
bioinstruction-solubility-open-baselines-emnlp-2025vemnlp-2025

Evaluated models / systems: Alpaca-7B (Biology-Instructions label), BioMedGPT-LM-7B (Biology-Instructions label), Galactica-1.3B (Biology-Instructions label), GLM-4-9B-Chat (Biology-Instructions label), InstructProtein-1.3B (Biology-Instructions label), Llama-molinst-protein-7B (Mol-Ins), Llama2-7B-Chat (Biology-Instructions label), LLaMA3.1-8B-Instruct (Biology-Instructions label), Qwen2-7B (Biology-Instructions label), Vicuna-v1.5-7B (Biology-Instructions label)

Scopesubset · n=2001
Shots0
Turnssingle-turn
System prompt publicNo
Reasoning / effortThe prompt requests the task-formatted answer and says not to explain or repeat.
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNot reported
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
RepeatsNot reported
Graderkeyword-first binary extraction with sentiment-model fallback, followed by accuracy · model: cardiffnlp/twitter-roberta-base-sentiment-latest fallback · human review: no
StatisticsPoint metric over the complete published test split, scaled by 100 and rounded to two decimals; no confidence interval is reported.
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
AccuracyabsolutepercentExact binary accuracy across held-out test examples, scaled by 100Not reported

Results

ModelMetricValuen
LLaMA3.1-8B-Instruct (Biology-Instructions label)Accuracy50.27 percent
Sol creator-paper result; metric scaled by 100.
2001
Qwen2-7B (Biology-Instructions label)Accuracy52.52 percent
Sol creator-paper result; metric scaled by 100.
2001
Llama2-7B-Chat (Biology-Instructions label)Accuracy49.48 percent
Sol creator-paper result; metric scaled by 100.
2001
Alpaca-7B (Biology-Instructions label)Accuracy50.12 percent
Sol creator-paper result; metric scaled by 100.
2001
GLM-4-9B-Chat (Biology-Instructions label)Accuracy50.72 percent
Sol creator-paper result; metric scaled by 100.
2001
Vicuna-v1.5-7B (Biology-Instructions label)Accuracy51.57 percent
Sol creator-paper result; metric scaled by 100.
2001
Galactica-1.3B (Biology-Instructions label)Accuracy46.78 percent
Sol creator-paper result; metric scaled by 100.
2001
InstructProtein-1.3B (Biology-Instructions label)Accuracy47.88 percent
Sol creator-paper result; metric scaled by 100.
2001
Llama-molinst-protein-7B (Mol-Ins)Accuracy48.33 percent
Sol creator-paper result; metric scaled by 100.
2001
BioMedGPT-LM-7B (Biology-Instructions label)Accuracy49.78 percent
Sol creator-paper result; metric scaled by 100.
2001

Evidence

  • table: Table 2 (Sol test split); Appendix A.3; Table 8; Table 9 open-source prompt (Scope, prompt, output parser, grader, scaling, and aggregation.) — supports /scope, /benchmark_version, /model_ids, /protocol, /metrics
  • table: Table 6 (Sol column) (All registered model values; literature-SOTA row omitted.) — supports /results
Biology-Instructions Protein Stability Prediction3 runs

Open benchmark record →

bioinstruction-stability-closed-baselines-emnlp-2025vemnlp-2025

Evaluated models / systems: GPT-4o (Biology-Instructions snapshot not reported), GPT-4o-mini (Biology-Instructions snapshot not reported)

Scopesubset · n=12851
Shots0
Turnssingle-turn
System prompt publicYes
Reasoning / effortThe prompt requests a direct JSON answer and no chain-of-thought.
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNot reported
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
RepeatsNot reported
Graderfirst-number extraction followed by Spearman rank correlation · human review: no
StatisticsPoint metric over the complete published test split, scaled by 100 and rounded to two decimals; no confidence interval is reported.
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Spearman's ρabsolutepercentSpearman rank correlation across held-out test examples, scaled by 100Not reported

Results

ModelMetricValuen
GPT-4o-mini (Biology-Instructions snapshot not reported)Spearman's ρ-1.52 percent
Sta creator-paper result; metric scaled by 100.
12851
GPT-4o (Biology-Instructions snapshot not reported)Spearman's ρ0.09 percent
Sta creator-paper result; metric scaled by 100.
12851

Evidence

  • table: Table 2 (Sta test split); Appendix A.3; Table 8; Table 9 closed-source prompt (Scope, prompt, output parser, grader, scaling, and aggregation.) — supports /scope, /benchmark_version, /model_ids, /protocol, /metrics
  • table: Table 6 (Sta column) (All registered model values; literature-SOTA row omitted.) — supports /results
bioinstruction-stability-creator-systems-emnlp-2025vemnlp-2025

Evaluated models / systems: ChatMultiOmics stage 1 + balanced stage 2, ChatMultiOmics stage 1 + stage 2, ChatMultiOmics, ChatMultiOmics stage 2 only

Scopesubset · n=12851
Shots0
Turnssingle-turn
System prompt publicYes
Reasoning / effortPsc requests clear, concise task answers and numeric output for regression tasks.
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNot reported
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
RepeatsNot reported
Graderfirst-number extraction followed by Spearman rank correlation · human review: no
StatisticsPoint metric over the complete published test split, scaled by 100 and rounded to two decimals; no confidence interval is reported.
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Spearman's ρabsolutepercentSpearman rank correlation across held-out test examples, scaled by 100Not reported

Results

ModelMetricValuen
ChatMultiOmics stage 1 + balanced stage 2Spearman's ρ0.48 percent
Sta creator-paper result; metric scaled by 100.
12851
ChatMultiOmics stage 2 onlySpearman's ρ0.23 percent
Sta creator-paper result; metric scaled by 100.
12851
ChatMultiOmics stage 1 + stage 2Spearman's ρ56.76 percent
Sta creator-paper result; metric scaled by 100.
12851
ChatMultiOmicsSpearman's ρ60.25 percent
Sta creator-paper result; metric scaled by 100.
12851

Evidence

  • table: Table 2 (Sta test split); Appendix A.3; Table 8; Section 4.2 Psc prompt (Scope, prompt, output parser, grader, scaling, and aggregation.) — supports /scope, /benchmark_version, /model_ids, /protocol, /metrics
  • table: Table 6 (Sta column) (All registered model values; literature-SOTA row omitted.) — supports /results
bioinstruction-stability-open-baselines-emnlp-2025vemnlp-2025

Evaluated models / systems: Alpaca-7B (Biology-Instructions label), BioMedGPT-LM-7B (Biology-Instructions label), Galactica-1.3B (Biology-Instructions label), GLM-4-9B-Chat (Biology-Instructions label), InstructProtein-1.3B (Biology-Instructions label), Llama-molinst-protein-7B (Mol-Ins), Llama2-7B-Chat (Biology-Instructions label), LLaMA3.1-8B-Instruct (Biology-Instructions label), Qwen2-7B (Biology-Instructions label), Vicuna-v1.5-7B (Biology-Instructions label)

Scopesubset · n=12851
Shots0
Turnssingle-turn
System prompt publicNo
Reasoning / effortThe prompt requests the task-formatted answer and says not to explain or repeat.
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNot reported
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
RepeatsNot reported
Graderfirst-number extraction followed by Spearman rank correlation · human review: no
StatisticsPoint metric over the complete published test split, scaled by 100 and rounded to two decimals; no confidence interval is reported.
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Spearman's ρabsolutepercentSpearman rank correlation across held-out test examples, scaled by 100Not reported

Results

ModelMetricValuen
LLaMA3.1-8B-Instruct (Biology-Instructions label)Spearman's ρ-0.61 percent
Sta creator-paper result; metric scaled by 100.
12851
Qwen2-7B (Biology-Instructions label)Spearman's ρ-5.86 percent
Sta creator-paper result; metric scaled by 100.
12851
Llama2-7B-Chat (Biology-Instructions label)Spearman's ρ-0.51 percent
Sta creator-paper result; metric scaled by 100.
12851
Alpaca-7B (Biology-Instructions label)Spearman's ρ2.05 percent
Sta creator-paper result; metric scaled by 100.
12851
GLM-4-9B-Chat (Biology-Instructions label)Spearman's ρ-2.72 percent
Sta creator-paper result; metric scaled by 100.
12851
Vicuna-v1.5-7B (Biology-Instructions label)Spearman's ρ5.65 percent
Sta creator-paper result; metric scaled by 100.
12851
Galactica-1.3B (Biology-Instructions label)Spearman's ρ-0.52 percent
Sta creator-paper result; metric scaled by 100.
12851
InstructProtein-1.3B (Biology-Instructions label)Spearman's ρ0.35 percent
Sta creator-paper result; metric scaled by 100.
12851
Llama-molinst-protein-7B (Mol-Ins)Spearman's ρ0.05 percent
Sta creator-paper result; metric scaled by 100.
12851
BioMedGPT-LM-7B (Biology-Instructions label)Spearman's ρ-0.92 percent
Sta creator-paper result; metric scaled by 100.
12851

Evidence

  • table: Table 2 (Sta test split); Appendix A.3; Table 8; Table 9 open-source prompt (Scope, prompt, output parser, grader, scaling, and aggregation.) — supports /scope, /benchmark_version, /model_ids, /protocol, /metrics
  • table: Table 6 (Sta column) (All registered model values; literature-SOTA row omitted.) — supports /results
Biology-Instructions Human Transcription Binding Sites Detection3 runs

Open benchmark record →

bioinstruction-tb-human-closed-baselines-emnlp-2025vemnlp-2025

Evaluated models / systems: GPT-4o (Biology-Instructions snapshot not reported), GPT-4o-mini (Biology-Instructions snapshot not reported)

Scopesubset · n=5000
Shots0
Turnssingle-turn
System prompt publicYes
Reasoning / effortThe prompt requests a direct JSON answer and no chain-of-thought.
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNot reported
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
RepeatsNot reported
Graderkeyword-first binary extraction with sentiment-model fallback, followed by MCC · model: cardiffnlp/twitter-roberta-base-sentiment-latest fallback · human review: no
StatisticsPoint metric over the complete published test split, scaled by 100 and rounded to two decimals; no confidence interval is reported.
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
MCCabsolutepercentMatthews correlation coefficient across held-out test examples, scaled by 100Not reported

Results

ModelMetricValuen
GPT-4o-mini (Biology-Instructions snapshot not reported)MCC0.14 percent
TB-H creator-paper result; metric scaled by 100.
5000
GPT-4o (Biology-Instructions snapshot not reported)MCC-1.7 percent
TB-H creator-paper result; metric scaled by 100.
5000

Evidence

  • table: Table 2 (TB-H test split); Appendix A.3; Table 8; Table 9 closed-source prompt (Scope, prompt, output parser, grader, scaling, and aggregation.) — supports /scope, /benchmark_version, /model_ids, /protocol, /metrics
  • table: Table 4 (TB-H column) (All registered model values; literature-SOTA row omitted.) — supports /results
bioinstruction-tb-human-creator-systems-emnlp-2025vemnlp-2025

Evaluated models / systems: ChatMultiOmics stage 1 + balanced stage 2, ChatMultiOmics stage 1 + stage 2, ChatMultiOmics, ChatMultiOmics stage 2 only

Scopesubset · n=5000
Shots0
Turnssingle-turn
System prompt publicYes
Reasoning / effortPsc requests clear, concise task answers and numeric output for regression tasks.
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNot reported
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
RepeatsNot reported
Graderkeyword-first binary extraction with sentiment-model fallback, followed by MCC · model: cardiffnlp/twitter-roberta-base-sentiment-latest fallback · human review: no
StatisticsPoint metric over the complete published test split, scaled by 100 and rounded to two decimals; no confidence interval is reported.
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
MCCabsolutepercentMatthews correlation coefficient across held-out test examples, scaled by 100Not reported

Results

ModelMetricValuen
ChatMultiOmics stage 1 + balanced stage 2MCC2.46 percent
TB-H creator-paper result; metric scaled by 100.
5000
ChatMultiOmics stage 2 onlyMCC0.86 percent
TB-H creator-paper result; metric scaled by 100.
5000
ChatMultiOmics stage 1 + stage 2MCC19.07 percent
TB-H creator-paper result; metric scaled by 100.
5000
ChatMultiOmicsMCC24.45 percent
TB-H creator-paper result; metric scaled by 100.
5000

Evidence

  • table: Table 2 (TB-H test split); Appendix A.3; Table 8; Section 4.2 Psc prompt (Scope, prompt, output parser, grader, scaling, and aggregation.) — supports /scope, /benchmark_version, /model_ids, /protocol, /metrics
  • table: Table 4 (TB-H column) (All registered model values; literature-SOTA row omitted.) — supports /results
bioinstruction-tb-human-open-baselines-emnlp-2025vemnlp-2025

Evaluated models / systems: Alpaca-7B (Biology-Instructions label), Galactica-1.3B (Biology-Instructions label), GLM-4-9B-Chat (Biology-Instructions label), InstructProtein-1.3B (Biology-Instructions label), Llama-molinst-protein-7B (Mol-Ins), Llama2-7B-Chat (Biology-Instructions label), LLaMA3.1-8B-Instruct (Biology-Instructions label), Qwen2-7B (Biology-Instructions label), Vicuna-v1.5-7B (Biology-Instructions label)

Scopesubset · n=5000
Shots0
Turnssingle-turn
System prompt publicNo
Reasoning / effortThe prompt requests the task-formatted answer and says not to explain or repeat.
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNot reported
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
RepeatsNot reported
Graderkeyword-first binary extraction with sentiment-model fallback, followed by MCC · model: cardiffnlp/twitter-roberta-base-sentiment-latest fallback · human review: no
StatisticsPoint metric over the complete published test split, scaled by 100 and rounded to two decimals; no confidence interval is reported.
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
MCCabsolutepercentMatthews correlation coefficient across held-out test examples, scaled by 100Not reported

Results

ModelMetricValuen
LLaMA3.1-8B-Instruct (Biology-Instructions label)MCC0 percent
TB-H creator-paper result; metric scaled by 100.
5000
Qwen2-7B (Biology-Instructions label)MCC-0.21 percent
TB-H creator-paper result; metric scaled by 100.
5000
Llama2-7B-Chat (Biology-Instructions label)MCC1.84 percent
TB-H creator-paper result; metric scaled by 100.
5000
Alpaca-7B (Biology-Instructions label)MCC2 percent
TB-H creator-paper result; metric scaled by 100.
5000
GLM-4-9B-Chat (Biology-Instructions label)MCC0 percent
TB-H creator-paper result; metric scaled by 100.
5000
Vicuna-v1.5-7B (Biology-Instructions label)MCC0 percent
TB-H creator-paper result; metric scaled by 100.
5000
Galactica-1.3B (Biology-Instructions label)MCC3 percent
TB-H creator-paper result; metric scaled by 100.
5000
InstructProtein-1.3B (Biology-Instructions label)MCC-1.29 percent
TB-H creator-paper result; metric scaled by 100.
5000
Llama-molinst-protein-7B (Mol-Ins)MCC2.4 percent
TB-H creator-paper result; metric scaled by 100.
5000

Evidence

  • table: Table 2 (TB-H test split); Appendix A.3; Table 8; Table 9 open-source prompt (Scope, prompt, output parser, grader, scaling, and aggregation.) — supports /scope, /benchmark_version, /model_ids, /protocol, /metrics
  • table: Table 4 (TB-H column) (All registered model values; literature-SOTA row omitted.) — supports /results
Biology-Instructions Mouse Transcription Binding Sites Detection3 runs

Open benchmark record →

bioinstruction-tb-mouse-closed-baselines-emnlp-2025vemnlp-2025

Evaluated models / systems: GPT-4o (Biology-Instructions snapshot not reported), GPT-4o-mini (Biology-Instructions snapshot not reported)

Scopesubset · n=10005
Shots0
Turnssingle-turn
System prompt publicYes
Reasoning / effortThe prompt requests a direct JSON answer and no chain-of-thought.
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNot reported
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
RepeatsNot reported
Graderkeyword-first binary extraction with sentiment-model fallback, followed by MCC · model: cardiffnlp/twitter-roberta-base-sentiment-latest fallback · human review: no
StatisticsPoint metric over the complete published test split, scaled by 100 and rounded to two decimals; no confidence interval is reported.
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
MCCabsolutepercentMatthews correlation coefficient across held-out test examples, scaled by 100Not reported

Results

ModelMetricValuen
GPT-4o-mini (Biology-Instructions snapshot not reported)MCC-0.31 percent
TB-M creator-paper result; metric scaled by 100.
10005
GPT-4o (Biology-Instructions snapshot not reported)MCC-1.38 percent
TB-M creator-paper result; metric scaled by 100.
10005

Evidence

  • table: Table 2 (TB-M test split); Appendix A.3; Table 8; Table 9 closed-source prompt (Scope, prompt, output parser, grader, scaling, and aggregation.) — supports /scope, /benchmark_version, /model_ids, /protocol, /metrics
  • table: Table 4 (TB-M column) (All registered model values; literature-SOTA row omitted.) — supports /results
bioinstruction-tb-mouse-creator-systems-emnlp-2025vemnlp-2025

Evaluated models / systems: ChatMultiOmics stage 1 + balanced stage 2, ChatMultiOmics stage 1 + stage 2, ChatMultiOmics, ChatMultiOmics stage 2 only

Scopesubset · n=10005
Shots0
Turnssingle-turn
System prompt publicYes
Reasoning / effortPsc requests clear, concise task answers and numeric output for regression tasks.
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNot reported
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
RepeatsNot reported
Graderkeyword-first binary extraction with sentiment-model fallback, followed by MCC · model: cardiffnlp/twitter-roberta-base-sentiment-latest fallback · human review: no
StatisticsPoint metric over the complete published test split, scaled by 100 and rounded to two decimals; no confidence interval is reported.
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
MCCabsolutepercentMatthews correlation coefficient across held-out test examples, scaled by 100Not reported

Results

ModelMetricValuen
ChatMultiOmics stage 1 + balanced stage 2MCC0.88 percent
TB-M creator-paper result; metric scaled by 100.
10005
ChatMultiOmics stage 2 onlyMCC0.13 percent
TB-M creator-paper result; metric scaled by 100.
10005
ChatMultiOmics stage 1 + stage 2MCC27.94 percent
TB-M creator-paper result; metric scaled by 100.
10005
ChatMultiOmicsMCC39.91 percent
TB-M creator-paper result; metric scaled by 100.
10005

Evidence

  • table: Table 2 (TB-M test split); Appendix A.3; Table 8; Section 4.2 Psc prompt (Scope, prompt, output parser, grader, scaling, and aggregation.) — supports /scope, /benchmark_version, /model_ids, /protocol, /metrics
  • table: Table 4 (TB-M column) (All registered model values; literature-SOTA row omitted.) — supports /results
bioinstruction-tb-mouse-open-baselines-emnlp-2025vemnlp-2025

Evaluated models / systems: Alpaca-7B (Biology-Instructions label), Galactica-1.3B (Biology-Instructions label), GLM-4-9B-Chat (Biology-Instructions label), InstructProtein-1.3B (Biology-Instructions label), Llama-molinst-protein-7B (Mol-Ins), Llama2-7B-Chat (Biology-Instructions label), LLaMA3.1-8B-Instruct (Biology-Instructions label), Qwen2-7B (Biology-Instructions label), Vicuna-v1.5-7B (Biology-Instructions label)

Scopesubset · n=10005
Shots0
Turnssingle-turn
System prompt publicNo
Reasoning / effortThe prompt requests the task-formatted answer and says not to explain or repeat.
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNot reported
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
RepeatsNot reported
Graderkeyword-first binary extraction with sentiment-model fallback, followed by MCC · model: cardiffnlp/twitter-roberta-base-sentiment-latest fallback · human review: no
StatisticsPoint metric over the complete published test split, scaled by 100 and rounded to two decimals; no confidence interval is reported.
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
MCCabsolutepercentMatthews correlation coefficient across held-out test examples, scaled by 100Not reported

Results

ModelMetricValuen
LLaMA3.1-8B-Instruct (Biology-Instructions label)MCC-1.42 percent
TB-M creator-paper result; metric scaled by 100.
10005
Qwen2-7B (Biology-Instructions label)MCC-1.59 percent
TB-M creator-paper result; metric scaled by 100.
10005
Llama2-7B-Chat (Biology-Instructions label)MCC0.97 percent
TB-M creator-paper result; metric scaled by 100.
10005
Alpaca-7B (Biology-Instructions label)MCC0 percent
TB-M creator-paper result; metric scaled by 100.
10005
GLM-4-9B-Chat (Biology-Instructions label)MCC0 percent
TB-M creator-paper result; metric scaled by 100.
10005
Vicuna-v1.5-7B (Biology-Instructions label)MCC0 percent
TB-M creator-paper result; metric scaled by 100.
10005
Galactica-1.3B (Biology-Instructions label)MCC-2.81 percent
TB-M creator-paper result; metric scaled by 100.
10005
InstructProtein-1.3B (Biology-Instructions label)MCC1.19 percent
TB-M creator-paper result; metric scaled by 100.
10005
Llama-molinst-protein-7B (Mol-Ins)MCC0.33 percent
TB-M creator-paper result; metric scaled by 100.
10005

Evidence

  • table: Table 2 (TB-M test split); Appendix A.3; Table 8; Table 9 open-source prompt (Scope, prompt, output parser, grader, scaling, and aggregation.) — supports /scope, /benchmark_version, /model_ids, /protocol, /metrics
  • table: Table 4 (TB-M column) (All registered model values; literature-SOTA row omitted.) — supports /results
Biology-Instructions Protein Thermostability Prediction3 runs

Open benchmark record →

bioinstruction-thermostability-closed-baselines-emnlp-2025vemnlp-2025

Evaluated models / systems: GPT-4o (Biology-Instructions snapshot not reported), GPT-4o-mini (Biology-Instructions snapshot not reported)

Scopesubset · n=1336
Shots0
Turnssingle-turn
System prompt publicYes
Reasoning / effortThe prompt requests a direct JSON answer and no chain-of-thought.
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNot reported
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
RepeatsNot reported
Graderfirst-number extraction followed by Spearman rank correlation · human review: no
StatisticsPoint metric over the complete published test split, scaled by 100 and rounded to two decimals; no confidence interval is reported.
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Spearman's ρabsolutepercentSpearman rank correlation across held-out test examples, scaled by 100Not reported

Results

ModelMetricValuen
GPT-4o-mini (Biology-Instructions snapshot not reported)Spearman's ρ0.32 percent
Ther creator-paper result; metric scaled by 100.
1336
GPT-4o (Biology-Instructions snapshot not reported)Spearman's ρ3.5 percent
Ther creator-paper result; metric scaled by 100.
1336

Evidence

  • table: Table 2 (Ther test split); Appendix A.3; Table 8; Table 9 closed-source prompt (Scope, prompt, output parser, grader, scaling, and aggregation.) — supports /scope, /benchmark_version, /model_ids, /protocol, /metrics
  • table: Table 6 (Ther column) (All registered model values; literature-SOTA row omitted.) — supports /results
bioinstruction-thermostability-creator-systems-emnlp-2025vemnlp-2025

Evaluated models / systems: ChatMultiOmics stage 1 + balanced stage 2, ChatMultiOmics stage 1 + stage 2, ChatMultiOmics, ChatMultiOmics stage 2 only

Scopesubset · n=1336
Shots0
Turnssingle-turn
System prompt publicYes
Reasoning / effortPsc requests clear, concise task answers and numeric output for regression tasks.
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNot reported
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
RepeatsNot reported
Graderfirst-number extraction followed by Spearman rank correlation · human review: no
StatisticsPoint metric over the complete published test split, scaled by 100 and rounded to two decimals; no confidence interval is reported.
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Spearman's ρabsolutepercentSpearman rank correlation across held-out test examples, scaled by 100Not reported

Results

ModelMetricValuen
ChatMultiOmics stage 1 + balanced stage 2Spearman's ρ39.97 percent
Ther creator-paper result; metric scaled by 100.
1336
ChatMultiOmics stage 2 onlySpearman's ρ-0.51 percent
Ther creator-paper result; metric scaled by 100.
1336
ChatMultiOmics stage 1 + stage 2Spearman's ρ44.59 percent
Ther creator-paper result; metric scaled by 100.
1336
ChatMultiOmicsSpearman's ρ45.07 percent
Ther creator-paper result; metric scaled by 100.
1336

Evidence

  • table: Table 2 (Ther test split); Appendix A.3; Table 8; Section 4.2 Psc prompt (Scope, prompt, output parser, grader, scaling, and aggregation.) — supports /scope, /benchmark_version, /model_ids, /protocol, /metrics
  • table: Table 6 (Ther column) (All registered model values; literature-SOTA row omitted.) — supports /results
bioinstruction-thermostability-open-baselines-emnlp-2025vemnlp-2025

Evaluated models / systems: Alpaca-7B (Biology-Instructions label), BioMedGPT-LM-7B (Biology-Instructions label), Galactica-1.3B (Biology-Instructions label), GLM-4-9B-Chat (Biology-Instructions label), InstructProtein-1.3B (Biology-Instructions label), Llama-molinst-protein-7B (Mol-Ins), Llama2-7B-Chat (Biology-Instructions label), LLaMA3.1-8B-Instruct (Biology-Instructions label), Qwen2-7B (Biology-Instructions label), Vicuna-v1.5-7B (Biology-Instructions label)

Scopesubset · n=1336
Shots0
Turnssingle-turn
System prompt publicNo
Reasoning / effortThe prompt requests the task-formatted answer and says not to explain or repeat.
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNot reported
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
RepeatsNot reported
Graderfirst-number extraction followed by Spearman rank correlation · human review: no
StatisticsPoint metric over the complete published test split, scaled by 100 and rounded to two decimals; no confidence interval is reported.
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Spearman's ρabsolutepercentSpearman rank correlation across held-out test examples, scaled by 100Not reported

Results

ModelMetricValuen
LLaMA3.1-8B-Instruct (Biology-Instructions label)Spearman's ρ4.67 percent
Ther creator-paper result; metric scaled by 100.
1336
Qwen2-7B (Biology-Instructions label)Spearman's ρ-0.93 percent
Ther creator-paper result; metric scaled by 100.
1336
Llama2-7B-Chat (Biology-Instructions label)Spearman's ρ0.4 percent
Ther creator-paper result; metric scaled by 100.
1336
Alpaca-7B (Biology-Instructions label)Spearman's ρ2.27 percent
Ther creator-paper result; metric scaled by 100.
1336
GLM-4-9B-Chat (Biology-Instructions label)Spearman's ρ1.4 percent
Ther creator-paper result; metric scaled by 100.
1336
Vicuna-v1.5-7B (Biology-Instructions label)Spearman's ρ0.9 percent
Ther creator-paper result; metric scaled by 100.
1336
Galactica-1.3B (Biology-Instructions label)Spearman's ρ-0.58 percent
Ther creator-paper result; metric scaled by 100.
1336
InstructProtein-1.3B (Biology-Instructions label)Spearman's ρ-0.5 percent
Ther creator-paper result; metric scaled by 100.
1336
Llama-molinst-protein-7B (Mol-Ins)Spearman's ρ1.07 percent
Ther creator-paper result; metric scaled by 100.
1336
BioMedGPT-LM-7B (Biology-Instructions label)Spearman's ρ-0.72 percent
Ther creator-paper result; metric scaled by 100.
1336

Evidence

  • table: Table 2 (Ther test split); Appendix A.3; Table 8; Table 9 open-source prompt (Scope, prompt, output parser, grader, scaling, and aggregation.) — supports /scope, /benchmark_version, /model_ids, /protocol, /metrics
  • table: Table 6 (Ther column) (All registered model values; literature-SOTA row omitted.) — supports /results