track · audited · verified 2026-07-21

Biology-Instructions Protein Fluorescence Prediction

Regression of protein fluorescence from an amino-acid sequence, evaluated with Spearman rank correlation.

Benchmark definition

What is counted

Version
emnlp-2025
Total
54025 (distinct examples across the published train, validation, and test splits)
Task formats
protein-sequence regression
Capabilities
PredictionRegression
Modalities
TextProtein sequence

Version history

VersionStatusRelease / as-ofTotalFormal tracks
emnlp-2025
bioinstruction-fluorescence-emnlp-2025
current2025-11-0454025 (distinct examples across the published train, validation, and test splits)None registered

Tracks and subsets

IDCountBasisPartition?Notes
Training split
bioinstruction-fluorescence-train
21446examplesExclusive & exhaustivePublished training split.
Validation split
bioinstruction-fluorescence-validation
5362examplesExclusive & exhaustivePublished validation split.
Test split
bioinstruction-fluorescence-test
27217examplesExclusive & exhaustiveHeld-out split used for creator-paper Tables 4-7.

Scientific Task Atlas

Scientific task classification

complete for emnlp-2025. Single-purpose formal Biology-Instructions evaluation track.

Scientific taskCoverageCountMappingEvidence
Protein fluorescence predictionexplicitly-in-scope54025 examples
distinct examples across the published train, validation, and test splits
official-track
high confidence
bioinstruction-fluorescence-evidence-paper
Protein fluorescence regression.

Evaluation registry

Works and run settings

A setting change—scope, prompt, tools, budget, grader, or repeats—creates a separate run. Charts never cross a comparability group.

bioinstruction-fluorescence-closed-baselines-emnlp-2025vemnlp-2025

Evaluated models / systems: GPT-4o (Biology-Instructions snapshot not reported), GPT-4o-mini (Biology-Instructions snapshot not reported)

Scopesubset · n=27217
Shots0
Turnssingle-turn
System prompt publicYes
Reasoning / effortThe prompt requests a direct JSON answer and no chain-of-thought.
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNot reported
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
RepeatsNot reported
Graderfirst-number extraction followed by Spearman rank correlation · human review: no
StatisticsPoint metric over the complete published test split, scaled by 100 and rounded to two decimals; no confidence interval is reported.
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Spearman's ρabsolutepercentSpearman rank correlation across held-out test examples, scaled by 100Not reported

Results

ModelMetricValuen
GPT-4o-mini (Biology-Instructions snapshot not reported)Spearman's ρ-0.47 percent
Flu creator-paper result; metric scaled by 100.
27217
GPT-4o (Biology-Instructions snapshot not reported)Spearman's ρ0.69 percent
Flu creator-paper result; metric scaled by 100.
27217

Evidence

  • table: Table 2 (Flu test split); Appendix A.3; Table 8; Table 9 closed-source prompt (Scope, prompt, output parser, grader, scaling, and aggregation.) — supports /scope, /benchmark_version, /model_ids, /protocol, /metrics
  • table: Table 6 (Flu column) (All registered model values; literature-SOTA row omitted.) — supports /results
bioinstruction-fluorescence-creator-systems-emnlp-2025vemnlp-2025

Evaluated models / systems: ChatMultiOmics stage 1 + balanced stage 2, ChatMultiOmics stage 1 + stage 2, ChatMultiOmics, ChatMultiOmics stage 2 only

Scopesubset · n=27217
Shots0
Turnssingle-turn
System prompt publicYes
Reasoning / effortPsc requests clear, concise task answers and numeric output for regression tasks.
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNot reported
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
RepeatsNot reported
Graderfirst-number extraction followed by Spearman rank correlation · human review: no
StatisticsPoint metric over the complete published test split, scaled by 100 and rounded to two decimals; no confidence interval is reported.
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Spearman's ρabsolutepercentSpearman rank correlation across held-out test examples, scaled by 100Not reported

Results

ModelMetricValuen
ChatMultiOmics stage 1 + balanced stage 2Spearman's ρ0.55 percent
Flu creator-paper result; metric scaled by 100.
27217
ChatMultiOmics stage 2 onlySpearman's ρ0.37 percent
Flu creator-paper result; metric scaled by 100.
27217
ChatMultiOmics stage 1 + stage 2Spearman's ρ1.49 percent
Flu creator-paper result; metric scaled by 100.
27217
ChatMultiOmicsSpearman's ρ2.57 percent
Flu creator-paper result; metric scaled by 100.
27217

Evidence

  • table: Table 2 (Flu test split); Appendix A.3; Table 8; Section 4.2 Psc prompt (Scope, prompt, output parser, grader, scaling, and aggregation.) — supports /scope, /benchmark_version, /model_ids, /protocol, /metrics
  • table: Table 6 (Flu column) (All registered model values; literature-SOTA row omitted.) — supports /results
bioinstruction-fluorescence-open-baselines-emnlp-2025vemnlp-2025

Evaluated models / systems: Alpaca-7B (Biology-Instructions label), BioMedGPT-LM-7B (Biology-Instructions label), Galactica-1.3B (Biology-Instructions label), GLM-4-9B-Chat (Biology-Instructions label), InstructProtein-1.3B (Biology-Instructions label), Llama-molinst-protein-7B (Mol-Ins), Llama2-7B-Chat (Biology-Instructions label), LLaMA3.1-8B-Instruct (Biology-Instructions label), Qwen2-7B (Biology-Instructions label), Vicuna-v1.5-7B (Biology-Instructions label)

Scopesubset · n=27217
Shots0
Turnssingle-turn
System prompt publicNo
Reasoning / effortThe prompt requests the task-formatted answer and says not to explain or repeat.
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNot reported
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
RepeatsNot reported
Graderfirst-number extraction followed by Spearman rank correlation · human review: no
StatisticsPoint metric over the complete published test split, scaled by 100 and rounded to two decimals; no confidence interval is reported.
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Spearman's ρabsolutepercentSpearman rank correlation across held-out test examples, scaled by 100Not reported

Results

ModelMetricValuen
LLaMA3.1-8B-Instruct (Biology-Instructions label)Spearman's ρ0.91 percent
Flu creator-paper result; metric scaled by 100.
27217
Qwen2-7B (Biology-Instructions label)Spearman's ρ0.81 percent
Flu creator-paper result; metric scaled by 100.
27217
Llama2-7B-Chat (Biology-Instructions label)Spearman's ρ0.28 percent
Flu creator-paper result; metric scaled by 100.
27217
Alpaca-7B (Biology-Instructions label)Spearman's ρ-0.2 percent
Flu creator-paper result; metric scaled by 100.
27217
GLM-4-9B-Chat (Biology-Instructions label)Spearman's ρ0.63 percent
Flu creator-paper result; metric scaled by 100.
27217
Vicuna-v1.5-7B (Biology-Instructions label)Spearman's ρ-0.51 percent
Flu creator-paper result; metric scaled by 100.
27217
Galactica-1.3B (Biology-Instructions label)Spearman's ρ-0.73 percent
Flu creator-paper result; metric scaled by 100.
27217
InstructProtein-1.3B (Biology-Instructions label)Spearman's ρ-0.03 percent
Flu creator-paper result; metric scaled by 100.
27217
Llama-molinst-protein-7B (Mol-Ins)Spearman's ρ0.27 percent
Flu creator-paper result; metric scaled by 100.
27217
BioMedGPT-LM-7B (Biology-Instructions label)Spearman's ρ0.43 percent
Flu creator-paper result; metric scaled by 100.
27217

Evidence

  • table: Table 2 (Flu test split); Appendix A.3; Table 8; Table 9 open-source prompt (Scope, prompt, output parser, grader, scaling, and aggregation.) — supports /scope, /benchmark_version, /model_ids, /protocol, /metrics
  • table: Table 6 (Flu column) (All registered model values; literature-SOTA row omitted.) — supports /results

Comparable result views

Spearman's ρ

bioinstruction-fluorescence-closed-baselines · bioinstruction-fluorescence-closed-baselines-emnlp-2025

CSV ↓
Accessible data table
ModelValueComparability group
GPT-4o-mini (Biology-Instructions snapshot not reported)-0.47bioinstruction-fluorescence-closed-baselines-emnlp-2025
GPT-4o (Biology-Instructions snapshot not reported)0.69bioinstruction-fluorescence-closed-baselines-emnlp-2025

Spearman's ρ

bioinstruction-fluorescence-creator-systems · bioinstruction-fluorescence-creator-systems-emnlp-2025

CSV ↓
Accessible data table
ModelValueComparability group
ChatMultiOmics stage 1 + balanced stage 20.55bioinstruction-fluorescence-creator-systems-emnlp-2025
ChatMultiOmics stage 2 only0.37bioinstruction-fluorescence-creator-systems-emnlp-2025
ChatMultiOmics stage 1 + stage 21.49bioinstruction-fluorescence-creator-systems-emnlp-2025
ChatMultiOmics2.57bioinstruction-fluorescence-creator-systems-emnlp-2025

Spearman's ρ

bioinstruction-fluorescence-open-baselines · bioinstruction-fluorescence-open-baselines-emnlp-2025

CSV ↓
Accessible data table
ModelValueComparability group
LLaMA3.1-8B-Instruct (Biology-Instructions label)0.91bioinstruction-fluorescence-open-baselines-emnlp-2025
Qwen2-7B (Biology-Instructions label)0.81bioinstruction-fluorescence-open-baselines-emnlp-2025
Llama2-7B-Chat (Biology-Instructions label)0.28bioinstruction-fluorescence-open-baselines-emnlp-2025
Alpaca-7B (Biology-Instructions label)-0.2bioinstruction-fluorescence-open-baselines-emnlp-2025
GLM-4-9B-Chat (Biology-Instructions label)0.63bioinstruction-fluorescence-open-baselines-emnlp-2025
Vicuna-v1.5-7B (Biology-Instructions label)-0.51bioinstruction-fluorescence-open-baselines-emnlp-2025
Galactica-1.3B (Biology-Instructions label)-0.73bioinstruction-fluorescence-open-baselines-emnlp-2025
InstructProtein-1.3B (Biology-Instructions label)-0.03bioinstruction-fluorescence-open-baselines-emnlp-2025
Llama-molinst-protein-7B (Mol-Ins)0.27bioinstruction-fluorescence-open-baselines-emnlp-2025
BioMedGPT-LM-7B (Biology-Instructions label)0.43bioinstruction-fluorescence-open-baselines-emnlp-2025

Evidence and change history

Source locators remain visible; expand an item to inspect the exact Registry fields it supports.

Biology-Instructions: A Dataset and Benchmark for Multi-Omics Sequence Understanding Capability of Large Language Models · table: Table 2 (Flu row); Appendix A.2-A.3; Table 8; Tables 4-7 (Task definition, split counts, input/output format, metric, and creator evaluation.) · Supports 29 fields

Open source →

  • /name
  • /aliases
  • /summary
  • /kind
  • /parent_id
  • /organizations
  • /release_date
  • /latest_version
  • /domains
  • /capabilities
  • /modalities
  • /task_formats
  • /task_counts/total
  • /task_counts/basis
  • /task_counts/subsets
  • /access/level
  • /access/tasks
  • /access/artifacts
  • /access/grader
  • /access/license
  • /access/biosafety_notes
  • /resources
  • /implementations
  • /versions/0/release_date
  • /versions/0/as_of
  • /versions/0/task_counts/total
  • /versions/0/task_counts/basis
  • /versions/0/task_counts/subsets
  • /scientific_task_classification/entries/0
bioinstruction-fluorescence-repository-resource · repository-path: evaluation/evaluate.py and evaluation/register_tasks.json at 600acaa08c0302e8f5ce86de0fe041f21c13b53e (Public grader implementation, partial artifact release, and absent repository license.) · Supports 7 fields

Open source →

  • /access/level
  • /access/tasks
  • /access/artifacts
  • /access/grader
  • /access/license
  • /resources
  • /implementations

View source-level modification history on GitHub →