track · audited · verified 2026-07-21
Biology-Instructions Epigenetic Marks Prediction
Binary prediction of whether a DNA sequence carries an epigenetic mark, evaluated with Matthews correlation coefficient.
Benchmark definition
What is counted
- Version
- emnlp-2025
- Total
- 287367 (distinct examples across the published train, validation, and test splits)
- Task formats
- DNA-sequence binary classification
- Capabilities
PredictionClassification
- Modalities
TextDNA or RNA sequence
Version history
| Version | Status | Release / as-of | Total | Formal tracks |
|---|
emnlp-2025
bioinstruction-emp-emnlp-2025 | current | 2025-11-04 | 287367 (distinct examples across the published train, validation, and test splits) | None registered |
Tracks and subsets
| ID | Count | Basis | Partition? | Notes |
|---|
Training split
bioinstruction-emp-train | 229885 | examples | Exclusive & exhaustive | Published training split. |
Validation split
bioinstruction-emp-validation | 28741 | examples | Exclusive & exhaustive | Published validation split. |
Test split
bioinstruction-emp-test | 28741 | examples | Exclusive & exhaustive | Held-out split used for creator-paper Tables 4-7. |
Scientific Task Atlas
Scientific task classification
complete for emnlp-2025. Single-purpose formal Biology-Instructions evaluation track.
| Scientific task | Coverage | Count | Mapping | Evidence |
|---|
| Epigenetic-mark prediction | explicitly-in-scope | 287367 examples distinct examples across the published train, validation, and test splits | official-track high confidence | bioinstruction-emp-evidence-paper Epigenetic-mark prediction. |
Evaluation registry
Works and run settings
A setting change—scope, prompt, tools, budget, grader, or repeats—creates a separate run. Charts never cross a comparability group.
bioinstruction-emp-closed-baselines-emnlp-2025vemnlp-2025
Evaluated models / systems: GPT-4o (Biology-Instructions snapshot not reported), GPT-4o-mini (Biology-Instructions snapshot not reported)
Scopesubset · n=28741
Shots0
Turnssingle-turn
System prompt publicYes
Reasoning / effortThe prompt requests a direct JSON answer and no chain-of-thought.
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNot reported
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
RepeatsNot reported
Graderkeyword-first binary extraction with sentiment-model fallback, followed by MCC · model: cardiffnlp/twitter-roberta-base-sentiment-latest fallback · human review: no
StatisticsPoint metric over the complete published test split, scaled by 100 and rounded to two decimals; no confidence interval is reported.
ContaminationNot reported
Metrics, results, and full protocol
Metrics
| Metric | Kind / baseline | Unit | Aggregation | Threshold / tolerance |
|---|
| MCC | absolute | percent | Matthews correlation coefficient across held-out test examples, scaled by 100 | Not reported |
Results
Evidence
- table: Table 2 (EMP test split); Appendix A.3; Table 8; Table 9 closed-source prompt (Scope, prompt, output parser, grader, scaling, and aggregation.) — supports /scope, /benchmark_version, /model_ids, /protocol, /metrics
- table: Table 4 (EMP column) (All registered model values; literature-SOTA row omitted.) — supports /results
bioinstruction-emp-creator-systems-emnlp-2025vemnlp-2025
Evaluated models / systems: ChatMultiOmics stage 1 + balanced stage 2, ChatMultiOmics stage 1 + stage 2, ChatMultiOmics, ChatMultiOmics stage 2 only
Scopesubset · n=28741
Shots0
Turnssingle-turn
System prompt publicYes
Reasoning / effortPsc requests clear, concise task answers and numeric output for regression tasks.
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNot reported
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
RepeatsNot reported
Graderkeyword-first binary extraction with sentiment-model fallback, followed by MCC · model: cardiffnlp/twitter-roberta-base-sentiment-latest fallback · human review: no
StatisticsPoint metric over the complete published test split, scaled by 100 and rounded to two decimals; no confidence interval is reported.
ContaminationNot reported
Metrics, results, and full protocol
Metrics
| Metric | Kind / baseline | Unit | Aggregation | Threshold / tolerance |
|---|
| MCC | absolute | percent | Matthews correlation coefficient across held-out test examples, scaled by 100 | Not reported |
Results
Evidence
- table: Table 2 (EMP test split); Appendix A.3; Table 8; Section 4.2 Psc prompt (Scope, prompt, output parser, grader, scaling, and aggregation.) — supports /scope, /benchmark_version, /model_ids, /protocol, /metrics
- table: Table 4 (EMP column) (All registered model values; literature-SOTA row omitted.) — supports /results
bioinstruction-emp-open-baselines-emnlp-2025vemnlp-2025
Evaluated models / systems: Alpaca-7B (Biology-Instructions label), Galactica-1.3B (Biology-Instructions label), GLM-4-9B-Chat (Biology-Instructions label), InstructProtein-1.3B (Biology-Instructions label), Llama-molinst-protein-7B (Mol-Ins), Llama2-7B-Chat (Biology-Instructions label), LLaMA3.1-8B-Instruct (Biology-Instructions label), Qwen2-7B (Biology-Instructions label), Vicuna-v1.5-7B (Biology-Instructions label)
Scopesubset · n=28741
Shots0
Turnssingle-turn
System prompt publicNo
Reasoning / effortThe prompt requests the task-formatted answer and says not to explain or repeat.
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNot reported
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
RepeatsNot reported
Graderkeyword-first binary extraction with sentiment-model fallback, followed by MCC · model: cardiffnlp/twitter-roberta-base-sentiment-latest fallback · human review: no
StatisticsPoint metric over the complete published test split, scaled by 100 and rounded to two decimals; no confidence interval is reported.
ContaminationNot reported
Metrics, results, and full protocol
Metrics
| Metric | Kind / baseline | Unit | Aggregation | Threshold / tolerance |
|---|
| MCC | absolute | percent | Matthews correlation coefficient across held-out test examples, scaled by 100 | Not reported |
Results
Evidence
- table: Table 2 (EMP test split); Appendix A.3; Table 8; Table 9 open-source prompt (Scope, prompt, output parser, grader, scaling, and aggregation.) — supports /scope, /benchmark_version, /model_ids, /protocol, /metrics
- table: Table 4 (EMP column) (All registered model values; literature-SOTA row omitted.) — supports /results
Comparable result views
Accessible data table
| Model | Value | Comparability group |
|---|
| GPT-4o-mini (Biology-Instructions snapshot not reported) | -0.91 | bioinstruction-emp-closed-baselines-emnlp-2025 |
| GPT-4o (Biology-Instructions snapshot not reported) | -0.49 | bioinstruction-emp-closed-baselines-emnlp-2025 |
Accessible data table
| Model | Value | Comparability group |
|---|
| ChatMultiOmics stage 1 + balanced stage 2 | 1.4 | bioinstruction-emp-creator-systems-emnlp-2025 |
| ChatMultiOmics stage 2 only | 0.31 | bioinstruction-emp-creator-systems-emnlp-2025 |
| ChatMultiOmics stage 1 + stage 2 | 8.1 | bioinstruction-emp-creator-systems-emnlp-2025 |
| ChatMultiOmics | 3.64 | bioinstruction-emp-creator-systems-emnlp-2025 |
Accessible data table
| Model | Value | Comparability group |
|---|
| LLaMA3.1-8B-Instruct (Biology-Instructions label) | -0.37 | bioinstruction-emp-open-baselines-emnlp-2025 |
| Qwen2-7B (Biology-Instructions label) | -0.66 | bioinstruction-emp-open-baselines-emnlp-2025 |
| Llama2-7B-Chat (Biology-Instructions label) | 0.94 | bioinstruction-emp-open-baselines-emnlp-2025 |
| Alpaca-7B (Biology-Instructions label) | -0.36 | bioinstruction-emp-open-baselines-emnlp-2025 |
| GLM-4-9B-Chat (Biology-Instructions label) | -0.22 | bioinstruction-emp-open-baselines-emnlp-2025 |
| Vicuna-v1.5-7B (Biology-Instructions label) | 0 | bioinstruction-emp-open-baselines-emnlp-2025 |
| Galactica-1.3B (Biology-Instructions label) | 0.07 | bioinstruction-emp-open-baselines-emnlp-2025 |
| InstructProtein-1.3B (Biology-Instructions label) | 0.22 | bioinstruction-emp-open-baselines-emnlp-2025 |
| Llama-molinst-protein-7B (Mol-Ins) | -0.29 | bioinstruction-emp-open-baselines-emnlp-2025 |
Evidence and change history
Source locators remain visible; expand an item to inspect the exact Registry fields it supports.
Biology-Instructions: A Dataset and Benchmark for Multi-Omics Sequence Understanding Capability of Large Language Models · table: Table 2 (EMP row); Appendix A.2-A.3; Table 8; Tables 4-7 (Task definition, split counts, input/output format, metric, and creator evaluation.) · Supports 29 fields
Open source →
/name/aliases/summary/kind/parent_id/organizations/release_date/latest_version/domains/capabilities/modalities/task_formats/task_counts/total/task_counts/basis/task_counts/subsets/access/level/access/tasks/access/artifacts/access/grader/access/license/access/biosafety_notes/resources/implementations/versions/0/release_date/versions/0/as_of/versions/0/task_counts/total/versions/0/task_counts/basis/versions/0/task_counts/subsets/scientific_task_classification/entries/0
bioinstruction-emp-repository-resource · repository-path: evaluation/evaluate.py and evaluation/register_tasks.json at 600acaa08c0302e8f5ce86de0fe041f21c13b53e (Public grader implementation, partial artifact release, and absent repository license.) · Supports 7 fields
Open source →
/access/level/access/tasks/access/artifacts/access/grader/access/license/resources/implementations
View source-level modification history on GitHub →