Biology-Instructions
A multi-omics sequence-instruction suite with 21 formal predictive tasks across DNA, RNA, protein, and multi-molecule inputs; the final paper reports task-specific held-out evaluations rather than a single aggregate score.
Scientific domain
Amino-acid sequence understanding and modeling.
蛋白质序列
A multi-omics sequence-instruction suite with 21 formal predictive tasks across DNA, RNA, protein, and multi-molecule inputs; the final paper reports task-specific held-out evaluations rather than a single aggregate score.
Binary neutralization prediction for an antibody-antigen protein-sequence pair, evaluated with Matthews correlation coefficient.
Multi-label Enzyme Commission number prediction from a protein sequence, evaluated with the creator's Fmax implementation.
Regression of protein fluorescence from an amino-acid sequence, evaluated with Spearman rank correlation.
Binary interaction prediction for an RNA and protein sequence pair, evaluated with Matthews correlation coefficient.
Binary protein-solubility prediction from an amino-acid sequence, evaluated with accuracy.
Regression of protein stability from an amino-acid sequence, evaluated with Spearman rank correlation.
Regression of protein thermostability from an amino-acid sequence, evaluated with Spearman rank correlation.
A multistate protein sequence-design benchmark spanning CaM conformations and binding modes.
A supervised protein sequence-to-fitness benchmark that turns three experimental landscapes into 15 biologically motivated dataset splits for testing generalization in protein engineering.
Seven supervised splits over sampled and machine-designed AAV2 VP-1 capsid variants, measuring generalization across mutation depth, fitness, and sampled-versus-designed pools.
Five supervised splits over a downsampled, highly epistatic four-site GB1 immunoglobulin-binding landscape, designed to test mutation-depth and low-to-high-fitness generalization.
Three supervised sequence-to-melting-temperature splits spanning all species, human proteins, and a single human cell line, with sequence-cluster-aware train/test separation.
A practical biology-research suite of 2,457 multiple-choice questions across eight broad categories and 31 versioned task files, with public and private contamination-monitoring splits.
Database-retrieval category spanning 10 genomics, clinical, protein, regulatory, vaccine-response, and viral-PPI tasks.
Uses a protein sequence and ClinVar lookup to identify benign or pathogenic variants.
Identifies ClinVar variant pathogenicity while reasoning across multiple protein sequences.
Sequence-comprehension and manipulation category spanning 15 formal tasks involving PCR, restriction digestion, ORFs, translation, GC content, and DNA–protein relationships.
Finds the amino acid encoded at a specified position in the longest ORF of a DNA sequence.
Translates the longest ORF in a DNA sequence to its amino-acid sequence.
Counts open reading frames encoding proteins above a specified amino-acid length.
Selects an RNA sequence whose ORF context is most likely to yield high translation efficiency.
Expert-authored, artifact-rich free-response tasks that evaluate realistic research judgment across applied life-science workflows.
A multistate protein sequence-design benchmark targeting the multispecific PapD binding interface.
Versioned deep-mutational-scanning and clinical-variant benchmarks for protein fitness prediction and design in zero-shot and supervised regimes.
ProteinGym track for classifying short human clinical insertion and deletion variants against ClinVar and gnomAD-derived labels.
ProteinGym track for classifying expert-annotated human clinical substitution variants on a per-protein basis.
ProteinGym track for predicting experimental fitness measurements of insertion and deletion mutants across deep-mutational-scanning assays.
ProteinGym track for predicting experimental fitness measurements of substitution mutants across deep-mutational-scanning assays.
A creator-curated set of 944 protein-science multiple-choice questions with answer explanations, generated from research literature and released for evaluating text LLM protein understanding.
A multistate protein sequence-design benchmark using the fold-switching conformations of RfaH.
A five-task benchmark for protein representation learning spanning secondary structure, residue contacts, remote homology, fluorescence, and stability.
Registry records tagged Protein sequence, counted by capability.
| Capability | Records |
|---|---|
| Knowledge | 3 |
| Evidence synthesis | 2 |
| Retrieval | 5 |
| Prediction | 23 |
| Classification | 15 |
| Regression | 12 |
| Design | 9 |
| Generation | 1 |
| Optimization | 7 |
| Data analysis | 6 |
| Tool use | 2 |
| Experiment planning | 2 |
| Troubleshooting | 2 |
| Scientific reasoning | 12 |
| Scientific communication | 1 |
Task mappings are evidence-backed and may be partial for mixed suites.