ProteinLMBench
A creator-curated set of 944 protein-science multiple-choice questions with answer explanations, generated from research literature and released for evaluating text LLM protein understanding.
Scientific task · Design and generation
Generate or optimize proteins under structural
蛋白质设计
protein-designCoverage
A creator-curated set of 944 protein-science multiple-choice questions with answer explanations, generated from research literature and released for evaluating text LLM protein understanding.
A multistate protein sequence-design benchmark spanning CaM conformations and binding modes.
A multistate protein sequence-design benchmark targeting the multispecific PapD binding interface.
A multistate protein sequence-design benchmark using the fold-switching conformations of RfaH.
Each row keeps its original unit and basis. Rows with different units or overlapping mappings are never added.
| Benchmark | Mapped task | Coverage | Count | Version | Evidence |
|---|---|---|---|---|---|
| LifeSciBench root: lifescibench | Protein design official-taxonomy · high | explicitly-in-scope | 62 tasks Expert-authored tasks in the Protein primary domain and Design / Optimization workflow cell. | initial-release | lifescibench-evidence-counts |
| ProteinLMBench root: proteinlmbench | Protein design official-taxonomy · high | explicitly-in-scope | Not reported Released ProteinLMBench question records. | hf-f139796 | proteinlmbench-evidence-paper |
| Biology-Instructions root: bioinstruction | Protein sequence design official-taxonomy · high | not-in-scope | 0 tracks Formal evaluation tracks in the final creator paper. | emnlp-2025 | bioinstruction-evidence-paper |
| CaM benchmark root: cam-benchmark | Protein sequence design official-taxonomy · high | explicitly-in-scope | Not reported Creator-defined multistate protein sequence-design benchmark system. | initial-release | cam-benchmark-automated-metadata-2-evidence |
| LAB-Bench root: lab-bench | Protein sequence design official-taxonomy · high | not-in-scope | 0 questions Released LAB-Bench questions. | repository-998a8e0 | lab-bench-evidence-paper |
| PapD benchmark root: papd-benchmark | Protein sequence design official-taxonomy · high | explicitly-in-scope | Not reported Creator-defined multistate protein sequence-design benchmark system. | initial-release | papd-benchmark-automated-metadata-2-evidence |
| RfaH benchmark root: rfah-benchmark | Protein sequence design official-taxonomy · high | explicitly-in-scope | Not reported Creator-defined multistate protein sequence-design benchmark system. | initial-release | rfah-benchmark-automated-metadata-2-evidence |
Runs are included only for benchmark records mapped here (and formal child tracks when a mapped suite is the root). A task mapping does not imply that every run isolates this task.
| Work | Provider / class | Related runs |
|---|---|---|
| A Fine-tuning Dataset and Benchmark for Large Language Models for Protein Understanding | Toursun Synbio, Johns Hopkins University, University of Cambridge, Shanghai Institute for Biomedical and Pharmaceutical Technologies, Shanghai AI Laboratory, Shanghai Jiao Tong University, UNSW Sydney benchmark_creator | proteinlmbench-creator-full |