dataset · audited-with-caveats · verified 2026-07-21

ProteinLMBench

A creator-curated set of 944 protein-science multiple-choice questions with answer explanations, generated from research literature and released for evaluating text LLM protein understanding.

+3 more
Audited with caveats: 4 field(s) are marked provisional or conflicted. Warnings are shown next to affected values and these claims are excluded from unqualified comparisons.
Released-data audit: the current pinned JSON contains 944 questions, but only 871 have six choices; the remaining 73 have 2–5, 7, 8, or 10 choices. This conflicts with the paper's “944 six-choice questions” description. The released file also has no topical-category field, so design and binding counts remain Not reported.

Benchmark definition

What is counted

Version
hf-f139796
Total
944 (question records in ProteinLMBench.json)
Task formats
variable-choice multiple choice with 2 to 10 options in the current snapshot
Capabilities
KnowledgeScientific reasoningPrediction
Modalities
Text

Version history

VersionStatusRelease / as-ofTotalFormal tracks
hf-c59f90c
proteinlmbench-hf-c59f90c
superseded2024-04-29944 (question rows in the initial public CSV)None registered
hf-f139796
proteinlmbench-hf-f139796
current2024-05-23944 (question records in ProteinLMBench.json)None registered
paper-v2
proteinlmbench-paper-v2
active2024-07-08944 (questions described in creator preprint v2)None registered

Tracks and subsets

IDCountBasisPartition?Notes
Two-choice questions
proteinlmbench-two-choice
3questions grouped by options-array lengthExclusive & exhaustive
Three-choice questions
proteinlmbench-three-choice
21questions grouped by options-array lengthExclusive & exhaustive
Four-choice questions
proteinlmbench-four-choice
42questions grouped by options-array lengthExclusive & exhaustive
Five-choice questions
proteinlmbench-five-choice
1questions grouped by options-array lengthExclusive & exhaustive
Six-choice questions
proteinlmbench-six-choice
871Conflicted · highquestions grouped by options-array lengthExclusive & exhaustiveThe paper describes all 944 questions as six-choice; the current versioned JSON contains 871 six-choice records.
Seven-choice questions
proteinlmbench-seven-choice
2questions grouped by options-array lengthExclusive & exhaustiveBoth have answer option 7.
Eight-choice questions
proteinlmbench-eight-choice
1questions grouped by options-array lengthExclusive & exhaustive
Ten-choice questions
proteinlmbench-ten-choice
3questions grouped by options-array lengthExclusive & exhaustive

Scientific Task Atlas

Scientific task classification

partial for hf-f139796. Official materials claim broad topical coverage, but the released question file has no official topic labels; no topical counts are inferred.

Scientific taskCoverageCountMappingEvidence
Protein structure predictionexplicitly-in-scopeNot reported
Released ProteinLMBench question records.
official-taxonomy
high confidence
proteinlmbench-evidence-paper
Broad structure coverage only; folding is not inferred.
Protein designexplicitly-in-scopeNot reported
Released ProteinLMBench question records.
official-taxonomy
high confidence
proteinlmbench-evidence-paper
Broad protein-design topic; no sequence-generation count is available.
Molecular interaction and bindingobservedNot reported
Released ProteinLMBench question records.
official-taxonomy
high confidence
proteinlmbench-evidence-paper
Official questions discuss multiple binding contexts without a topic field or binding-subtype counts.

Scientific coverage notes

DomainCoverageCountInterpretation
Protein designexplicitly-in-scopeNot reportedThe paper dataset card names Protein design as a possible benchmark task, but the released evaluation file has no topical category field, so no design-question count is inferred.
Protein sequenceexplicitly-in-scopeNot reportedThe paper claims sequence-understanding coverage, but the current evaluation JSON exposes only question, options, answer, and explanation text fields and contains no explicit raw-sequence input field.
Protein-protein bindingobservedNot reportedOfficial questions discuss protein-protein interactions and antibody binding, but no official topical labels support a standalone count.
Protein-ligand bindingobservedNot reportedOfficial questions discuss ligand binding, but no official topical labels support a standalone count.

Evaluation registry

Works and run settings

A setting change—scope, prompt, tools, budget, grader, or repeats—creates a separate run. Charts never cross a comparability group.

proteinlmbench-paper-v2-full-officialvpaper-v2

Evaluated models / systems: Baichuan2-7B, ChatGLM3-6B, Falcon-7B, Falcon-7B-Instruct, GPT3.5-turbo (ProteinLMBench label), GPT4.0-turbo (ProteinLMBench label), InternLM-Chat-20B, InternLM2-20B, InternLM2-7B, InternLM2-Chat-20B, InternLM2-Chat-7B, InternLM2-Protein-7B (w/o SSL), Llama-2-7B-Chat-hf, Mistral-7B-Instruct-v0.2, Moonshot (ProteinLMBench label), Qwen1.5-7B, Yi-6B-Chat, InternLM2-Protein-7B

Scopefull · n=944
Shots0
Turnssingle-turn
System prompt publicNot applicable
Reasoning / effortThink step by step.
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNot reported
External toolsNo
Token budget20 generated tokens per question
Time / cost budgetNot reported
Temperature0.1
Seednot set
Repeats1
Graderfirst-integer exact option match · human review: no
Statisticssingle-run problem-weighted accuracy and total wall-clock inference minutes; no confidence intervals or error bars
ContaminationNo decontamination analysis reported; the paper says RAG generated questions and GPT-4 validated answers, while the later official repository says Mixtral-8x7B was used at each generation stage followed by expert review.
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Correct Rateabsolutepercentproblem-weighted over 944 questionsexact first-integer option match
Inference Timeabsoluteminutestotal wall-clock time over 944 questionsNot reported

Results

ModelMetricValuen
GPT4.0-turbo (ProteinLMBench label)Correct Rate57.94 percent
Exact API snapshot not reported.
944
GPT4.0-turbo (ProteinLMBench label)Inference Time15.52 minutes
Observed total; common hardware/service conditions not reported.
944
InternLM2-20BCorrect Rate57.52 percent
944
InternLM2-20BInference Time47.2 minutes
944
GPT3.5-turbo (ProteinLMBench label)Correct Rate55.19 percent
Exact API snapshot not reported.
944
GPT3.5-turbo (ProteinLMBench label)Inference Time21.03 minutes
944
InternLM2-7BCorrect Rate54.98 percent
944
InternLM2-7BInference Time19.23 minutes
944
InternLM2-Chat-7BCorrect Rate54.76 percent
944
InternLM2-Chat-7BInference Time35.58 minutes
944
InternLM2-Chat-20BCorrect Rate51.38 percent
944
InternLM2-Chat-20BInference Time31.11 minutes
944
Yi-6B-ChatCorrect Rate50.85 percent
944
Yi-6B-ChatInference Time59.05 minutes
944
Mistral-7B-Instruct-v0.2Correct Rate50.11 percent
944
Mistral-7B-Instruct-v0.2Inference Time13 minutes
944
ChatGLM3-6BCorrect Rate48.94 percent
944
ChatGLM3-6BInference Time8 minutes
944
Baichuan2-7BCorrect Rate44.49 percent
944
Baichuan2-7BInference Time16.37 minutes
944
InternLM-Chat-20BCorrect Rate40.54 percent
944
InternLM-Chat-20BInference Time66 minutes
944
Llama-2-7B-Chat-hfCorrect Rate39.64 percent
944
Llama-2-7B-Chat-hfInference Time64 minutes
944
Moonshot (ProteinLMBench label)Correct Rate38.26 percent
Exact provider model/version not reported.
944
Moonshot (ProteinLMBench label)Inference Time16.25 minutes
944
Qwen1.5-7BCorrect Rate21.73 percent
944
Qwen1.5-7BInference Time13 minutes
944
Falcon-7B-InstructCorrect Rate20.55 percent
944
Falcon-7B-InstructInference Time25.42 minutes
944
Falcon-7BCorrect Rate19.17 percent
944
Falcon-7BInference Time15.55 minutes
944
InternLM2-Protein-7BCorrect Rate62.18 percent
InternLM2-Protein-7B with SSL then SFT.
944
InternLM2-Protein-7BInference Time22.34 minutes
944
InternLM2-Protein-7B (w/o SSL)Correct Rate58.26 percent
SFT only; no ProteinLMDataset self-supervised phase.
944
InternLM2-Protein-7B (w/o SSL)Inference Time21.36 minutes
944

Evidence

  • section: Abstract; Sections 3.2, 4.3, 6; Appendix C.3 (944-question full paper evaluation and 18 evaluated model/system labels.) — supports /scope, /benchmark_version, /model_ids
  • repository-path: benchmark/benchmark_your_model.py at d8586e22ff85f6805edea0bbc23002aaccf525c4 (No demonstrations/tools, single prompt, think-step-by-step instruction, temperature 0.1, 20-token cap, first-integer parser, and exact scorer.) — supports /protocol
  • table: p. 23, Table 3; p. 14 checklist item 3(c) (All accuracy and inference-time values; checklist confirms no repeated experiments or random seeds and no error bars.) — supports /protocol/seed, /protocol/repeats, /protocol/statistical, /metrics, /results
  • section: Section 4.3 and Appendix B.3 Q22 (RAG/GPT-4 generation/validation and machine-plus-human verification; no decontamination analysis.) — supports /protocol/contamination

Comparable result views

Correct Rate

proteinlmbench-creator-full · proteinlmbench-paper-v2-full-official

CSV ↓
Accessible data table
ModelValueComparability group
GPT4.0-turbo (ProteinLMBench label)57.94proteinlmbench-paper-v2-full-official
InternLM2-20B57.52proteinlmbench-paper-v2-full-official
GPT3.5-turbo (ProteinLMBench label)55.19proteinlmbench-paper-v2-full-official
InternLM2-7B54.98proteinlmbench-paper-v2-full-official
InternLM2-Chat-7B54.76proteinlmbench-paper-v2-full-official
InternLM2-Chat-20B51.38proteinlmbench-paper-v2-full-official
Yi-6B-Chat50.85proteinlmbench-paper-v2-full-official
Mistral-7B-Instruct-v0.250.11proteinlmbench-paper-v2-full-official
ChatGLM3-6B48.94proteinlmbench-paper-v2-full-official
Baichuan2-7B44.49proteinlmbench-paper-v2-full-official
InternLM-Chat-20B40.54proteinlmbench-paper-v2-full-official
Llama-2-7B-Chat-hf39.64proteinlmbench-paper-v2-full-official
Moonshot (ProteinLMBench label)38.26proteinlmbench-paper-v2-full-official
Qwen1.5-7B21.73proteinlmbench-paper-v2-full-official
Falcon-7B-Instruct20.55proteinlmbench-paper-v2-full-official
Falcon-7B19.17proteinlmbench-paper-v2-full-official
InternLM2-Protein-7B62.18proteinlmbench-paper-v2-full-official
InternLM2-Protein-7B (w/o SSL)58.26proteinlmbench-paper-v2-full-official

Inference Time

proteinlmbench-creator-full · proteinlmbench-paper-v2-full-official

CSV ↓
Accessible data table
ModelValueComparability group
GPT4.0-turbo (ProteinLMBench label)15.52proteinlmbench-paper-v2-full-official
InternLM2-20B47.2proteinlmbench-paper-v2-full-official
GPT3.5-turbo (ProteinLMBench label)21.03proteinlmbench-paper-v2-full-official
InternLM2-7B19.23proteinlmbench-paper-v2-full-official
InternLM2-Chat-7B35.58proteinlmbench-paper-v2-full-official
InternLM2-Chat-20B31.11proteinlmbench-paper-v2-full-official
Yi-6B-Chat59.05proteinlmbench-paper-v2-full-official
Mistral-7B-Instruct-v0.213proteinlmbench-paper-v2-full-official
ChatGLM3-6B8proteinlmbench-paper-v2-full-official
Baichuan2-7B16.37proteinlmbench-paper-v2-full-official
InternLM-Chat-20B66proteinlmbench-paper-v2-full-official
Llama-2-7B-Chat-hf64proteinlmbench-paper-v2-full-official
Moonshot (ProteinLMBench label)16.25proteinlmbench-paper-v2-full-official
Qwen1.5-7B13proteinlmbench-paper-v2-full-official
Falcon-7B-Instruct25.42proteinlmbench-paper-v2-full-official
Falcon-7B15.55proteinlmbench-paper-v2-full-official
InternLM2-Protein-7B22.34proteinlmbench-paper-v2-full-official
InternLM2-Protein-7B (w/o SSL)21.36proteinlmbench-paper-v2-full-official

Evidence and change history

Source locators remain visible; expand an item to inspect the exact Registry fields it supports.

A Fine-tuning Dataset and Benchmark for Large Language Models for Protein Understanding · section: Abstract; Sections 3.2, 4.3, 5-6; Table 3; Appendix B.5-B.7 and C (944 count, claimed six-choice format, topic scope, construction/verification, evaluated models/results, no repeats/seeds, license statement, and paper-version snapshot.) · Supports 27 fields

Open source →

  • /name
  • /aliases
  • /summary
  • /kind
  • /organizations
  • /domains
  • /capabilities
  • /modalities
  • /task_formats
  • /task_counts/total
  • /task_counts/basis
  • /task_counts/subsets/4/count
  • /coverage_notes
  • /access/level
  • /access/tasks
  • /access/artifacts
  • /access/grader
  • /access/license
  • /access/biosafety_notes
  • /versions/1/task_counts/subsets/4/count
  • /versions/2/release_date
  • /versions/2/task_counts/total
  • /versions/2/task_counts/basis
  • /versions/2/task_counts/subsets
  • /scientific_task_classification/entries/0
  • /scientific_task_classification/entries/1
  • /scientific_task_classification/entries/2
proteinlmbench-dataset-resource · release: Hugging Face commits c59f90c91e215bae673574c57c9a6c4a9f6aa87b and f1397963c7f727a4a2f00cdd691e6e219c36e992 (Initial CSV release date/counts and later JSON restoration.) · Supports 6 fields

Open source →

  • /release_date
  • /versions/0/release_date
  • /versions/0/as_of
  • /versions/0/task_counts/total
  • /versions/0/task_counts/basis
  • /versions/0/task_counts/subsets
proteinlmbench-dataset-resource · repository-path: ProteinLMBench.json and README.md at f1397963c7f727a4a2f00cdd691e6e219c36e992 (944 records; fields question/options/answer/explanation; option counts 2:3, 3:21, 4:42, 5:1, 6:871, 7:2, 8:1, 10:3; Apache-2.0 card.) · Supports 18 fields

Open source →

  • /latest_version
  • /modalities
  • /task_formats
  • /task_counts/total
  • /task_counts/basis
  • /task_counts/subsets
  • /task_counts/subsets/4/count
  • /access/level
  • /access/tasks
  • /access/artifacts
  • /access/license
  • /resources
  • /versions/1/release_date
  • /versions/1/as_of
  • /versions/1/task_counts/total
  • /versions/1/task_counts/basis
  • /versions/1/task_counts/subsets
  • /versions/1/task_counts/subsets/4/count
proteinlmbench-repository-resource · repository-path: README.md; LICENSE; benchmark/benchmark_your_model.py at d8586e22ff85f6805edea0bbc23002aaccf525c4 (Manual expert review statement, public prompt-only runner, parser/grader, and Apache-2.0 license.) · Supports 5 fields

Open source →

  • /access/tasks
  • /access/grader
  • /access/license
  • /resources
  • /implementations

Unresolved field claims

  • /task_formatsConflicted · high — The paper calls every item six-choice, while the current official JSON contains 2-10 options per record.
    Evidence: proteinlmbench-evidence-paper, proteinlmbench-evidence-current-dataset
  • /task_counts/subsets/4/countConflicted · high — Current JSON has 871 six-choice records; paper v2 claims all 944 are six-choice.
    Evidence: proteinlmbench-evidence-paper, proteinlmbench-evidence-current-dataset
  • /versions/1/task_counts/subsets/4/countConflicted · high — Current JSON has 871 six-choice records; paper v2 claims all 944 are six-choice.
    Evidence: proteinlmbench-evidence-paper, proteinlmbench-evidence-current-dataset
  • /access/licenseConflicted · high — Current official dataset card and repository say Apache-2.0; the paper dataset card says Toursun Synbio metadata are CC BY 4.0. Both statements are retained.
    Evidence: proteinlmbench-evidence-paper, proteinlmbench-evidence-current-dataset

View source-level modification history on GitHub →