preprint · benchmark creator

A Fine-tuning Dataset and Benchmark for Large Language Models for Protein Understanding

Toursun Synbio · Johns Hopkins University · University of Cambridge · Shanghai Institute for Biomedical and Pharmaceutical Technologies · Shanghai AI Laboratory · Shanghai Jiao Tong University · UNSW Sydney · 2024-06-08

Relationship layer

Benchmark usage

This table records what the work did with each benchmark before attempting to normalize a run. Partial claims remain visible without being treated as comparable evaluations.

No BenchmarkUse relation is normalized for this legacy work yet. Existing EvaluationRuns remain available below.

Normalized evaluation runs

ProteinLMBench1 run

Open benchmark record →

proteinlmbench-paper-v2-full-officialvpaper-v2

Evaluated models / systems: Baichuan2-7B, ChatGLM3-6B, Falcon-7B, Falcon-7B-Instruct, GPT3.5-turbo (ProteinLMBench label), GPT4.0-turbo (ProteinLMBench label), InternLM-Chat-20B, InternLM2-20B, InternLM2-7B, InternLM2-Chat-20B, InternLM2-Chat-7B, InternLM2-Protein-7B (w/o SSL), Llama-2-7B-Chat-hf, Mistral-7B-Instruct-v0.2, Moonshot (ProteinLMBench label), Qwen1.5-7B, Yi-6B-Chat, InternLM2-Protein-7B

Scopefull · n=944
Shots0
Turnssingle-turn
System prompt publicNot applicable
Reasoning / effortThink step by step.
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNot reported
External toolsNo
Token budget20 generated tokens per question
Time / cost budgetNot reported
Temperature0.1
Seednot set
Repeats1
Graderfirst-integer exact option match · human review: no
Statisticssingle-run problem-weighted accuracy and total wall-clock inference minutes; no confidence intervals or error bars
ContaminationNo decontamination analysis reported; the paper says RAG generated questions and GPT-4 validated answers, while the later official repository says Mixtral-8x7B was used at each generation stage followed by expert review.
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Correct Rateabsolutepercentproblem-weighted over 944 questionsexact first-integer option match
Inference Timeabsoluteminutestotal wall-clock time over 944 questionsNot reported

Results

ModelMetricValuen
GPT4.0-turbo (ProteinLMBench label)Correct Rate57.94 percent
Exact API snapshot not reported.
944
GPT4.0-turbo (ProteinLMBench label)Inference Time15.52 minutes
Observed total; common hardware/service conditions not reported.
944
InternLM2-20BCorrect Rate57.52 percent
944
InternLM2-20BInference Time47.2 minutes
944
GPT3.5-turbo (ProteinLMBench label)Correct Rate55.19 percent
Exact API snapshot not reported.
944
GPT3.5-turbo (ProteinLMBench label)Inference Time21.03 minutes
944
InternLM2-7BCorrect Rate54.98 percent
944
InternLM2-7BInference Time19.23 minutes
944
InternLM2-Chat-7BCorrect Rate54.76 percent
944
InternLM2-Chat-7BInference Time35.58 minutes
944
InternLM2-Chat-20BCorrect Rate51.38 percent
944
InternLM2-Chat-20BInference Time31.11 minutes
944
Yi-6B-ChatCorrect Rate50.85 percent
944
Yi-6B-ChatInference Time59.05 minutes
944
Mistral-7B-Instruct-v0.2Correct Rate50.11 percent
944
Mistral-7B-Instruct-v0.2Inference Time13 minutes
944
ChatGLM3-6BCorrect Rate48.94 percent
944
ChatGLM3-6BInference Time8 minutes
944
Baichuan2-7BCorrect Rate44.49 percent
944
Baichuan2-7BInference Time16.37 minutes
944
InternLM-Chat-20BCorrect Rate40.54 percent
944
InternLM-Chat-20BInference Time66 minutes
944
Llama-2-7B-Chat-hfCorrect Rate39.64 percent
944
Llama-2-7B-Chat-hfInference Time64 minutes
944
Moonshot (ProteinLMBench label)Correct Rate38.26 percent
Exact provider model/version not reported.
944
Moonshot (ProteinLMBench label)Inference Time16.25 minutes
944
Qwen1.5-7BCorrect Rate21.73 percent
944
Qwen1.5-7BInference Time13 minutes
944
Falcon-7B-InstructCorrect Rate20.55 percent
944
Falcon-7B-InstructInference Time25.42 minutes
944
Falcon-7BCorrect Rate19.17 percent
944
Falcon-7BInference Time15.55 minutes
944
InternLM2-Protein-7BCorrect Rate62.18 percent
InternLM2-Protein-7B with SSL then SFT.
944
InternLM2-Protein-7BInference Time22.34 minutes
944
InternLM2-Protein-7B (w/o SSL)Correct Rate58.26 percent
SFT only; no ProteinLMDataset self-supervised phase.
944
InternLM2-Protein-7B (w/o SSL)Inference Time21.36 minutes
944

Evidence

  • section: Abstract; Sections 3.2, 4.3, 6; Appendix C.3 (944-question full paper evaluation and 18 evaluated model/system labels.) — supports /scope, /benchmark_version, /model_ids
  • repository-path: benchmark/benchmark_your_model.py at d8586e22ff85f6805edea0bbc23002aaccf525c4 (No demonstrations/tools, single prompt, think-step-by-step instruction, temperature 0.1, 20-token cap, first-integer parser, and exact scorer.) — supports /protocol
  • table: p. 23, Table 3; p. 14 checklist item 3(c) (All accuracy and inference-time values; checklist confirms no repeated experiments or random seeds and no error bars.) — supports /protocol/seed, /protocol/repeats, /protocol/statistical, /metrics, /results
  • section: Section 4.3 and Appendix B.3 Q22 (RAG/GPT-4 generation/validation and machine-plus-human verification; no decontamination analysis.) — supports /protocol/contamination