Evaluation run
proteinlmbench-creator-full
From A Fine-tuning Dataset and Benchmark for Large Language Models for Protein Understanding
Evaluated models / systems: Baichuan2-7B, ChatGLM3-6B, Falcon-7B, Falcon-7B-Instruct, GPT3.5-turbo (ProteinLMBench label), GPT4.0-turbo (ProteinLMBench label), InternLM-Chat-20B, InternLM2-20B, InternLM2-7B, InternLM2-Chat-20B, InternLM2-Chat-7B, InternLM2-Protein-7B (w/o SSL), Llama-2-7B-Chat-hf, Mistral-7B-Instruct-v0.2, Moonshot (ProteinLMBench label), Qwen1.5-7B, Yi-6B-Chat, InternLM2-Protein-7B
Metrics, results, and full protocol
Metrics
| Metric | Kind / baseline | Unit | Aggregation | Threshold / tolerance |
|---|---|---|---|---|
| Correct Rate | absolute | percent | problem-weighted over 944 questions | exact first-integer option match |
| Inference Time | absolute | minutes | total wall-clock time over 944 questions | Not reported |
Results
| Model | Metric | Value | n |
|---|---|---|---|
| GPT4.0-turbo (ProteinLMBench label) | Correct Rate | 57.94 percent Exact API snapshot not reported. | 944 |
| GPT4.0-turbo (ProteinLMBench label) | Inference Time | 15.52 minutes Observed total; common hardware/service conditions not reported. | 944 |
| InternLM2-20B | Correct Rate | 57.52 percent | 944 |
| InternLM2-20B | Inference Time | 47.2 minutes | 944 |
| GPT3.5-turbo (ProteinLMBench label) | Correct Rate | 55.19 percent Exact API snapshot not reported. | 944 |
| GPT3.5-turbo (ProteinLMBench label) | Inference Time | 21.03 minutes | 944 |
| InternLM2-7B | Correct Rate | 54.98 percent | 944 |
| InternLM2-7B | Inference Time | 19.23 minutes | 944 |
| InternLM2-Chat-7B | Correct Rate | 54.76 percent | 944 |
| InternLM2-Chat-7B | Inference Time | 35.58 minutes | 944 |
| InternLM2-Chat-20B | Correct Rate | 51.38 percent | 944 |
| InternLM2-Chat-20B | Inference Time | 31.11 minutes | 944 |
| Yi-6B-Chat | Correct Rate | 50.85 percent | 944 |
| Yi-6B-Chat | Inference Time | 59.05 minutes | 944 |
| Mistral-7B-Instruct-v0.2 | Correct Rate | 50.11 percent | 944 |
| Mistral-7B-Instruct-v0.2 | Inference Time | 13 minutes | 944 |
| ChatGLM3-6B | Correct Rate | 48.94 percent | 944 |
| ChatGLM3-6B | Inference Time | 8 minutes | 944 |
| Baichuan2-7B | Correct Rate | 44.49 percent | 944 |
| Baichuan2-7B | Inference Time | 16.37 minutes | 944 |
| InternLM-Chat-20B | Correct Rate | 40.54 percent | 944 |
| InternLM-Chat-20B | Inference Time | 66 minutes | 944 |
| Llama-2-7B-Chat-hf | Correct Rate | 39.64 percent | 944 |
| Llama-2-7B-Chat-hf | Inference Time | 64 minutes | 944 |
| Moonshot (ProteinLMBench label) | Correct Rate | 38.26 percent Exact provider model/version not reported. | 944 |
| Moonshot (ProteinLMBench label) | Inference Time | 16.25 minutes | 944 |
| Qwen1.5-7B | Correct Rate | 21.73 percent | 944 |
| Qwen1.5-7B | Inference Time | 13 minutes | 944 |
| Falcon-7B-Instruct | Correct Rate | 20.55 percent | 944 |
| Falcon-7B-Instruct | Inference Time | 25.42 minutes | 944 |
| Falcon-7B | Correct Rate | 19.17 percent | 944 |
| Falcon-7B | Inference Time | 15.55 minutes | 944 |
| InternLM2-Protein-7B | Correct Rate | 62.18 percent InternLM2-Protein-7B with SSL then SFT. | 944 |
| InternLM2-Protein-7B | Inference Time | 22.34 minutes | 944 |
| InternLM2-Protein-7B (w/o SSL) | Correct Rate | 58.26 percent SFT only; no ProteinLMDataset self-supervised phase. | 944 |
| InternLM2-Protein-7B (w/o SSL) | Inference Time | 21.36 minutes | 944 |
Evidence
- section: Abstract; Sections 3.2, 4.3, 6; Appendix C.3 (944-question full paper evaluation and 18 evaluated model/system labels.) — supports /scope, /benchmark_version, /model_ids
- repository-path: benchmark/benchmark_your_model.py at d8586e22ff85f6805edea0bbc23002aaccf525c4 (No demonstrations/tools, single prompt, think-step-by-step instruction, temperature 0.1, 20-token cap, first-integer parser, and exact scorer.) — supports /protocol
- table: p. 23, Table 3; p. 14 checklist item 3(c) (All accuracy and inference-time values; checklist confirms no repeated experiments or random seeds and no error bars.) — supports /protocol/seed, /protocol/repeats, /protocol/statistical, /metrics, /results
- section: Section 4.3 and Appendix B.3 Q22 (RAG/GPT-4 generation/validation and machine-plus-human verification; no decontamination analysis.) — supports /protocol/contamination