Evaluation run
lifescibench-official-full
From LifeSciBench: Evaluating Language Models on Realistic, Expert-Level Tasks in the Life Sciences
Evaluated models / systems: Gemini 3.1 Pro, GPT-5.4, GPT-5.5, GPT-Rosalind, Grok 4.3
Scopefull · n=750
ShotsNot reported
Turnssingle-turn
System prompt publicNot reported
Reasoning / effortNot reported
BrowserYes
InternetYes
DatabasesNot reported
Code executionNot reported
ContainerNot reported
External toolsNot reported
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
RepeatsNot reported
Gradertask-specific expert rubric; automated or model-assisted where used; expert spot-validation on a stratified response subset · human review: yes
Statisticsproblem-weighted mean with each task weighted equally
ContaminationNot reported
Metrics, results, and full protocol
Metrics
| Metric | Kind / baseline | Unit | Aggregation | Threshold / tolerance |
|---|---|---|---|---|
| Normalized rubric score | absolute | proportion | problem-weighted mean with each task weighted equally | Not reported |
| Task pass rate | absolute | percent | fraction of tasks whose normalized rubric score is at least 70% | 70 |
Results
| Model | Metric | Value | n |
|---|---|---|---|
| GPT-Rosalind | Normalized rubric score | 0.576 proportion n is the number of tasks; repeats are not reported. | 750 |
| GPT-Rosalind | Task pass rate | 36.1 percent n is the number of tasks; repeats are not reported. | 750 |
| GPT-5.5 | Normalized rubric score | 0.519 proportion n is the number of tasks; repeats are not reported. | 750 |
| GPT-5.5 | Task pass rate | 25.7 percent n is the number of tasks; repeats are not reported. | 750 |
| Gemini 3.1 Pro | Normalized rubric score | 0.515 proportion n is the number of tasks; repeats are not reported. | 750 |
| Gemini 3.1 Pro | Task pass rate | 23.6 percent n is the number of tasks; repeats are not reported. | 750 |
| GPT-5.4 | Normalized rubric score | 0.479 proportion n is the number of tasks; repeats are not reported. | 750 |
| GPT-5.4 | Task pass rate | 20.7 percent n is the number of tasks; repeats are not reported. | 750 |
| Grok 4.3 | Normalized rubric score | 0.399 proportion n is the number of tasks; repeats are not reported. | 750 |
| Grok 4.3 | Task pass rate | 13 percent n is the number of tasks; repeats are not reported. | 750 |
Evidence
- section: pp. 7 and 15, Sections 5.1 and 8 (States single-turn evaluation, unrestricted Internet browsing, and evaluation of five models across all 750 questions.) — supports /scope, /benchmark_version, /protocol/turns, /protocol/tools/browser, /protocol/tools/internet
- section: pp. 7–8, Sections 5.1–5.3 (Defines rubric grading, normalized score, 70% task threshold, problem weighting, automated/model-assisted grading, and expert spot-validation.) — supports /protocol/grader, /protocol/statistical, /metrics
- section: pp. 7–8 and 16, Sections 5.1–5.3 and Appendix A (Complete public protocol and disclosure sections do not state shots, system prompt, model effort, non-browser tools, budgets, temperature, seed, repeat count, or contamination analysis.) — supports /protocol/shots, /protocol/system_prompt_public, /protocol/reasoning, /protocol/tools/databases, /protocol/tools/code_execution, /protocol/tools/container, /protocol/tools/external_tools, /protocol/token_budget, /protocol/time_budget, /protocol/temperature, /protocol/seed, /protocol/repeats, /protocol/contamination
- section: p. 9, Section 6.1 and Figure 4 (Reports overall normalized score and task pass rate for all five models.) — supports /results