preprint · benchmark creator

LifeSciBench: Evaluating Language Models on Realistic, Expert-Level Tasks in the Life Sciences

OpenAI · Tacit Labs · 2026-06-17

Relationship layer

Benchmark usage

This table records what the work did with each benchmark before attempting to normalize a run. Partial claims remain visible without being treated as comparable evaluations.

No BenchmarkUse relation is normalized for this legacy work yet. Existing EvaluationRuns remain available below.

Normalized evaluation runs

LifeSciBench1 run

Open benchmark record →

lifescibench-initial-release-full-officialvinitial-release

Evaluated models / systems: Gemini 3.1 Pro, GPT-5.4, GPT-5.5, GPT-Rosalind, Grok 4.3

Scopefull · n=750
ShotsNot reported
Turnssingle-turn
System prompt publicNot reported
Reasoning / effortNot reported
BrowserYes
InternetYes
DatabasesNot reported
Code executionNot reported
ContainerNot reported
External toolsNot reported
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
RepeatsNot reported
Gradertask-specific expert rubric; automated or model-assisted where used; expert spot-validation on a stratified response subset · human review: yes
Statisticsproblem-weighted mean with each task weighted equally
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Normalized rubric scoreabsoluteproportionproblem-weighted mean with each task weighted equallyNot reported
Task pass rateabsolutepercentfraction of tasks whose normalized rubric score is at least 70%70

Results

ModelMetricValuen
GPT-RosalindNormalized rubric score0.576 proportion
n is the number of tasks; repeats are not reported.
750
GPT-RosalindTask pass rate36.1 percent
n is the number of tasks; repeats are not reported.
750
GPT-5.5Normalized rubric score0.519 proportion
n is the number of tasks; repeats are not reported.
750
GPT-5.5Task pass rate25.7 percent
n is the number of tasks; repeats are not reported.
750
Gemini 3.1 ProNormalized rubric score0.515 proportion
n is the number of tasks; repeats are not reported.
750
Gemini 3.1 ProTask pass rate23.6 percent
n is the number of tasks; repeats are not reported.
750
GPT-5.4Normalized rubric score0.479 proportion
n is the number of tasks; repeats are not reported.
750
GPT-5.4Task pass rate20.7 percent
n is the number of tasks; repeats are not reported.
750
Grok 4.3Normalized rubric score0.399 proportion
n is the number of tasks; repeats are not reported.
750
Grok 4.3Task pass rate13 percent
n is the number of tasks; repeats are not reported.
750

Evidence

  • section: pp. 7 and 15, Sections 5.1 and 8 (States single-turn evaluation, unrestricted Internet browsing, and evaluation of five models across all 750 questions.) — supports /scope, /benchmark_version, /protocol/turns, /protocol/tools/browser, /protocol/tools/internet
  • section: pp. 7–8, Sections 5.1–5.3 (Defines rubric grading, normalized score, 70% task threshold, problem weighting, automated/model-assisted grading, and expert spot-validation.) — supports /protocol/grader, /protocol/statistical, /metrics
  • section: pp. 7–8 and 16, Sections 5.1–5.3 and Appendix A (Complete public protocol and disclosure sections do not state shots, system prompt, model effort, non-browser tools, budgets, temperature, seed, repeat count, or contamination analysis.) — supports /protocol/shots, /protocol/system_prompt_public, /protocol/reasoning, /protocol/tools/databases, /protocol/tools/code_execution, /protocol/tools/container, /protocol/tools/external_tools, /protocol/token_budget, /protocol/time_budget, /protocol/temperature, /protocol/seed, /protocol/repeats, /protocol/contamination
  • section: p. 9, Section 6.1 and Figure 4 (Reports overall normalized score and task pass rate for all five models.) — supports /results