agentic-eval · audited · verified 2026-07-21

LifeSciBench

Expert-authored, artifact-rich free-response tasks that evaluate realistic research judgment across applied life-science workflows.

+12 more
Protein & binding audit: 136 tasks use Protein and Structural Biology as the primary domain; 62 are in Design, Optimization & Prediction. Protein binding is explicitly in scope, but its standalone task count is Not reported.
Private/internal benchmark: tasks, counts, and core protocol are not public. This record is excluded from public and runnable coverage statistics. Exact labeled deltas remain inspectable, but are not converted into absolute scores or plotted as a ranking.

Benchmark definition

What is counted

Version
initial-release
Total
750 (expert-authored tasks)
Task formats
expert free response; artifact-grounded analysis
Capabilities
Evidence synthesisRetrievalDesignGenerationOptimizationData analysisExperiment planningTroubleshootingScientific reasoningScientific communication
Modalities
TextPaper or documentTableFigureImageDNA or RNA sequenceProtein sequence3D structureRaw omicsWebWet-lab output

Version history

VersionStatusRelease / as-ofTotalFormal tracks
initial-release
lifescibench-initial-release
current2026-06-17750 (expert-authored tasks)None registered

Tracks and subsets

IDCountBasisPartition?Notes
Protein and structural biology primary domain
protein-primary-domain
136tasksNoFigure 13 column total for the Protein + Structural Biology domain.
Protein-domain tasks in design and optimization workflow
protein-design-optimization
62tasksNoFigure 13 cell at Design and optimization × Protein + Structural Biology.

Scientific Task Atlas

Scientific task classification

partial for initial-release. Official sources identify broad protein design and binding coverage but do not publish an exhaustive task-level taxonomy or binding subtype counts.

Scientific taskCoverageCountMappingEvidence
Protein designexplicitly-in-scope62 tasks
Expert-authored tasks in the Protein primary domain and Design / Optimization workflow cell.
official-taxonomy
high confidence
lifescibench-evidence-counts
This is a broad design-or-optimization count; it is not relabeled as sequence generation.
Protein-protein interaction predictionexplicitly-in-scopeNot reported
Expert-authored benchmark tasks.
official-taxonomy
high confidence
lifescibench-evidence-taxonomy
The official protein-domain definition and examples include protein-protein binding, without a standalone count.
Protein-ligand binding predictionexplicitly-in-scopeNot reported
Expert-authored benchmark tasks.
official-taxonomy
high confidence
lifescibench-evidence-taxonomy
The source does not distinguish pose from affinity, so the broad task is retained.

Scientific coverage notes

DomainCoverageCountInterpretation
Protein scienceobserved136Figure 13 column total for the primary Protein + Structural Biology domain.
Protein designobserved62Figure 13 cross-tabulation of the protein domain and design/optimization workflow.
Protein-protein bindingexplicitly-in-scopeNot reportedThe official protein-domain definition includes binding and a protein-protein binding example, but no standalone binding-type count is published.
Protein-ligand bindingexplicitly-in-scopeNot reportedThe official protein-domain definition includes binding and ligand-related examples, but no standalone binding-type count is published.

Evaluation registry

Works and run settings

A setting change—scope, prompt, tools, budget, grader, or repeats—creates a separate run. Charts never cross a comparability group.

lifescibench-initial-release-full-officialvinitial-release

Evaluated models / systems: Gemini 3.1 Pro, GPT-5.4, GPT-5.5, GPT-Rosalind, Grok 4.3

Scopefull · n=750
ShotsNot reported
Turnssingle-turn
System prompt publicNot reported
Reasoning / effortNot reported
BrowserYes
InternetYes
DatabasesNot reported
Code executionNot reported
ContainerNot reported
External toolsNot reported
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
RepeatsNot reported
Gradertask-specific expert rubric; automated or model-assisted where used; expert spot-validation on a stratified response subset · human review: yes
Statisticsproblem-weighted mean with each task weighted equally
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Normalized rubric scoreabsoluteproportionproblem-weighted mean with each task weighted equallyNot reported
Task pass rateabsolutepercentfraction of tasks whose normalized rubric score is at least 70%70

Results

ModelMetricValuen
GPT-RosalindNormalized rubric score0.576 proportion
n is the number of tasks; repeats are not reported.
750
GPT-RosalindTask pass rate36.1 percent
n is the number of tasks; repeats are not reported.
750
GPT-5.5Normalized rubric score0.519 proportion
n is the number of tasks; repeats are not reported.
750
GPT-5.5Task pass rate25.7 percent
n is the number of tasks; repeats are not reported.
750
Gemini 3.1 ProNormalized rubric score0.515 proportion
n is the number of tasks; repeats are not reported.
750
Gemini 3.1 ProTask pass rate23.6 percent
n is the number of tasks; repeats are not reported.
750
GPT-5.4Normalized rubric score0.479 proportion
n is the number of tasks; repeats are not reported.
750
GPT-5.4Task pass rate20.7 percent
n is the number of tasks; repeats are not reported.
750
Grok 4.3Normalized rubric score0.399 proportion
n is the number of tasks; repeats are not reported.
750
Grok 4.3Task pass rate13 percent
n is the number of tasks; repeats are not reported.
750

Evidence

  • section: pp. 7 and 15, Sections 5.1 and 8 (States single-turn evaluation, unrestricted Internet browsing, and evaluation of five models across all 750 questions.) — supports /scope, /benchmark_version, /protocol/turns, /protocol/tools/browser, /protocol/tools/internet
  • section: pp. 7–8, Sections 5.1–5.3 (Defines rubric grading, normalized score, 70% task threshold, problem weighting, automated/model-assisted grading, and expert spot-validation.) — supports /protocol/grader, /protocol/statistical, /metrics
  • section: pp. 7–8 and 16, Sections 5.1–5.3 and Appendix A (Complete public protocol and disclosure sections do not state shots, system prompt, model effort, non-browser tools, budgets, temperature, seed, repeat count, or contamination analysis.) — supports /protocol/shots, /protocol/system_prompt_public, /protocol/reasoning, /protocol/tools/databases, /protocol/tools/code_execution, /protocol/tools/container, /protocol/tools/external_tools, /protocol/token_budget, /protocol/time_budget, /protocol/temperature, /protocol/seed, /protocol/repeats, /protocol/contamination
  • section: p. 9, Section 6.1 and Figure 4 (Reports overall normalized score and task pass rate for all five models.) — supports /results

Evidence and change history

Source locators remain visible; expand an item to inspect the exact Registry fields it supports.

lifescibench-launch-resource · section: Release date; introduction; What LifeSciBench measures; Dataset construction; Grading and rubric breakdown (Official release page.) · Supports 3 fields

Open source →

  • /release_date
  • /latest_version
  • /resources
LifeSciBench: Evaluating Language Models on Realistic, Expert-Level Tasks in the Life Sciences · section: pp. 1–4, title, affiliations, abstract, introduction, and Sections 3.1–3.3 (Official identity, scope, task structure, and creator affiliations.) · Supports 5 fields

Open source →

  • /name
  • /organizations
  • /kind
  • /summary
  • /task_formats
LifeSciBench: Evaluating Language Models on Realistic, Expert-Level Tasks in the Life Sciences · figure: pp. 6 and 18, Table 2 and Figure 13 (Explicit labels report 750 total, 136 Protein + Structural Biology, and 62 at Design and optimization × Protein + Structural Biology.) · Supports 9 fields

Open source →

  • /task_counts/total
  • /task_counts/basis
  • /task_counts/subsets
  • /coverage_notes/0
  • /coverage_notes/1
  • /versions/0/task_counts/total
  • /versions/0/task_counts/basis
  • /versions/0/task_counts/subsets
  • /scientific_task_classification/entries/0
LifeSciBench: Evaluating Language Models on Realistic, Expert-Level Tasks in the Life Sciences · section: pp. 3–4 and 17–18, Sections 3.1–3.3 and Appendix B.1–B.4; pp. 20–24, example tasks (Official workflow, domain, evidence-source taxonomies and public examples.) · Supports 7 fields

Open source →

  • /domains
  • /capabilities
  • /modalities
  • /coverage_notes/2
  • /coverage_notes/3
  • /scientific_task_classification/entries/1
  • /scientific_task_classification/entries/2
LifeSciBench: Evaluating Language Models on Realistic, Expert-Level Tasks in the Life Sciences · section: p. 16, Appendix A.5 Data Availability and Safety Disclosure (Describes licensing, privacy, proprietary-information, and biosafety release restrictions.) · Supports 6 fields

Open source →

  • /access/level
  • /access/tasks
  • /access/artifacts
  • /access/grader
  • /access/license
  • /access/biosafety_notes

View source-level modification history on GitHub →