Protein & binding audit: 136 tasks use Protein and Structural Biology as the primary domain; 62 are in Design, Optimization & Prediction. Protein binding is explicitly in scope, but its standalone task count is Not reported.
Private/internal benchmark: tasks, counts, and core protocol are not public. This record is excluded from public and runnable coverage statistics. Exact labeled deltas remain inspectable, but are not converted into absolute scores or plotted as a ranking.
Benchmark definition
What is counted
- Version
- initial-release
- Total
- 750 (expert-authored tasks)
- Task formats
- expert free response; artifact-grounded analysis
- Capabilities
Evidence synthesisRetrievalDesignGenerationOptimizationData analysisExperiment planningTroubleshootingScientific reasoningScientific communication
- Modalities
TextPaper or documentTableFigureImageDNA or RNA sequenceProtein sequence3D structureRaw omicsWebWet-lab output
Version history
| Version | Status | Release / as-of | Total | Formal tracks |
|---|
initial-release
lifescibench-initial-release | current | 2026-06-17 | 750 (expert-authored tasks) | None registered |
Tracks and subsets
| ID | Count | Basis | Partition? | Notes |
|---|
Protein and structural biology primary domain
protein-primary-domain | 136 | tasks | No | Figure 13 column total for the Protein + Structural Biology domain. |
Protein-domain tasks in design and optimization workflow
protein-design-optimization | 62 | tasks | No | Figure 13 cell at Design and optimization × Protein + Structural Biology. |
Scientific Task Atlas
Scientific task classification
partial for initial-release. Official sources identify broad protein design and binding coverage but do not publish an exhaustive task-level taxonomy or binding subtype counts.
| Scientific task | Coverage | Count | Mapping | Evidence |
|---|
| Protein design | explicitly-in-scope | 62 tasks Expert-authored tasks in the Protein primary domain and Design / Optimization workflow cell. | official-taxonomy high confidence | lifescibench-evidence-counts This is a broad design-or-optimization count; it is not relabeled as sequence generation. |
| Protein-protein interaction prediction | explicitly-in-scope | Not reported Expert-authored benchmark tasks. | official-taxonomy high confidence | lifescibench-evidence-taxonomy The official protein-domain definition and examples include protein-protein binding, without a standalone count. |
| Protein-ligand binding prediction | explicitly-in-scope | Not reported Expert-authored benchmark tasks. | official-taxonomy high confidence | lifescibench-evidence-taxonomy The source does not distinguish pose from affinity, so the broad task is retained. |
Scientific coverage notes
| Domain | Coverage | Count | Interpretation |
|---|
| Protein science | observed | 136 | Figure 13 column total for the primary Protein + Structural Biology domain. |
| Protein design | observed | 62 | Figure 13 cross-tabulation of the protein domain and design/optimization workflow. |
| Protein-protein binding | explicitly-in-scope | Not reported | The official protein-domain definition includes binding and a protein-protein binding example, but no standalone binding-type count is published. |
| Protein-ligand binding | explicitly-in-scope | Not reported | The official protein-domain definition includes binding and ligand-related examples, but no standalone binding-type count is published. |
Evaluation registry
Works and run settings
A setting change—scope, prompt, tools, budget, grader, or repeats—creates a separate run. Charts never cross a comparability group.
lifescibench-initial-release-full-officialvinitial-release
Evaluated models / systems: Gemini 3.1 Pro, GPT-5.4, GPT-5.5, GPT-Rosalind, Grok 4.3
Scopefull · n=750
ShotsNot reported
Turnssingle-turn
System prompt publicNot reported
Reasoning / effortNot reported
BrowserYes
InternetYes
DatabasesNot reported
Code executionNot reported
ContainerNot reported
External toolsNot reported
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
RepeatsNot reported
Gradertask-specific expert rubric; automated or model-assisted where used; expert spot-validation on a stratified response subset · human review: yes
Statisticsproblem-weighted mean with each task weighted equally
ContaminationNot reported
Metrics, results, and full protocol
Metrics
| Metric | Kind / baseline | Unit | Aggregation | Threshold / tolerance |
|---|
| Normalized rubric score | absolute | proportion | problem-weighted mean with each task weighted equally | Not reported |
| Task pass rate | absolute | percent | fraction of tasks whose normalized rubric score is at least 70% | 70 |
Results
| Model | Metric | Value | n |
|---|
| GPT-Rosalind | Normalized rubric score | 0.576 proportion n is the number of tasks; repeats are not reported. | 750 |
| GPT-Rosalind | Task pass rate | 36.1 percent n is the number of tasks; repeats are not reported. | 750 |
| GPT-5.5 | Normalized rubric score | 0.519 proportion n is the number of tasks; repeats are not reported. | 750 |
| GPT-5.5 | Task pass rate | 25.7 percent n is the number of tasks; repeats are not reported. | 750 |
| Gemini 3.1 Pro | Normalized rubric score | 0.515 proportion n is the number of tasks; repeats are not reported. | 750 |
| Gemini 3.1 Pro | Task pass rate | 23.6 percent n is the number of tasks; repeats are not reported. | 750 |
| GPT-5.4 | Normalized rubric score | 0.479 proportion n is the number of tasks; repeats are not reported. | 750 |
| GPT-5.4 | Task pass rate | 20.7 percent n is the number of tasks; repeats are not reported. | 750 |
| Grok 4.3 | Normalized rubric score | 0.399 proportion n is the number of tasks; repeats are not reported. | 750 |
| Grok 4.3 | Task pass rate | 13 percent n is the number of tasks; repeats are not reported. | 750 |
Evidence
- section: pp. 7 and 15, Sections 5.1 and 8 (States single-turn evaluation, unrestricted Internet browsing, and evaluation of five models across all 750 questions.) — supports /scope, /benchmark_version, /protocol/turns, /protocol/tools/browser, /protocol/tools/internet
- section: pp. 7–8, Sections 5.1–5.3 (Defines rubric grading, normalized score, 70% task threshold, problem weighting, automated/model-assisted grading, and expert spot-validation.) — supports /protocol/grader, /protocol/statistical, /metrics
- section: pp. 7–8 and 16, Sections 5.1–5.3 and Appendix A (Complete public protocol and disclosure sections do not state shots, system prompt, model effort, non-browser tools, budgets, temperature, seed, repeat count, or contamination analysis.) — supports /protocol/shots, /protocol/system_prompt_public, /protocol/reasoning, /protocol/tools/databases, /protocol/tools/code_execution, /protocol/tools/container, /protocol/tools/external_tools, /protocol/token_budget, /protocol/time_budget, /protocol/temperature, /protocol/seed, /protocol/repeats, /protocol/contamination
- section: p. 9, Section 6.1 and Figure 4 (Reports overall normalized score and task pass rate for all five models.) — supports /results
Evidence and change history
Source locators remain visible; expand an item to inspect the exact Registry fields it supports.
lifescibench-launch-resource · section: Release date; introduction; What LifeSciBench measures; Dataset construction; Grading and rubric breakdown (Official release page.) · Supports 3 fields
Open source →
/release_date/latest_version/resources
LifeSciBench: Evaluating Language Models on Realistic, Expert-Level Tasks in the Life Sciences · section: pp. 1–4, title, affiliations, abstract, introduction, and Sections 3.1–3.3 (Official identity, scope, task structure, and creator affiliations.) · Supports 5 fields
Open source →
/name/organizations/kind/summary/task_formats
LifeSciBench: Evaluating Language Models on Realistic, Expert-Level Tasks in the Life Sciences · figure: pp. 6 and 18, Table 2 and Figure 13 (Explicit labels report 750 total, 136 Protein + Structural Biology, and 62 at Design and optimization × Protein + Structural Biology.) · Supports 9 fields
Open source →
/task_counts/total/task_counts/basis/task_counts/subsets/coverage_notes/0/coverage_notes/1/versions/0/task_counts/total/versions/0/task_counts/basis/versions/0/task_counts/subsets/scientific_task_classification/entries/0
LifeSciBench: Evaluating Language Models on Realistic, Expert-Level Tasks in the Life Sciences · section: pp. 3–4 and 17–18, Sections 3.1–3.3 and Appendix B.1–B.4; pp. 20–24, example tasks (Official workflow, domain, evidence-source taxonomies and public examples.) · Supports 7 fields
Open source →
/domains/capabilities/modalities/coverage_notes/2/coverage_notes/3/scientific_task_classification/entries/1/scientific_task_classification/entries/2
LifeSciBench: Evaluating Language Models on Realistic, Expert-Level Tasks in the Life Sciences · section: p. 16, Appendix A.5 Data Availability and Safety Disclosure (Describes licensing, privacy, proprietary-information, and biosafety release restrictions.) · Supports 6 fields
Open source →
/access/level/access/tasks/access/artifacts/access/grader/access/license/access/biosafety_notes
View source-level modification history on GitHub →