LAB-Bench DbQA — Disease gene associations
Identifies genes associated with a phenotype in DisGeNET but not OMIM.
1 evaluation run(s)
track · audited · verified 2026-07-21
Database-retrieval category spanning 10 genomics, clinical, protein, regulatory, vaccine-response, and viral-PPI tasks.
Benchmark definition
| Version | Status | Release / as-of | Total | Formal tracks |
|---|---|---|---|---|
paper-v3lab-bench-dbqa-paper-v3 | active | 2024-07-17 | 650 (questions across formal child tasks) | lab-bench-dbqa-dga, lab-bench-dbqa-gene-location, lab-bench-dbqa-mirna-targets, lab-bench-dbqa-mouse-tumor-gene-sets, lab-bench-dbqa-oncogenic-signatures, lab-bench-dbqa-tfbs-gtrd, lab-bench-dbqa-variant-from-sequence, lab-bench-dbqa-variant-multi-sequence, lab-bench-dbqa-vax-response, lab-bench-dbqa-viral-ppi |
repository-998a8e0lab-bench-dbqa-repository-998a8e0 | current | 2025-09-27 | 650 (questions across formal child tasks) | lab-bench-dbqa-dga, lab-bench-dbqa-gene-location, lab-bench-dbqa-mirna-targets, lab-bench-dbqa-mouse-tumor-gene-sets, lab-bench-dbqa-oncogenic-signatures, lab-bench-dbqa-tfbs-gtrd, lab-bench-dbqa-variant-from-sequence, lab-bench-dbqa-variant-multi-sequence, lab-bench-dbqa-vax-response, lab-bench-dbqa-viral-ppi |
| ID | Count | Basis | Partition? | Notes |
|---|---|---|---|---|
Disease gene associationslab-bench-dbqa-dga | 50 | questions | Exclusive & exhaustive | — |
Gene locationlab-bench-dbqa-gene-location | 50 | questions | Exclusive & exhaustive | — |
miRNA targetslab-bench-dbqa-mirna-targets | 50 | questions | Exclusive & exhaustive | — |
Mouse tumor gene setslab-bench-dbqa-mouse-tumor-gene-sets | 100 | questions | Exclusive & exhaustive | — |
Oncogenic signatureslab-bench-dbqa-oncogenic-signatures | 50 | questions | Exclusive & exhaustive | — |
GTRD transcription-factor binding siteslab-bench-dbqa-tfbs-gtrd | 50 | questions | Exclusive & exhaustive | — |
Protein variant from sequencelab-bench-dbqa-variant-from-sequence | 100 | questions | Exclusive & exhaustive | — |
Protein variant with multiple sequenceslab-bench-dbqa-variant-multi-sequence | 100 | questions | Exclusive & exhaustive | — |
Vaccine response gene setslab-bench-dbqa-vax-response | 50 | questions | Exclusive & exhaustive | — |
Viral protein–protein interactionslab-bench-dbqa-viral-ppi | 50 | questions | Exclusive & exhaustive | — |
Identifies genes associated with a phenotype in DisGeNET but not OMIM.
1 evaluation run(s)
Retrieves human-gene cytogenetic locations from the stated Ensembl release.
1 evaluation run(s)
Retrieves computationally predicted human miRNA targets from miRDB.
1 evaluation run(s)
Retrieves genes in Mammalian Phenotype Tumor Ontology gene sets.
1 evaluation run(s)
Retrieves membership in MSigDB C6 oncogenic-signature gene sets.
1 evaluation run(s)
Retrieves promoter-region transcription-factor binding-site annotations from GTRD.
1 evaluation run(s)
Uses a protein sequence and ClinVar lookup to identify benign or pathogenic variants.
1 evaluation run(s)
Identifies ClinVar variant pathogenicity while reasoning across multiple protein sequences.
1 evaluation run(s)
Retrieves membership in MSigDB vaccine-response gene sets.
1 evaluation run(s)
Retrieves predicted human interaction partners of viral proteins from P-HIPSter.
1 evaluation run(s)
Scientific Task Atlas
complete for repository-998a8e0. Single-purpose formal LAB-Bench track.
| Scientific task | Coverage | Count | Mapping | Evidence |
|---|---|---|---|---|
| Scientific database retrieval | explicitly-in-scope | 650 questions questions across formal child tasks | official-track high confidence | lab-bench-dbqa-evidence-repositoryComplete formal DbQA category. |
| Domain | Coverage | Count | Interpretation |
|---|---|---|---|
| Protein sequence | explicitly-in-scope | 200 | Two ClinVar protein-sequence tasks contain 100 questions each. |
| Protein-protein binding | explicitly-in-scope | 50 | Viral PPI contains 50 database-retrieval questions. |
Evaluation registry
A setting change—scope, prompt, tools, budget, grader, or repeats—creates a separate run. Charts never cross a comparability group.
Evaluation run
From LAB-Bench: Measuring Capabilities of Language Models for Biology Research
Evaluated models / systems: Claude 3.5 Sonnet (claude-3-5-sonnet-20240620), Claude 3 Haiku (claude-3-haiku-20240307), Claude 3 Opus (claude-3-opus-20240229), Gemini 1.5 Pro (gemini-1.5-pro-001), GPT-4 Turbo (LAB-Bench snapshot not reported), GPT-4o (LAB-Bench snapshot not reported), Meta-Llama-3-70B-Instruct (Anyscale API)
| Metric | Kind / baseline | Unit | Aggregation | Threshold / tolerance |
|---|---|---|---|---|
| Accuracy | absolute | proportion | correct / all questions; mean over three runs | exact choice |
| Precision (selective accuracy) | absolute | proportion | correct / attempted questions; mean over three runs | insufficient-information responses excluded from denominator |
| Coverage | absolute | proportion | attempted / all questions; mean over three runs | insufficient-information responses treated as not attempted |
| Model | Metric | Value | n |
|---|---|---|---|
| Claude 3.5 Sonnet (claude-3-5-sonnet-20240620) | Accuracy | 0 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Claude 3.5 Sonnet (claude-3-5-sonnet-20240620) | Precision (selective accuracy) | 0 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Claude 3.5 Sonnet (claude-3-5-sonnet-20240620) | Coverage | 0 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Claude 3 Opus (claude-3-opus-20240229) | Accuracy | 0 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Claude 3 Opus (claude-3-opus-20240229) | Precision (selective accuracy) | 0 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Claude 3 Opus (claude-3-opus-20240229) | Coverage | 0.01 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Gemini 1.5 Pro (gemini-1.5-pro-001) | Accuracy | 0 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Gemini 1.5 Pro (gemini-1.5-pro-001) | Precision (selective accuracy) | 0 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Gemini 1.5 Pro (gemini-1.5-pro-001) | Coverage | 0 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| GPT-4o (LAB-Bench snapshot not reported) | Accuracy | 0.16 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| GPT-4o (LAB-Bench snapshot not reported) | Precision (selective accuracy) | 0.17 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| GPT-4o (LAB-Bench snapshot not reported) | Coverage | 0.95 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| GPT-4 Turbo (LAB-Bench snapshot not reported) | Accuracy | 0.05 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| GPT-4 Turbo (LAB-Bench snapshot not reported) | Precision (selective accuracy) | 0.17 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| GPT-4 Turbo (LAB-Bench snapshot not reported) | Coverage | 0.32 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Claude 3 Haiku (claude-3-haiku-20240307) | Accuracy | 0.15 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Claude 3 Haiku (claude-3-haiku-20240307) | Precision (selective accuracy) | 0.17 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Claude 3 Haiku (claude-3-haiku-20240307) | Coverage | 0.9 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Meta-Llama-3-70B-Instruct (Anyscale API) | Accuracy | 0.25 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Meta-Llama-3-70B-Instruct (Anyscale API) | Precision (selective accuracy) | 0.25 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Meta-Llama-3-70B-Instruct (Anyscale API) | Coverage | 0.99 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
Evaluation run
From LAB-Bench: Measuring Capabilities of Language Models for Biology Research
Evaluated models / systems: Claude 3.5 Sonnet (claude-3-5-sonnet-20240620), Claude 3 Haiku (claude-3-haiku-20240307), Claude 3 Opus (claude-3-opus-20240229), Gemini 1.5 Pro (gemini-1.5-pro-001), GPT-4 Turbo (LAB-Bench snapshot not reported), GPT-4o (LAB-Bench snapshot not reported), Meta-Llama-3-70B-Instruct (Anyscale API)
| Metric | Kind / baseline | Unit | Aggregation | Threshold / tolerance |
|---|---|---|---|---|
| Accuracy | absolute | proportion | correct / all questions; mean over three runs | exact choice |
| Precision (selective accuracy) | absolute | proportion | correct / attempted questions; mean over three runs | insufficient-information responses excluded from denominator |
| Coverage | absolute | proportion | attempted / all questions; mean over three runs | insufficient-information responses treated as not attempted |
| Model | Metric | Value | n |
|---|---|---|---|
| Claude 3.5 Sonnet (claude-3-5-sonnet-20240620) | Accuracy | 0 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Claude 3.5 Sonnet (claude-3-5-sonnet-20240620) | Precision (selective accuracy) | 0 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Claude 3.5 Sonnet (claude-3-5-sonnet-20240620) | Coverage | 0 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Claude 3 Opus (claude-3-opus-20240229) | Accuracy | 0 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Claude 3 Opus (claude-3-opus-20240229) | Precision (selective accuracy) | 0 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Claude 3 Opus (claude-3-opus-20240229) | Coverage | 0.01 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Gemini 1.5 Pro (gemini-1.5-pro-001) | Accuracy | 0.03 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Gemini 1.5 Pro (gemini-1.5-pro-001) | Precision (selective accuracy) | 0.22 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Gemini 1.5 Pro (gemini-1.5-pro-001) | Coverage | 0.09 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| GPT-4o (LAB-Bench snapshot not reported) | Accuracy | 0.31 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| GPT-4o (LAB-Bench snapshot not reported) | Precision (selective accuracy) | 0.38 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| GPT-4o (LAB-Bench snapshot not reported) | Coverage | 0.81 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| GPT-4 Turbo (LAB-Bench snapshot not reported) | Accuracy | 0.12 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| GPT-4 Turbo (LAB-Bench snapshot not reported) | Precision (selective accuracy) | 0.22 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| GPT-4 Turbo (LAB-Bench snapshot not reported) | Coverage | 0.55 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Claude 3 Haiku (claude-3-haiku-20240307) | Accuracy | 0.18 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Claude 3 Haiku (claude-3-haiku-20240307) | Precision (selective accuracy) | 0.22 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Claude 3 Haiku (claude-3-haiku-20240307) | Coverage | 0.81 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Meta-Llama-3-70B-Instruct (Anyscale API) | Accuracy | 0.2 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Meta-Llama-3-70B-Instruct (Anyscale API) | Precision (selective accuracy) | 0.21 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Meta-Llama-3-70B-Instruct (Anyscale API) | Coverage | 0.97 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
Evaluation run
From LAB-Bench: Measuring Capabilities of Language Models for Biology Research
Evaluated models / systems: Claude 3.5 Sonnet (claude-3-5-sonnet-20240620), Claude 3 Haiku (claude-3-haiku-20240307), Claude 3 Opus (claude-3-opus-20240229), Gemini 1.5 Pro (gemini-1.5-pro-001), GPT-4 Turbo (LAB-Bench snapshot not reported), GPT-4o (LAB-Bench snapshot not reported), Meta-Llama-3-70B-Instruct (Anyscale API)
| Metric | Kind / baseline | Unit | Aggregation | Threshold / tolerance |
|---|---|---|---|---|
| Accuracy | absolute | proportion | correct / all questions; mean over three runs | exact choice |
| Precision (selective accuracy) | absolute | proportion | correct / attempted questions; mean over three runs | insufficient-information responses excluded from denominator |
| Coverage | absolute | proportion | attempted / all questions; mean over three runs | insufficient-information responses treated as not attempted |
| Model | Metric | Value | n |
|---|---|---|---|
| Claude 3.5 Sonnet (claude-3-5-sonnet-20240620) | Accuracy | 0 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Claude 3.5 Sonnet (claude-3-5-sonnet-20240620) | Precision (selective accuracy) | 0 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Claude 3.5 Sonnet (claude-3-5-sonnet-20240620) | Coverage | 0 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Claude 3 Opus (claude-3-opus-20240229) | Accuracy | 0 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Claude 3 Opus (claude-3-opus-20240229) | Precision (selective accuracy) | 0 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Claude 3 Opus (claude-3-opus-20240229) | Coverage | 0 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Gemini 1.5 Pro (gemini-1.5-pro-001) | Accuracy | 0 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Gemini 1.5 Pro (gemini-1.5-pro-001) | Precision (selective accuracy) | 0 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Gemini 1.5 Pro (gemini-1.5-pro-001) | Coverage | 0 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| GPT-4o (LAB-Bench snapshot not reported) | Accuracy | 0.1 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| GPT-4o (LAB-Bench snapshot not reported) | Precision (selective accuracy) | 0.3 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| GPT-4o (LAB-Bench snapshot not reported) | Coverage | 0.35 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| GPT-4 Turbo (LAB-Bench snapshot not reported) | Accuracy | 0.01 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| GPT-4 Turbo (LAB-Bench snapshot not reported) | Precision (selective accuracy) | 0.03 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| GPT-4 Turbo (LAB-Bench snapshot not reported) | Coverage | 0.15 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Claude 3 Haiku (claude-3-haiku-20240307) | Accuracy | 0.15 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Claude 3 Haiku (claude-3-haiku-20240307) | Precision (selective accuracy) | 0.26 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Claude 3 Haiku (claude-3-haiku-20240307) | Coverage | 0.6 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Meta-Llama-3-70B-Instruct (Anyscale API) | Accuracy | 0.29 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Meta-Llama-3-70B-Instruct (Anyscale API) | Precision (selective accuracy) | 0.29 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Meta-Llama-3-70B-Instruct (Anyscale API) | Coverage | 0.99 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
Evaluation run
From LAB-Bench: Measuring Capabilities of Language Models for Biology Research
Evaluated models / systems: Claude 3.5 Sonnet (claude-3-5-sonnet-20240620), Claude 3 Haiku (claude-3-haiku-20240307), Claude 3 Opus (claude-3-opus-20240229), Gemini 1.5 Pro (gemini-1.5-pro-001), GPT-4 Turbo (LAB-Bench snapshot not reported), GPT-4o (LAB-Bench snapshot not reported), Meta-Llama-3-70B-Instruct (Anyscale API)
| Metric | Kind / baseline | Unit | Aggregation | Threshold / tolerance |
|---|---|---|---|---|
| Accuracy | absolute | proportion | correct / all questions; mean over three runs | exact choice |
| Precision (selective accuracy) | absolute | proportion | correct / attempted questions; mean over three runs | insufficient-information responses excluded from denominator |
| Coverage | absolute | proportion | attempted / all questions; mean over three runs | insufficient-information responses treated as not attempted |
| Model | Metric | Value | n |
|---|---|---|---|
| Claude 3.5 Sonnet (claude-3-5-sonnet-20240620) | Accuracy | 0.55 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 100 |
| Claude 3.5 Sonnet (claude-3-5-sonnet-20240620) | Precision (selective accuracy) | 0.73 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 100 |
| Claude 3.5 Sonnet (claude-3-5-sonnet-20240620) | Coverage | 0.76 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 100 |
| Claude 3 Opus (claude-3-opus-20240229) | Accuracy | 0.36 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 100 |
| Claude 3 Opus (claude-3-opus-20240229) | Precision (selective accuracy) | 0.71 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 100 |
| Claude 3 Opus (claude-3-opus-20240229) | Coverage | 0.51 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 100 |
| Gemini 1.5 Pro (gemini-1.5-pro-001) | Accuracy | 0.1 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 100 |
| Gemini 1.5 Pro (gemini-1.5-pro-001) | Precision (selective accuracy) | 0.76 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 100 |
| Gemini 1.5 Pro (gemini-1.5-pro-001) | Coverage | 0.14 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 100 |
| GPT-4o (LAB-Bench snapshot not reported) | Accuracy | 0.57 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 100 |
| GPT-4o (LAB-Bench snapshot not reported) | Precision (selective accuracy) | 0.59 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 100 |
| GPT-4o (LAB-Bench snapshot not reported) | Coverage | 0.96 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 100 |
| GPT-4 Turbo (LAB-Bench snapshot not reported) | Accuracy | 0.16 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 100 |
| GPT-4 Turbo (LAB-Bench snapshot not reported) | Precision (selective accuracy) | 0.58 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 100 |
| GPT-4 Turbo (LAB-Bench snapshot not reported) | Coverage | 0.28 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 100 |
| Claude 3 Haiku (claude-3-haiku-20240307) | Accuracy | 0.48 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 100 |
| Claude 3 Haiku (claude-3-haiku-20240307) | Precision (selective accuracy) | 0.55 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 100 |
| Claude 3 Haiku (claude-3-haiku-20240307) | Coverage | 0.87 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 100 |
| Meta-Llama-3-70B-Instruct (Anyscale API) | Accuracy | 0.53 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 100 |
| Meta-Llama-3-70B-Instruct (Anyscale API) | Precision (selective accuracy) | 0.58 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 100 |
| Meta-Llama-3-70B-Instruct (Anyscale API) | Coverage | 0.93 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 100 |
Evaluation run
From LAB-Bench: Measuring Capabilities of Language Models for Biology Research
Evaluated models / systems: Claude 3.5 Sonnet (claude-3-5-sonnet-20240620), Claude 3 Haiku (claude-3-haiku-20240307), Claude 3 Opus (claude-3-opus-20240229), Gemini 1.5 Pro (gemini-1.5-pro-001), GPT-4 Turbo (LAB-Bench snapshot not reported), GPT-4o (LAB-Bench snapshot not reported), Meta-Llama-3-70B-Instruct (Anyscale API)
| Metric | Kind / baseline | Unit | Aggregation | Threshold / tolerance |
|---|---|---|---|---|
| Accuracy | absolute | proportion | correct / all questions; mean over three runs | exact choice |
| Precision (selective accuracy) | absolute | proportion | correct / attempted questions; mean over three runs | insufficient-information responses excluded from denominator |
| Coverage | absolute | proportion | attempted / all questions; mean over three runs | insufficient-information responses treated as not attempted |
| Model | Metric | Value | n |
|---|---|---|---|
| Claude 3.5 Sonnet (claude-3-5-sonnet-20240620) | Accuracy | 0.14 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Claude 3.5 Sonnet (claude-3-5-sonnet-20240620) | Precision (selective accuracy) | 0.59 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Claude 3.5 Sonnet (claude-3-5-sonnet-20240620) | Coverage | 0.24 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Claude 3 Opus (claude-3-opus-20240229) | Accuracy | 0.12 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Claude 3 Opus (claude-3-opus-20240229) | Precision (selective accuracy) | 0.57 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Claude 3 Opus (claude-3-opus-20240229) | Coverage | 0.2 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Gemini 1.5 Pro (gemini-1.5-pro-001) | Accuracy | 0.04 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Gemini 1.5 Pro (gemini-1.5-pro-001) | Precision (selective accuracy) | 0.75 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Gemini 1.5 Pro (gemini-1.5-pro-001) | Coverage | 0.05 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| GPT-4o (LAB-Bench snapshot not reported) | Accuracy | 0.25 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| GPT-4o (LAB-Bench snapshot not reported) | Precision (selective accuracy) | 0.32 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| GPT-4o (LAB-Bench snapshot not reported) | Coverage | 0.79 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| GPT-4 Turbo (LAB-Bench snapshot not reported) | Accuracy | 0.06 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| GPT-4 Turbo (LAB-Bench snapshot not reported) | Precision (selective accuracy) | 0.61 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| GPT-4 Turbo (LAB-Bench snapshot not reported) | Coverage | 0.1 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Claude 3 Haiku (claude-3-haiku-20240307) | Accuracy | 0.24 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Claude 3 Haiku (claude-3-haiku-20240307) | Precision (selective accuracy) | 0.3 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Claude 3 Haiku (claude-3-haiku-20240307) | Coverage | 0.8 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Meta-Llama-3-70B-Instruct (Anyscale API) | Accuracy | 0.2 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Meta-Llama-3-70B-Instruct (Anyscale API) | Precision (selective accuracy) | 0.47 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Meta-Llama-3-70B-Instruct (Anyscale API) | Coverage | 0.42 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
Evaluation run
From LAB-Bench: Measuring Capabilities of Language Models for Biology Research
Evaluated models / systems: Claude 3.5 Sonnet (claude-3-5-sonnet-20240620), Claude 3 Haiku (claude-3-haiku-20240307), Claude 3 Opus (claude-3-opus-20240229), Gemini 1.5 Pro (gemini-1.5-pro-001), GPT-4 Turbo (LAB-Bench snapshot not reported), GPT-4o (LAB-Bench snapshot not reported), Meta-Llama-3-70B-Instruct (Anyscale API)
| Metric | Kind / baseline | Unit | Aggregation | Threshold / tolerance |
|---|---|---|---|---|
| Accuracy | absolute | proportion | correct / all questions; mean over three runs | exact choice |
| Precision (selective accuracy) | absolute | proportion | correct / attempted questions; mean over three runs | insufficient-information responses excluded from denominator |
| Coverage | absolute | proportion | attempted / all questions; mean over three runs | insufficient-information responses treated as not attempted |
| Model | Metric | Value | n |
|---|---|---|---|
| Claude 3.5 Sonnet (claude-3-5-sonnet-20240620) | Accuracy | 0 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Claude 3.5 Sonnet (claude-3-5-sonnet-20240620) | Precision (selective accuracy) | 0 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Claude 3.5 Sonnet (claude-3-5-sonnet-20240620) | Coverage | 0 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Claude 3 Opus (claude-3-opus-20240229) | Accuracy | 0 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Claude 3 Opus (claude-3-opus-20240229) | Precision (selective accuracy) | 0 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Claude 3 Opus (claude-3-opus-20240229) | Coverage | 0 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Gemini 1.5 Pro (gemini-1.5-pro-001) | Accuracy | 0 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Gemini 1.5 Pro (gemini-1.5-pro-001) | Precision (selective accuracy) | 0 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Gemini 1.5 Pro (gemini-1.5-pro-001) | Coverage | 0.01 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| GPT-4o (LAB-Bench snapshot not reported) | Accuracy | 0.01 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| GPT-4o (LAB-Bench snapshot not reported) | Precision (selective accuracy) | 0.33 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| GPT-4o (LAB-Bench snapshot not reported) | Coverage | 0.01 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| GPT-4 Turbo (LAB-Bench snapshot not reported) | Accuracy | 0 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| GPT-4 Turbo (LAB-Bench snapshot not reported) | Precision (selective accuracy) | 0 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| GPT-4 Turbo (LAB-Bench snapshot not reported) | Coverage | 0.11 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Claude 3 Haiku (claude-3-haiku-20240307) | Accuracy | 0.11 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Claude 3 Haiku (claude-3-haiku-20240307) | Precision (selective accuracy) | 0.31 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Claude 3 Haiku (claude-3-haiku-20240307) | Coverage | 0.35 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Meta-Llama-3-70B-Instruct (Anyscale API) | Accuracy | 0.19 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Meta-Llama-3-70B-Instruct (Anyscale API) | Precision (selective accuracy) | 0.19 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Meta-Llama-3-70B-Instruct (Anyscale API) | Coverage | 1 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
Evaluation run
From LAB-Bench: Measuring Capabilities of Language Models for Biology Research
Evaluated models / systems: Claude 3.5 Sonnet (claude-3-5-sonnet-20240620), Claude 3 Haiku (claude-3-haiku-20240307), Claude 3 Opus (claude-3-opus-20240229), Gemini 1.5 Pro (gemini-1.5-pro-001), GPT-4 Turbo (LAB-Bench snapshot not reported), GPT-4o (LAB-Bench snapshot not reported), Meta-Llama-3-70B-Instruct (Anyscale API)
| Metric | Kind / baseline | Unit | Aggregation | Threshold / tolerance |
|---|---|---|---|---|
| Accuracy | absolute | proportion | correct / all questions; mean over three runs | exact choice |
| Precision (selective accuracy) | absolute | proportion | correct / attempted questions; mean over three runs | insufficient-information responses excluded from denominator |
| Coverage | absolute | proportion | attempted / all questions; mean over three runs | insufficient-information responses treated as not attempted |
| Model | Metric | Value | n |
|---|---|---|---|
| Claude 3.5 Sonnet (claude-3-5-sonnet-20240620) | Accuracy | 0.02 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 100 |
| Claude 3.5 Sonnet (claude-3-5-sonnet-20240620) | Precision (selective accuracy) | 0.75 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 100 |
| Claude 3.5 Sonnet (claude-3-5-sonnet-20240620) | Coverage | 0.03 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 100 |
| Claude 3 Opus (claude-3-opus-20240229) | Accuracy | 0 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 100 |
| Claude 3 Opus (claude-3-opus-20240229) | Precision (selective accuracy) | 0.11 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 100 |
| Claude 3 Opus (claude-3-opus-20240229) | Coverage | 0.03 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 100 |
| Gemini 1.5 Pro (gemini-1.5-pro-001) | Accuracy | 0 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 100 |
| Gemini 1.5 Pro (gemini-1.5-pro-001) | Precision (selective accuracy) | 0.17 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 100 |
| Gemini 1.5 Pro (gemini-1.5-pro-001) | Coverage | 0.01 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 100 |
| GPT-4o (LAB-Bench snapshot not reported) | Accuracy | 0.42 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 100 |
| GPT-4o (LAB-Bench snapshot not reported) | Precision (selective accuracy) | 0.44 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 100 |
| GPT-4o (LAB-Bench snapshot not reported) | Coverage | 0.95 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 100 |
| GPT-4 Turbo (LAB-Bench snapshot not reported) | Accuracy | 0 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 100 |
| GPT-4 Turbo (LAB-Bench snapshot not reported) | Precision (selective accuracy) | 0 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 100 |
| GPT-4 Turbo (LAB-Bench snapshot not reported) | Coverage | 0.01 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 100 |
| Claude 3 Haiku (claude-3-haiku-20240307) | Accuracy | 0.22 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 100 |
| Claude 3 Haiku (claude-3-haiku-20240307) | Precision (selective accuracy) | 0.36 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 100 |
| Claude 3 Haiku (claude-3-haiku-20240307) | Coverage | 0.62 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 100 |
| Meta-Llama-3-70B-Instruct (Anyscale API) | Accuracy | 0.3 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 100 |
| Meta-Llama-3-70B-Instruct (Anyscale API) | Precision (selective accuracy) | 0.35 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 100 |
| Meta-Llama-3-70B-Instruct (Anyscale API) | Coverage | 0.85 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 100 |
Evaluation run
From LAB-Bench: Measuring Capabilities of Language Models for Biology Research
Evaluated models / systems: Claude 3.5 Sonnet (claude-3-5-sonnet-20240620), Claude 3 Haiku (claude-3-haiku-20240307), Claude 3 Opus (claude-3-opus-20240229), Gemini 1.5 Pro (gemini-1.5-pro-001), GPT-4 Turbo (LAB-Bench snapshot not reported), GPT-4o (LAB-Bench snapshot not reported), Meta-Llama-3-70B-Instruct (Anyscale API)
| Metric | Kind / baseline | Unit | Aggregation | Threshold / tolerance |
|---|---|---|---|---|
| Accuracy | absolute | proportion | correct / all questions; mean over three runs | exact choice |
| Precision (selective accuracy) | absolute | proportion | correct / attempted questions; mean over three runs | insufficient-information responses excluded from denominator |
| Coverage | absolute | proportion | attempted / all questions; mean over three runs | insufficient-information responses treated as not attempted |
| Model | Metric | Value | n |
|---|---|---|---|
| Claude 3.5 Sonnet (claude-3-5-sonnet-20240620) | Accuracy | 0.09 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 100 |
| Claude 3.5 Sonnet (claude-3-5-sonnet-20240620) | Precision (selective accuracy) | 0.16 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 100 |
| Claude 3.5 Sonnet (claude-3-5-sonnet-20240620) | Coverage | 0.56 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 100 |
| Claude 3 Opus (claude-3-opus-20240229) | Accuracy | 0.05 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 100 |
| Claude 3 Opus (claude-3-opus-20240229) | Precision (selective accuracy) | 0.18 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 100 |
| Claude 3 Opus (claude-3-opus-20240229) | Coverage | 0.26 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 100 |
| Gemini 1.5 Pro (gemini-1.5-pro-001) | Accuracy | 0 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 100 |
| Gemini 1.5 Pro (gemini-1.5-pro-001) | Precision (selective accuracy) | 0.33 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 100 |
| Gemini 1.5 Pro (gemini-1.5-pro-001) | Coverage | 0.04 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 100 |
| GPT-4o (LAB-Bench snapshot not reported) | Accuracy | 0.06 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 100 |
| GPT-4o (LAB-Bench snapshot not reported) | Precision (selective accuracy) | 0.12 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 100 |
| GPT-4o (LAB-Bench snapshot not reported) | Coverage | 0.49 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 100 |
| GPT-4 Turbo (LAB-Bench snapshot not reported) | Accuracy | 0 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 100 |
| GPT-4 Turbo (LAB-Bench snapshot not reported) | Precision (selective accuracy) | 0 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 100 |
| GPT-4 Turbo (LAB-Bench snapshot not reported) | Coverage | 0.02 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 100 |
| Claude 3 Haiku (claude-3-haiku-20240307) | Accuracy | 0.14 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 100 |
| Claude 3 Haiku (claude-3-haiku-20240307) | Precision (selective accuracy) | 0.16 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 100 |
| Claude 3 Haiku (claude-3-haiku-20240307) | Coverage | 0.84 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 100 |
| Meta-Llama-3-70B-Instruct (Anyscale API) | Accuracy | 0.07 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 100 |
| Meta-Llama-3-70B-Instruct (Anyscale API) | Precision (selective accuracy) | 0.12 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 100 |
| Meta-Llama-3-70B-Instruct (Anyscale API) | Coverage | 0.63 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 100 |
Evaluation run
From LAB-Bench: Measuring Capabilities of Language Models for Biology Research
Evaluated models / systems: Claude 3.5 Sonnet (claude-3-5-sonnet-20240620), Claude 3 Haiku (claude-3-haiku-20240307), Claude 3 Opus (claude-3-opus-20240229), Gemini 1.5 Pro (gemini-1.5-pro-001), GPT-4 Turbo (LAB-Bench snapshot not reported), GPT-4o (LAB-Bench snapshot not reported), Meta-Llama-3-70B-Instruct (Anyscale API)
| Metric | Kind / baseline | Unit | Aggregation | Threshold / tolerance |
|---|---|---|---|---|
| Accuracy | absolute | proportion | correct / all questions; mean over three runs | exact choice |
| Precision (selective accuracy) | absolute | proportion | correct / attempted questions; mean over three runs | insufficient-information responses excluded from denominator |
| Coverage | absolute | proportion | attempted / all questions; mean over three runs | insufficient-information responses treated as not attempted |
| Model | Metric | Value | n |
|---|---|---|---|
| Claude 3.5 Sonnet (claude-3-5-sonnet-20240620) | Accuracy | 0.21 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Claude 3.5 Sonnet (claude-3-5-sonnet-20240620) | Precision (selective accuracy) | 0.57 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Claude 3.5 Sonnet (claude-3-5-sonnet-20240620) | Coverage | 0.36 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Claude 3 Opus (claude-3-opus-20240229) | Accuracy | 0.14 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Claude 3 Opus (claude-3-opus-20240229) | Precision (selective accuracy) | 0.65 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Claude 3 Opus (claude-3-opus-20240229) | Coverage | 0.21 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Gemini 1.5 Pro (gemini-1.5-pro-001) | Accuracy | 0.13 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Gemini 1.5 Pro (gemini-1.5-pro-001) | Precision (selective accuracy) | 0.61 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Gemini 1.5 Pro (gemini-1.5-pro-001) | Coverage | 0.21 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| GPT-4o (LAB-Bench snapshot not reported) | Accuracy | 0.33 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| GPT-4o (LAB-Bench snapshot not reported) | Precision (selective accuracy) | 0.45 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| GPT-4o (LAB-Bench snapshot not reported) | Coverage | 0.73 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| GPT-4 Turbo (LAB-Bench snapshot not reported) | Accuracy | 0.06 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| GPT-4 Turbo (LAB-Bench snapshot not reported) | Precision (selective accuracy) | 0.65 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| GPT-4 Turbo (LAB-Bench snapshot not reported) | Coverage | 0.09 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Claude 3 Haiku (claude-3-haiku-20240307) | Accuracy | 0.32 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Claude 3 Haiku (claude-3-haiku-20240307) | Precision (selective accuracy) | 0.38 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Claude 3 Haiku (claude-3-haiku-20240307) | Coverage | 0.83 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Meta-Llama-3-70B-Instruct (Anyscale API) | Accuracy | 0.15 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Meta-Llama-3-70B-Instruct (Anyscale API) | Precision (selective accuracy) | 0.54 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Meta-Llama-3-70B-Instruct (Anyscale API) | Coverage | 0.29 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
Evaluation run
From LAB-Bench: Measuring Capabilities of Language Models for Biology Research
Evaluated models / systems: Claude 3.5 Sonnet (claude-3-5-sonnet-20240620), Claude 3 Haiku (claude-3-haiku-20240307), Claude 3 Opus (claude-3-opus-20240229), Gemini 1.5 Pro (gemini-1.5-pro-001), GPT-4 Turbo (LAB-Bench snapshot not reported), GPT-4o (LAB-Bench snapshot not reported), Meta-Llama-3-70B-Instruct (Anyscale API)
| Metric | Kind / baseline | Unit | Aggregation | Threshold / tolerance |
|---|---|---|---|---|
| Accuracy | absolute | proportion | correct / all questions; mean over three runs | exact choice |
| Precision (selective accuracy) | absolute | proportion | correct / attempted questions; mean over three runs | insufficient-information responses excluded from denominator |
| Coverage | absolute | proportion | attempted / all questions; mean over three runs | insufficient-information responses treated as not attempted |
| Model | Metric | Value | n |
|---|---|---|---|
| Claude 3.5 Sonnet (claude-3-5-sonnet-20240620) | Accuracy | 0 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Claude 3.5 Sonnet (claude-3-5-sonnet-20240620) | Precision (selective accuracy) | 0 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Claude 3.5 Sonnet (claude-3-5-sonnet-20240620) | Coverage | 0 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Claude 3 Opus (claude-3-opus-20240229) | Accuracy | 0 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Claude 3 Opus (claude-3-opus-20240229) | Precision (selective accuracy) | 0 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Claude 3 Opus (claude-3-opus-20240229) | Coverage | 0 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Gemini 1.5 Pro (gemini-1.5-pro-001) | Accuracy | 0 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Gemini 1.5 Pro (gemini-1.5-pro-001) | Precision (selective accuracy) | 0 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Gemini 1.5 Pro (gemini-1.5-pro-001) | Coverage | 0 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| GPT-4o (LAB-Bench snapshot not reported) | Accuracy | 0.26 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| GPT-4o (LAB-Bench snapshot not reported) | Precision (selective accuracy) | 0.69 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| GPT-4o (LAB-Bench snapshot not reported) | Coverage | 0.38 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| GPT-4 Turbo (LAB-Bench snapshot not reported) | Accuracy | 0.02 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| GPT-4 Turbo (LAB-Bench snapshot not reported) | Precision (selective accuracy) | 0.28 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| GPT-4 Turbo (LAB-Bench snapshot not reported) | Coverage | 0.06 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Claude 3 Haiku (claude-3-haiku-20240307) | Accuracy | 0.27 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Claude 3 Haiku (claude-3-haiku-20240307) | Precision (selective accuracy) | 0.49 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Claude 3 Haiku (claude-3-haiku-20240307) | Coverage | 0.55 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Meta-Llama-3-70B-Instruct (Anyscale API) | Accuracy | 0.4 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Meta-Llama-3-70B-Instruct (Anyscale API) | Precision (selective accuracy) | 0.4 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
| Meta-Llama-3-70B-Instruct (Anyscale API) | Coverage | 1 proportion Creator-paper full-snapshot result; human-baseline row is not encoded as a model result. | 50 |
lab-bench-dbqa-dga-creator-mcq · lab-bench-dbqa-dga-paper-v3-zero-shot-cot-no-tools
| Model | Value | Comparability group |
|---|---|---|
| Claude 3.5 Sonnet (claude-3-5-sonnet-20240620) | 0 | lab-bench-dbqa-dga-paper-v3-zero-shot-cot-no-tools |
| Claude 3 Opus (claude-3-opus-20240229) | 0 | lab-bench-dbqa-dga-paper-v3-zero-shot-cot-no-tools |
| Gemini 1.5 Pro (gemini-1.5-pro-001) | 0 | lab-bench-dbqa-dga-paper-v3-zero-shot-cot-no-tools |
| GPT-4o (LAB-Bench snapshot not reported) | 0.16 | lab-bench-dbqa-dga-paper-v3-zero-shot-cot-no-tools |
| GPT-4 Turbo (LAB-Bench snapshot not reported) | 0.05 | lab-bench-dbqa-dga-paper-v3-zero-shot-cot-no-tools |
| Claude 3 Haiku (claude-3-haiku-20240307) | 0.15 | lab-bench-dbqa-dga-paper-v3-zero-shot-cot-no-tools |
| Meta-Llama-3-70B-Instruct (Anyscale API) | 0.25 | lab-bench-dbqa-dga-paper-v3-zero-shot-cot-no-tools |
lab-bench-dbqa-dga-creator-mcq · lab-bench-dbqa-dga-paper-v3-zero-shot-cot-no-tools
| Model | Value | Comparability group |
|---|---|---|
| Claude 3.5 Sonnet (claude-3-5-sonnet-20240620) | 0 | lab-bench-dbqa-dga-paper-v3-zero-shot-cot-no-tools |
| Claude 3 Opus (claude-3-opus-20240229) | 0 | lab-bench-dbqa-dga-paper-v3-zero-shot-cot-no-tools |
| Gemini 1.5 Pro (gemini-1.5-pro-001) | 0 | lab-bench-dbqa-dga-paper-v3-zero-shot-cot-no-tools |
| GPT-4o (LAB-Bench snapshot not reported) | 0.17 | lab-bench-dbqa-dga-paper-v3-zero-shot-cot-no-tools |
| GPT-4 Turbo (LAB-Bench snapshot not reported) | 0.17 | lab-bench-dbqa-dga-paper-v3-zero-shot-cot-no-tools |
| Claude 3 Haiku (claude-3-haiku-20240307) | 0.17 | lab-bench-dbqa-dga-paper-v3-zero-shot-cot-no-tools |
| Meta-Llama-3-70B-Instruct (Anyscale API) | 0.25 | lab-bench-dbqa-dga-paper-v3-zero-shot-cot-no-tools |
lab-bench-dbqa-dga-creator-mcq · lab-bench-dbqa-dga-paper-v3-zero-shot-cot-no-tools
| Model | Value | Comparability group |
|---|---|---|
| Claude 3.5 Sonnet (claude-3-5-sonnet-20240620) | 0 | lab-bench-dbqa-dga-paper-v3-zero-shot-cot-no-tools |
| Claude 3 Opus (claude-3-opus-20240229) | 0.01 | lab-bench-dbqa-dga-paper-v3-zero-shot-cot-no-tools |
| Gemini 1.5 Pro (gemini-1.5-pro-001) | 0 | lab-bench-dbqa-dga-paper-v3-zero-shot-cot-no-tools |
| GPT-4o (LAB-Bench snapshot not reported) | 0.95 | lab-bench-dbqa-dga-paper-v3-zero-shot-cot-no-tools |
| GPT-4 Turbo (LAB-Bench snapshot not reported) | 0.32 | lab-bench-dbqa-dga-paper-v3-zero-shot-cot-no-tools |
| Claude 3 Haiku (claude-3-haiku-20240307) | 0.9 | lab-bench-dbqa-dga-paper-v3-zero-shot-cot-no-tools |
| Meta-Llama-3-70B-Instruct (Anyscale API) | 0.99 | lab-bench-dbqa-dga-paper-v3-zero-shot-cot-no-tools |
lab-bench-dbqa-gene-location-creator-mcq · lab-bench-dbqa-gene-location-paper-v3-zero-shot-cot-no-tools
| Model | Value | Comparability group |
|---|---|---|
| Claude 3.5 Sonnet (claude-3-5-sonnet-20240620) | 0 | lab-bench-dbqa-gene-location-paper-v3-zero-shot-cot-no-tools |
| Claude 3 Opus (claude-3-opus-20240229) | 0 | lab-bench-dbqa-gene-location-paper-v3-zero-shot-cot-no-tools |
| Gemini 1.5 Pro (gemini-1.5-pro-001) | 0.03 | lab-bench-dbqa-gene-location-paper-v3-zero-shot-cot-no-tools |
| GPT-4o (LAB-Bench snapshot not reported) | 0.31 | lab-bench-dbqa-gene-location-paper-v3-zero-shot-cot-no-tools |
| GPT-4 Turbo (LAB-Bench snapshot not reported) | 0.12 | lab-bench-dbqa-gene-location-paper-v3-zero-shot-cot-no-tools |
| Claude 3 Haiku (claude-3-haiku-20240307) | 0.18 | lab-bench-dbqa-gene-location-paper-v3-zero-shot-cot-no-tools |
| Meta-Llama-3-70B-Instruct (Anyscale API) | 0.2 | lab-bench-dbqa-gene-location-paper-v3-zero-shot-cot-no-tools |
lab-bench-dbqa-gene-location-creator-mcq · lab-bench-dbqa-gene-location-paper-v3-zero-shot-cot-no-tools
| Model | Value | Comparability group |
|---|---|---|
| Claude 3.5 Sonnet (claude-3-5-sonnet-20240620) | 0 | lab-bench-dbqa-gene-location-paper-v3-zero-shot-cot-no-tools |
| Claude 3 Opus (claude-3-opus-20240229) | 0 | lab-bench-dbqa-gene-location-paper-v3-zero-shot-cot-no-tools |
| Gemini 1.5 Pro (gemini-1.5-pro-001) | 0.22 | lab-bench-dbqa-gene-location-paper-v3-zero-shot-cot-no-tools |
| GPT-4o (LAB-Bench snapshot not reported) | 0.38 | lab-bench-dbqa-gene-location-paper-v3-zero-shot-cot-no-tools |
| GPT-4 Turbo (LAB-Bench snapshot not reported) | 0.22 | lab-bench-dbqa-gene-location-paper-v3-zero-shot-cot-no-tools |
| Claude 3 Haiku (claude-3-haiku-20240307) | 0.22 | lab-bench-dbqa-gene-location-paper-v3-zero-shot-cot-no-tools |
| Meta-Llama-3-70B-Instruct (Anyscale API) | 0.21 | lab-bench-dbqa-gene-location-paper-v3-zero-shot-cot-no-tools |
lab-bench-dbqa-gene-location-creator-mcq · lab-bench-dbqa-gene-location-paper-v3-zero-shot-cot-no-tools
| Model | Value | Comparability group |
|---|---|---|
| Claude 3.5 Sonnet (claude-3-5-sonnet-20240620) | 0 | lab-bench-dbqa-gene-location-paper-v3-zero-shot-cot-no-tools |
| Claude 3 Opus (claude-3-opus-20240229) | 0.01 | lab-bench-dbqa-gene-location-paper-v3-zero-shot-cot-no-tools |
| Gemini 1.5 Pro (gemini-1.5-pro-001) | 0.09 | lab-bench-dbqa-gene-location-paper-v3-zero-shot-cot-no-tools |
| GPT-4o (LAB-Bench snapshot not reported) | 0.81 | lab-bench-dbqa-gene-location-paper-v3-zero-shot-cot-no-tools |
| GPT-4 Turbo (LAB-Bench snapshot not reported) | 0.55 | lab-bench-dbqa-gene-location-paper-v3-zero-shot-cot-no-tools |
| Claude 3 Haiku (claude-3-haiku-20240307) | 0.81 | lab-bench-dbqa-gene-location-paper-v3-zero-shot-cot-no-tools |
| Meta-Llama-3-70B-Instruct (Anyscale API) | 0.97 | lab-bench-dbqa-gene-location-paper-v3-zero-shot-cot-no-tools |
lab-bench-dbqa-mirna-targets-creator-mcq · lab-bench-dbqa-mirna-targets-paper-v3-zero-shot-cot-no-tools
| Model | Value | Comparability group |
|---|---|---|
| Claude 3.5 Sonnet (claude-3-5-sonnet-20240620) | 0 | lab-bench-dbqa-mirna-targets-paper-v3-zero-shot-cot-no-tools |
| Claude 3 Opus (claude-3-opus-20240229) | 0 | lab-bench-dbqa-mirna-targets-paper-v3-zero-shot-cot-no-tools |
| Gemini 1.5 Pro (gemini-1.5-pro-001) | 0 | lab-bench-dbqa-mirna-targets-paper-v3-zero-shot-cot-no-tools |
| GPT-4o (LAB-Bench snapshot not reported) | 0.1 | lab-bench-dbqa-mirna-targets-paper-v3-zero-shot-cot-no-tools |
| GPT-4 Turbo (LAB-Bench snapshot not reported) | 0.01 | lab-bench-dbqa-mirna-targets-paper-v3-zero-shot-cot-no-tools |
| Claude 3 Haiku (claude-3-haiku-20240307) | 0.15 | lab-bench-dbqa-mirna-targets-paper-v3-zero-shot-cot-no-tools |
| Meta-Llama-3-70B-Instruct (Anyscale API) | 0.29 | lab-bench-dbqa-mirna-targets-paper-v3-zero-shot-cot-no-tools |
lab-bench-dbqa-mirna-targets-creator-mcq · lab-bench-dbqa-mirna-targets-paper-v3-zero-shot-cot-no-tools
| Model | Value | Comparability group |
|---|---|---|
| Claude 3.5 Sonnet (claude-3-5-sonnet-20240620) | 0 | lab-bench-dbqa-mirna-targets-paper-v3-zero-shot-cot-no-tools |
| Claude 3 Opus (claude-3-opus-20240229) | 0 | lab-bench-dbqa-mirna-targets-paper-v3-zero-shot-cot-no-tools |
| Gemini 1.5 Pro (gemini-1.5-pro-001) | 0 | lab-bench-dbqa-mirna-targets-paper-v3-zero-shot-cot-no-tools |
| GPT-4o (LAB-Bench snapshot not reported) | 0.3 | lab-bench-dbqa-mirna-targets-paper-v3-zero-shot-cot-no-tools |
| GPT-4 Turbo (LAB-Bench snapshot not reported) | 0.03 | lab-bench-dbqa-mirna-targets-paper-v3-zero-shot-cot-no-tools |
| Claude 3 Haiku (claude-3-haiku-20240307) | 0.26 | lab-bench-dbqa-mirna-targets-paper-v3-zero-shot-cot-no-tools |
| Meta-Llama-3-70B-Instruct (Anyscale API) | 0.29 | lab-bench-dbqa-mirna-targets-paper-v3-zero-shot-cot-no-tools |
lab-bench-dbqa-mirna-targets-creator-mcq · lab-bench-dbqa-mirna-targets-paper-v3-zero-shot-cot-no-tools
| Model | Value | Comparability group |
|---|---|---|
| Claude 3.5 Sonnet (claude-3-5-sonnet-20240620) | 0 | lab-bench-dbqa-mirna-targets-paper-v3-zero-shot-cot-no-tools |
| Claude 3 Opus (claude-3-opus-20240229) | 0 | lab-bench-dbqa-mirna-targets-paper-v3-zero-shot-cot-no-tools |
| Gemini 1.5 Pro (gemini-1.5-pro-001) | 0 | lab-bench-dbqa-mirna-targets-paper-v3-zero-shot-cot-no-tools |
| GPT-4o (LAB-Bench snapshot not reported) | 0.35 | lab-bench-dbqa-mirna-targets-paper-v3-zero-shot-cot-no-tools |
| GPT-4 Turbo (LAB-Bench snapshot not reported) | 0.15 | lab-bench-dbqa-mirna-targets-paper-v3-zero-shot-cot-no-tools |
| Claude 3 Haiku (claude-3-haiku-20240307) | 0.6 | lab-bench-dbqa-mirna-targets-paper-v3-zero-shot-cot-no-tools |
| Meta-Llama-3-70B-Instruct (Anyscale API) | 0.99 | lab-bench-dbqa-mirna-targets-paper-v3-zero-shot-cot-no-tools |
lab-bench-dbqa-mouse-tumor-gene-sets-creator-mcq · lab-bench-dbqa-mouse-tumor-gene-sets-paper-v3-zero-shot-cot-no-tools
| Model | Value | Comparability group |
|---|---|---|
| Claude 3.5 Sonnet (claude-3-5-sonnet-20240620) | 0.55 | lab-bench-dbqa-mouse-tumor-gene-sets-paper-v3-zero-shot-cot-no-tools |
| Claude 3 Opus (claude-3-opus-20240229) | 0.36 | lab-bench-dbqa-mouse-tumor-gene-sets-paper-v3-zero-shot-cot-no-tools |
| Gemini 1.5 Pro (gemini-1.5-pro-001) | 0.1 | lab-bench-dbqa-mouse-tumor-gene-sets-paper-v3-zero-shot-cot-no-tools |
| GPT-4o (LAB-Bench snapshot not reported) | 0.57 | lab-bench-dbqa-mouse-tumor-gene-sets-paper-v3-zero-shot-cot-no-tools |
| GPT-4 Turbo (LAB-Bench snapshot not reported) | 0.16 | lab-bench-dbqa-mouse-tumor-gene-sets-paper-v3-zero-shot-cot-no-tools |
| Claude 3 Haiku (claude-3-haiku-20240307) | 0.48 | lab-bench-dbqa-mouse-tumor-gene-sets-paper-v3-zero-shot-cot-no-tools |
| Meta-Llama-3-70B-Instruct (Anyscale API) | 0.53 | lab-bench-dbqa-mouse-tumor-gene-sets-paper-v3-zero-shot-cot-no-tools |
lab-bench-dbqa-mouse-tumor-gene-sets-creator-mcq · lab-bench-dbqa-mouse-tumor-gene-sets-paper-v3-zero-shot-cot-no-tools
| Model | Value | Comparability group |
|---|---|---|
| Claude 3.5 Sonnet (claude-3-5-sonnet-20240620) | 0.73 | lab-bench-dbqa-mouse-tumor-gene-sets-paper-v3-zero-shot-cot-no-tools |
| Claude 3 Opus (claude-3-opus-20240229) | 0.71 | lab-bench-dbqa-mouse-tumor-gene-sets-paper-v3-zero-shot-cot-no-tools |
| Gemini 1.5 Pro (gemini-1.5-pro-001) | 0.76 | lab-bench-dbqa-mouse-tumor-gene-sets-paper-v3-zero-shot-cot-no-tools |
| GPT-4o (LAB-Bench snapshot not reported) | 0.59 | lab-bench-dbqa-mouse-tumor-gene-sets-paper-v3-zero-shot-cot-no-tools |
| GPT-4 Turbo (LAB-Bench snapshot not reported) | 0.58 | lab-bench-dbqa-mouse-tumor-gene-sets-paper-v3-zero-shot-cot-no-tools |
| Claude 3 Haiku (claude-3-haiku-20240307) | 0.55 | lab-bench-dbqa-mouse-tumor-gene-sets-paper-v3-zero-shot-cot-no-tools |
| Meta-Llama-3-70B-Instruct (Anyscale API) | 0.58 | lab-bench-dbqa-mouse-tumor-gene-sets-paper-v3-zero-shot-cot-no-tools |
lab-bench-dbqa-mouse-tumor-gene-sets-creator-mcq · lab-bench-dbqa-mouse-tumor-gene-sets-paper-v3-zero-shot-cot-no-tools
| Model | Value | Comparability group |
|---|---|---|
| Claude 3.5 Sonnet (claude-3-5-sonnet-20240620) | 0.76 | lab-bench-dbqa-mouse-tumor-gene-sets-paper-v3-zero-shot-cot-no-tools |
| Claude 3 Opus (claude-3-opus-20240229) | 0.51 | lab-bench-dbqa-mouse-tumor-gene-sets-paper-v3-zero-shot-cot-no-tools |
| Gemini 1.5 Pro (gemini-1.5-pro-001) | 0.14 | lab-bench-dbqa-mouse-tumor-gene-sets-paper-v3-zero-shot-cot-no-tools |
| GPT-4o (LAB-Bench snapshot not reported) | 0.96 | lab-bench-dbqa-mouse-tumor-gene-sets-paper-v3-zero-shot-cot-no-tools |
| GPT-4 Turbo (LAB-Bench snapshot not reported) | 0.28 | lab-bench-dbqa-mouse-tumor-gene-sets-paper-v3-zero-shot-cot-no-tools |
| Claude 3 Haiku (claude-3-haiku-20240307) | 0.87 | lab-bench-dbqa-mouse-tumor-gene-sets-paper-v3-zero-shot-cot-no-tools |
| Meta-Llama-3-70B-Instruct (Anyscale API) | 0.93 | lab-bench-dbqa-mouse-tumor-gene-sets-paper-v3-zero-shot-cot-no-tools |
lab-bench-dbqa-oncogenic-signatures-creator-mcq · lab-bench-dbqa-oncogenic-signatures-paper-v3-zero-shot-cot-no-tools
| Model | Value | Comparability group |
|---|---|---|
| Claude 3.5 Sonnet (claude-3-5-sonnet-20240620) | 0.14 | lab-bench-dbqa-oncogenic-signatures-paper-v3-zero-shot-cot-no-tools |
| Claude 3 Opus (claude-3-opus-20240229) | 0.12 | lab-bench-dbqa-oncogenic-signatures-paper-v3-zero-shot-cot-no-tools |
| Gemini 1.5 Pro (gemini-1.5-pro-001) | 0.04 | lab-bench-dbqa-oncogenic-signatures-paper-v3-zero-shot-cot-no-tools |
| GPT-4o (LAB-Bench snapshot not reported) | 0.25 | lab-bench-dbqa-oncogenic-signatures-paper-v3-zero-shot-cot-no-tools |
| GPT-4 Turbo (LAB-Bench snapshot not reported) | 0.06 | lab-bench-dbqa-oncogenic-signatures-paper-v3-zero-shot-cot-no-tools |
| Claude 3 Haiku (claude-3-haiku-20240307) | 0.24 | lab-bench-dbqa-oncogenic-signatures-paper-v3-zero-shot-cot-no-tools |
| Meta-Llama-3-70B-Instruct (Anyscale API) | 0.2 | lab-bench-dbqa-oncogenic-signatures-paper-v3-zero-shot-cot-no-tools |
lab-bench-dbqa-oncogenic-signatures-creator-mcq · lab-bench-dbqa-oncogenic-signatures-paper-v3-zero-shot-cot-no-tools
| Model | Value | Comparability group |
|---|---|---|
| Claude 3.5 Sonnet (claude-3-5-sonnet-20240620) | 0.59 | lab-bench-dbqa-oncogenic-signatures-paper-v3-zero-shot-cot-no-tools |
| Claude 3 Opus (claude-3-opus-20240229) | 0.57 | lab-bench-dbqa-oncogenic-signatures-paper-v3-zero-shot-cot-no-tools |
| Gemini 1.5 Pro (gemini-1.5-pro-001) | 0.75 | lab-bench-dbqa-oncogenic-signatures-paper-v3-zero-shot-cot-no-tools |
| GPT-4o (LAB-Bench snapshot not reported) | 0.32 | lab-bench-dbqa-oncogenic-signatures-paper-v3-zero-shot-cot-no-tools |
| GPT-4 Turbo (LAB-Bench snapshot not reported) | 0.61 | lab-bench-dbqa-oncogenic-signatures-paper-v3-zero-shot-cot-no-tools |
| Claude 3 Haiku (claude-3-haiku-20240307) | 0.3 | lab-bench-dbqa-oncogenic-signatures-paper-v3-zero-shot-cot-no-tools |
| Meta-Llama-3-70B-Instruct (Anyscale API) | 0.47 | lab-bench-dbqa-oncogenic-signatures-paper-v3-zero-shot-cot-no-tools |
lab-bench-dbqa-oncogenic-signatures-creator-mcq · lab-bench-dbqa-oncogenic-signatures-paper-v3-zero-shot-cot-no-tools
| Model | Value | Comparability group |
|---|---|---|
| Claude 3.5 Sonnet (claude-3-5-sonnet-20240620) | 0.24 | lab-bench-dbqa-oncogenic-signatures-paper-v3-zero-shot-cot-no-tools |
| Claude 3 Opus (claude-3-opus-20240229) | 0.2 | lab-bench-dbqa-oncogenic-signatures-paper-v3-zero-shot-cot-no-tools |
| Gemini 1.5 Pro (gemini-1.5-pro-001) | 0.05 | lab-bench-dbqa-oncogenic-signatures-paper-v3-zero-shot-cot-no-tools |
| GPT-4o (LAB-Bench snapshot not reported) | 0.79 | lab-bench-dbqa-oncogenic-signatures-paper-v3-zero-shot-cot-no-tools |
| GPT-4 Turbo (LAB-Bench snapshot not reported) | 0.1 | lab-bench-dbqa-oncogenic-signatures-paper-v3-zero-shot-cot-no-tools |
| Claude 3 Haiku (claude-3-haiku-20240307) | 0.8 | lab-bench-dbqa-oncogenic-signatures-paper-v3-zero-shot-cot-no-tools |
| Meta-Llama-3-70B-Instruct (Anyscale API) | 0.42 | lab-bench-dbqa-oncogenic-signatures-paper-v3-zero-shot-cot-no-tools |
lab-bench-dbqa-tfbs-gtrd-creator-mcq · lab-bench-dbqa-tfbs-gtrd-paper-v3-zero-shot-cot-no-tools
| Model | Value | Comparability group |
|---|---|---|
| Claude 3.5 Sonnet (claude-3-5-sonnet-20240620) | 0 | lab-bench-dbqa-tfbs-gtrd-paper-v3-zero-shot-cot-no-tools |
| Claude 3 Opus (claude-3-opus-20240229) | 0 | lab-bench-dbqa-tfbs-gtrd-paper-v3-zero-shot-cot-no-tools |
| Gemini 1.5 Pro (gemini-1.5-pro-001) | 0 | lab-bench-dbqa-tfbs-gtrd-paper-v3-zero-shot-cot-no-tools |
| GPT-4o (LAB-Bench snapshot not reported) | 0.01 | lab-bench-dbqa-tfbs-gtrd-paper-v3-zero-shot-cot-no-tools |
| GPT-4 Turbo (LAB-Bench snapshot not reported) | 0 | lab-bench-dbqa-tfbs-gtrd-paper-v3-zero-shot-cot-no-tools |
| Claude 3 Haiku (claude-3-haiku-20240307) | 0.11 | lab-bench-dbqa-tfbs-gtrd-paper-v3-zero-shot-cot-no-tools |
| Meta-Llama-3-70B-Instruct (Anyscale API) | 0.19 | lab-bench-dbqa-tfbs-gtrd-paper-v3-zero-shot-cot-no-tools |
lab-bench-dbqa-tfbs-gtrd-creator-mcq · lab-bench-dbqa-tfbs-gtrd-paper-v3-zero-shot-cot-no-tools
| Model | Value | Comparability group |
|---|---|---|
| Claude 3.5 Sonnet (claude-3-5-sonnet-20240620) | 0 | lab-bench-dbqa-tfbs-gtrd-paper-v3-zero-shot-cot-no-tools |
| Claude 3 Opus (claude-3-opus-20240229) | 0 | lab-bench-dbqa-tfbs-gtrd-paper-v3-zero-shot-cot-no-tools |
| Gemini 1.5 Pro (gemini-1.5-pro-001) | 0 | lab-bench-dbqa-tfbs-gtrd-paper-v3-zero-shot-cot-no-tools |
| GPT-4o (LAB-Bench snapshot not reported) | 0.33 | lab-bench-dbqa-tfbs-gtrd-paper-v3-zero-shot-cot-no-tools |
| GPT-4 Turbo (LAB-Bench snapshot not reported) | 0 | lab-bench-dbqa-tfbs-gtrd-paper-v3-zero-shot-cot-no-tools |
| Claude 3 Haiku (claude-3-haiku-20240307) | 0.31 | lab-bench-dbqa-tfbs-gtrd-paper-v3-zero-shot-cot-no-tools |
| Meta-Llama-3-70B-Instruct (Anyscale API) | 0.19 | lab-bench-dbqa-tfbs-gtrd-paper-v3-zero-shot-cot-no-tools |
lab-bench-dbqa-tfbs-gtrd-creator-mcq · lab-bench-dbqa-tfbs-gtrd-paper-v3-zero-shot-cot-no-tools
| Model | Value | Comparability group |
|---|---|---|
| Claude 3.5 Sonnet (claude-3-5-sonnet-20240620) | 0 | lab-bench-dbqa-tfbs-gtrd-paper-v3-zero-shot-cot-no-tools |
| Claude 3 Opus (claude-3-opus-20240229) | 0 | lab-bench-dbqa-tfbs-gtrd-paper-v3-zero-shot-cot-no-tools |
| Gemini 1.5 Pro (gemini-1.5-pro-001) | 0.01 | lab-bench-dbqa-tfbs-gtrd-paper-v3-zero-shot-cot-no-tools |
| GPT-4o (LAB-Bench snapshot not reported) | 0.01 | lab-bench-dbqa-tfbs-gtrd-paper-v3-zero-shot-cot-no-tools |
| GPT-4 Turbo (LAB-Bench snapshot not reported) | 0.11 | lab-bench-dbqa-tfbs-gtrd-paper-v3-zero-shot-cot-no-tools |
| Claude 3 Haiku (claude-3-haiku-20240307) | 0.35 | lab-bench-dbqa-tfbs-gtrd-paper-v3-zero-shot-cot-no-tools |
| Meta-Llama-3-70B-Instruct (Anyscale API) | 1 | lab-bench-dbqa-tfbs-gtrd-paper-v3-zero-shot-cot-no-tools |
lab-bench-dbqa-variant-from-sequence-creator-mcq · lab-bench-dbqa-variant-from-sequence-paper-v3-zero-shot-cot-no-tools
| Model | Value | Comparability group |
|---|---|---|
| Claude 3.5 Sonnet (claude-3-5-sonnet-20240620) | 0.02 | lab-bench-dbqa-variant-from-sequence-paper-v3-zero-shot-cot-no-tools |
| Claude 3 Opus (claude-3-opus-20240229) | 0 | lab-bench-dbqa-variant-from-sequence-paper-v3-zero-shot-cot-no-tools |
| Gemini 1.5 Pro (gemini-1.5-pro-001) | 0 | lab-bench-dbqa-variant-from-sequence-paper-v3-zero-shot-cot-no-tools |
| GPT-4o (LAB-Bench snapshot not reported) | 0.42 | lab-bench-dbqa-variant-from-sequence-paper-v3-zero-shot-cot-no-tools |
| GPT-4 Turbo (LAB-Bench snapshot not reported) | 0 | lab-bench-dbqa-variant-from-sequence-paper-v3-zero-shot-cot-no-tools |
| Claude 3 Haiku (claude-3-haiku-20240307) | 0.22 | lab-bench-dbqa-variant-from-sequence-paper-v3-zero-shot-cot-no-tools |
| Meta-Llama-3-70B-Instruct (Anyscale API) | 0.3 | lab-bench-dbqa-variant-from-sequence-paper-v3-zero-shot-cot-no-tools |
lab-bench-dbqa-variant-from-sequence-creator-mcq · lab-bench-dbqa-variant-from-sequence-paper-v3-zero-shot-cot-no-tools
| Model | Value | Comparability group |
|---|---|---|
| Claude 3.5 Sonnet (claude-3-5-sonnet-20240620) | 0.75 | lab-bench-dbqa-variant-from-sequence-paper-v3-zero-shot-cot-no-tools |
| Claude 3 Opus (claude-3-opus-20240229) | 0.11 | lab-bench-dbqa-variant-from-sequence-paper-v3-zero-shot-cot-no-tools |
| Gemini 1.5 Pro (gemini-1.5-pro-001) | 0.17 | lab-bench-dbqa-variant-from-sequence-paper-v3-zero-shot-cot-no-tools |
| GPT-4o (LAB-Bench snapshot not reported) | 0.44 | lab-bench-dbqa-variant-from-sequence-paper-v3-zero-shot-cot-no-tools |
| GPT-4 Turbo (LAB-Bench snapshot not reported) | 0 | lab-bench-dbqa-variant-from-sequence-paper-v3-zero-shot-cot-no-tools |
| Claude 3 Haiku (claude-3-haiku-20240307) | 0.36 | lab-bench-dbqa-variant-from-sequence-paper-v3-zero-shot-cot-no-tools |
| Meta-Llama-3-70B-Instruct (Anyscale API) | 0.35 | lab-bench-dbqa-variant-from-sequence-paper-v3-zero-shot-cot-no-tools |
lab-bench-dbqa-variant-from-sequence-creator-mcq · lab-bench-dbqa-variant-from-sequence-paper-v3-zero-shot-cot-no-tools
| Model | Value | Comparability group |
|---|---|---|
| Claude 3.5 Sonnet (claude-3-5-sonnet-20240620) | 0.03 | lab-bench-dbqa-variant-from-sequence-paper-v3-zero-shot-cot-no-tools |
| Claude 3 Opus (claude-3-opus-20240229) | 0.03 | lab-bench-dbqa-variant-from-sequence-paper-v3-zero-shot-cot-no-tools |
| Gemini 1.5 Pro (gemini-1.5-pro-001) | 0.01 | lab-bench-dbqa-variant-from-sequence-paper-v3-zero-shot-cot-no-tools |
| GPT-4o (LAB-Bench snapshot not reported) | 0.95 | lab-bench-dbqa-variant-from-sequence-paper-v3-zero-shot-cot-no-tools |
| GPT-4 Turbo (LAB-Bench snapshot not reported) | 0.01 | lab-bench-dbqa-variant-from-sequence-paper-v3-zero-shot-cot-no-tools |
| Claude 3 Haiku (claude-3-haiku-20240307) | 0.62 | lab-bench-dbqa-variant-from-sequence-paper-v3-zero-shot-cot-no-tools |
| Meta-Llama-3-70B-Instruct (Anyscale API) | 0.85 | lab-bench-dbqa-variant-from-sequence-paper-v3-zero-shot-cot-no-tools |
lab-bench-dbqa-variant-multi-sequence-creator-mcq · lab-bench-dbqa-variant-multi-sequence-paper-v3-zero-shot-cot-no-tools
| Model | Value | Comparability group |
|---|---|---|
| Claude 3.5 Sonnet (claude-3-5-sonnet-20240620) | 0.09 | lab-bench-dbqa-variant-multi-sequence-paper-v3-zero-shot-cot-no-tools |
| Claude 3 Opus (claude-3-opus-20240229) | 0.05 | lab-bench-dbqa-variant-multi-sequence-paper-v3-zero-shot-cot-no-tools |
| Gemini 1.5 Pro (gemini-1.5-pro-001) | 0 | lab-bench-dbqa-variant-multi-sequence-paper-v3-zero-shot-cot-no-tools |
| GPT-4o (LAB-Bench snapshot not reported) | 0.06 | lab-bench-dbqa-variant-multi-sequence-paper-v3-zero-shot-cot-no-tools |
| GPT-4 Turbo (LAB-Bench snapshot not reported) | 0 | lab-bench-dbqa-variant-multi-sequence-paper-v3-zero-shot-cot-no-tools |
| Claude 3 Haiku (claude-3-haiku-20240307) | 0.14 | lab-bench-dbqa-variant-multi-sequence-paper-v3-zero-shot-cot-no-tools |
| Meta-Llama-3-70B-Instruct (Anyscale API) | 0.07 | lab-bench-dbqa-variant-multi-sequence-paper-v3-zero-shot-cot-no-tools |
lab-bench-dbqa-variant-multi-sequence-creator-mcq · lab-bench-dbqa-variant-multi-sequence-paper-v3-zero-shot-cot-no-tools
| Model | Value | Comparability group |
|---|---|---|
| Claude 3.5 Sonnet (claude-3-5-sonnet-20240620) | 0.16 | lab-bench-dbqa-variant-multi-sequence-paper-v3-zero-shot-cot-no-tools |
| Claude 3 Opus (claude-3-opus-20240229) | 0.18 | lab-bench-dbqa-variant-multi-sequence-paper-v3-zero-shot-cot-no-tools |
| Gemini 1.5 Pro (gemini-1.5-pro-001) | 0.33 | lab-bench-dbqa-variant-multi-sequence-paper-v3-zero-shot-cot-no-tools |
| GPT-4o (LAB-Bench snapshot not reported) | 0.12 | lab-bench-dbqa-variant-multi-sequence-paper-v3-zero-shot-cot-no-tools |
| GPT-4 Turbo (LAB-Bench snapshot not reported) | 0 | lab-bench-dbqa-variant-multi-sequence-paper-v3-zero-shot-cot-no-tools |
| Claude 3 Haiku (claude-3-haiku-20240307) | 0.16 | lab-bench-dbqa-variant-multi-sequence-paper-v3-zero-shot-cot-no-tools |
| Meta-Llama-3-70B-Instruct (Anyscale API) | 0.12 | lab-bench-dbqa-variant-multi-sequence-paper-v3-zero-shot-cot-no-tools |
lab-bench-dbqa-variant-multi-sequence-creator-mcq · lab-bench-dbqa-variant-multi-sequence-paper-v3-zero-shot-cot-no-tools
| Model | Value | Comparability group |
|---|---|---|
| Claude 3.5 Sonnet (claude-3-5-sonnet-20240620) | 0.56 | lab-bench-dbqa-variant-multi-sequence-paper-v3-zero-shot-cot-no-tools |
| Claude 3 Opus (claude-3-opus-20240229) | 0.26 | lab-bench-dbqa-variant-multi-sequence-paper-v3-zero-shot-cot-no-tools |
| Gemini 1.5 Pro (gemini-1.5-pro-001) | 0.04 | lab-bench-dbqa-variant-multi-sequence-paper-v3-zero-shot-cot-no-tools |
| GPT-4o (LAB-Bench snapshot not reported) | 0.49 | lab-bench-dbqa-variant-multi-sequence-paper-v3-zero-shot-cot-no-tools |
| GPT-4 Turbo (LAB-Bench snapshot not reported) | 0.02 | lab-bench-dbqa-variant-multi-sequence-paper-v3-zero-shot-cot-no-tools |
| Claude 3 Haiku (claude-3-haiku-20240307) | 0.84 | lab-bench-dbqa-variant-multi-sequence-paper-v3-zero-shot-cot-no-tools |
| Meta-Llama-3-70B-Instruct (Anyscale API) | 0.63 | lab-bench-dbqa-variant-multi-sequence-paper-v3-zero-shot-cot-no-tools |
lab-bench-dbqa-vax-response-creator-mcq · lab-bench-dbqa-vax-response-paper-v3-zero-shot-cot-no-tools
| Model | Value | Comparability group |
|---|---|---|
| Claude 3.5 Sonnet (claude-3-5-sonnet-20240620) | 0.21 | lab-bench-dbqa-vax-response-paper-v3-zero-shot-cot-no-tools |
| Claude 3 Opus (claude-3-opus-20240229) | 0.14 | lab-bench-dbqa-vax-response-paper-v3-zero-shot-cot-no-tools |
| Gemini 1.5 Pro (gemini-1.5-pro-001) | 0.13 | lab-bench-dbqa-vax-response-paper-v3-zero-shot-cot-no-tools |
| GPT-4o (LAB-Bench snapshot not reported) | 0.33 | lab-bench-dbqa-vax-response-paper-v3-zero-shot-cot-no-tools |
| GPT-4 Turbo (LAB-Bench snapshot not reported) | 0.06 | lab-bench-dbqa-vax-response-paper-v3-zero-shot-cot-no-tools |
| Claude 3 Haiku (claude-3-haiku-20240307) | 0.32 | lab-bench-dbqa-vax-response-paper-v3-zero-shot-cot-no-tools |
| Meta-Llama-3-70B-Instruct (Anyscale API) | 0.15 | lab-bench-dbqa-vax-response-paper-v3-zero-shot-cot-no-tools |
lab-bench-dbqa-vax-response-creator-mcq · lab-bench-dbqa-vax-response-paper-v3-zero-shot-cot-no-tools
| Model | Value | Comparability group |
|---|---|---|
| Claude 3.5 Sonnet (claude-3-5-sonnet-20240620) | 0.57 | lab-bench-dbqa-vax-response-paper-v3-zero-shot-cot-no-tools |
| Claude 3 Opus (claude-3-opus-20240229) | 0.65 | lab-bench-dbqa-vax-response-paper-v3-zero-shot-cot-no-tools |
| Gemini 1.5 Pro (gemini-1.5-pro-001) | 0.61 | lab-bench-dbqa-vax-response-paper-v3-zero-shot-cot-no-tools |
| GPT-4o (LAB-Bench snapshot not reported) | 0.45 | lab-bench-dbqa-vax-response-paper-v3-zero-shot-cot-no-tools |
| GPT-4 Turbo (LAB-Bench snapshot not reported) | 0.65 | lab-bench-dbqa-vax-response-paper-v3-zero-shot-cot-no-tools |
| Claude 3 Haiku (claude-3-haiku-20240307) | 0.38 | lab-bench-dbqa-vax-response-paper-v3-zero-shot-cot-no-tools |
| Meta-Llama-3-70B-Instruct (Anyscale API) | 0.54 | lab-bench-dbqa-vax-response-paper-v3-zero-shot-cot-no-tools |
lab-bench-dbqa-vax-response-creator-mcq · lab-bench-dbqa-vax-response-paper-v3-zero-shot-cot-no-tools
| Model | Value | Comparability group |
|---|---|---|
| Claude 3.5 Sonnet (claude-3-5-sonnet-20240620) | 0.36 | lab-bench-dbqa-vax-response-paper-v3-zero-shot-cot-no-tools |
| Claude 3 Opus (claude-3-opus-20240229) | 0.21 | lab-bench-dbqa-vax-response-paper-v3-zero-shot-cot-no-tools |
| Gemini 1.5 Pro (gemini-1.5-pro-001) | 0.21 | lab-bench-dbqa-vax-response-paper-v3-zero-shot-cot-no-tools |
| GPT-4o (LAB-Bench snapshot not reported) | 0.73 | lab-bench-dbqa-vax-response-paper-v3-zero-shot-cot-no-tools |
| GPT-4 Turbo (LAB-Bench snapshot not reported) | 0.09 | lab-bench-dbqa-vax-response-paper-v3-zero-shot-cot-no-tools |
| Claude 3 Haiku (claude-3-haiku-20240307) | 0.83 | lab-bench-dbqa-vax-response-paper-v3-zero-shot-cot-no-tools |
| Meta-Llama-3-70B-Instruct (Anyscale API) | 0.29 | lab-bench-dbqa-vax-response-paper-v3-zero-shot-cot-no-tools |
lab-bench-dbqa-viral-ppi-creator-mcq · lab-bench-dbqa-viral-ppi-paper-v3-zero-shot-cot-no-tools
| Model | Value | Comparability group |
|---|---|---|
| Claude 3.5 Sonnet (claude-3-5-sonnet-20240620) | 0 | lab-bench-dbqa-viral-ppi-paper-v3-zero-shot-cot-no-tools |
| Claude 3 Opus (claude-3-opus-20240229) | 0 | lab-bench-dbqa-viral-ppi-paper-v3-zero-shot-cot-no-tools |
| Gemini 1.5 Pro (gemini-1.5-pro-001) | 0 | lab-bench-dbqa-viral-ppi-paper-v3-zero-shot-cot-no-tools |
| GPT-4o (LAB-Bench snapshot not reported) | 0.26 | lab-bench-dbqa-viral-ppi-paper-v3-zero-shot-cot-no-tools |
| GPT-4 Turbo (LAB-Bench snapshot not reported) | 0.02 | lab-bench-dbqa-viral-ppi-paper-v3-zero-shot-cot-no-tools |
| Claude 3 Haiku (claude-3-haiku-20240307) | 0.27 | lab-bench-dbqa-viral-ppi-paper-v3-zero-shot-cot-no-tools |
| Meta-Llama-3-70B-Instruct (Anyscale API) | 0.4 | lab-bench-dbqa-viral-ppi-paper-v3-zero-shot-cot-no-tools |
lab-bench-dbqa-viral-ppi-creator-mcq · lab-bench-dbqa-viral-ppi-paper-v3-zero-shot-cot-no-tools
| Model | Value | Comparability group |
|---|---|---|
| Claude 3.5 Sonnet (claude-3-5-sonnet-20240620) | 0 | lab-bench-dbqa-viral-ppi-paper-v3-zero-shot-cot-no-tools |
| Claude 3 Opus (claude-3-opus-20240229) | 0 | lab-bench-dbqa-viral-ppi-paper-v3-zero-shot-cot-no-tools |
| Gemini 1.5 Pro (gemini-1.5-pro-001) | 0 | lab-bench-dbqa-viral-ppi-paper-v3-zero-shot-cot-no-tools |
| GPT-4o (LAB-Bench snapshot not reported) | 0.69 | lab-bench-dbqa-viral-ppi-paper-v3-zero-shot-cot-no-tools |
| GPT-4 Turbo (LAB-Bench snapshot not reported) | 0.28 | lab-bench-dbqa-viral-ppi-paper-v3-zero-shot-cot-no-tools |
| Claude 3 Haiku (claude-3-haiku-20240307) | 0.49 | lab-bench-dbqa-viral-ppi-paper-v3-zero-shot-cot-no-tools |
| Meta-Llama-3-70B-Instruct (Anyscale API) | 0.4 | lab-bench-dbqa-viral-ppi-paper-v3-zero-shot-cot-no-tools |
lab-bench-dbqa-viral-ppi-creator-mcq · lab-bench-dbqa-viral-ppi-paper-v3-zero-shot-cot-no-tools
| Model | Value | Comparability group |
|---|---|---|
| Claude 3.5 Sonnet (claude-3-5-sonnet-20240620) | 0 | lab-bench-dbqa-viral-ppi-paper-v3-zero-shot-cot-no-tools |
| Claude 3 Opus (claude-3-opus-20240229) | 0 | lab-bench-dbqa-viral-ppi-paper-v3-zero-shot-cot-no-tools |
| Gemini 1.5 Pro (gemini-1.5-pro-001) | 0 | lab-bench-dbqa-viral-ppi-paper-v3-zero-shot-cot-no-tools |
| GPT-4o (LAB-Bench snapshot not reported) | 0.38 | lab-bench-dbqa-viral-ppi-paper-v3-zero-shot-cot-no-tools |
| GPT-4 Turbo (LAB-Bench snapshot not reported) | 0.06 | lab-bench-dbqa-viral-ppi-paper-v3-zero-shot-cot-no-tools |
| Claude 3 Haiku (claude-3-haiku-20240307) | 0.55 | lab-bench-dbqa-viral-ppi-paper-v3-zero-shot-cot-no-tools |
| Meta-Llama-3-70B-Instruct (Anyscale API) | 1 | lab-bench-dbqa-viral-ppi-paper-v3-zero-shot-cot-no-tools |
Source locators remain visible; expand an item to inspect the exact Registry fields it supports.
/name/organizations/release_date/kind/domains/capabilities/modalities/task_counts/total/task_counts/basis/task_counts/subsets/access/level/access/license/versions/0/task_counts/total/versions/0/task_counts/basis/versions/0/task_counts/subsets/latest_version/task_counts/total/task_counts/basis/task_counts/subsets/access/level/access/tasks/access/artifacts/access/grader/access/license/resources/versions/1/task_counts/total/versions/1/task_counts/basis/versions/1/task_counts/subsets/scientific_task_classification/entries/0