track · audited · verified 2026-07-21

LAB-Bench DbQA

Database-retrieval category spanning 10 genomics, clinical, protein, regulatory, vaccine-response, and viral-PPI tasks.

+3 more

Benchmark definition

What is counted

Version
repository-998a8e0
Total
650 (questions across formal child tasks)
Task formats
multiple choice
Capabilities
RetrievalKnowledgePredictionClassificationScientific reasoning
Modalities
TextProtein sequenceDatabaseWeb

Version history

VersionStatusRelease / as-ofTotalFormal tracks
paper-v3
lab-bench-dbqa-paper-v3
active2024-07-17650 (questions across formal child tasks)lab-bench-dbqa-dga, lab-bench-dbqa-gene-location, lab-bench-dbqa-mirna-targets, lab-bench-dbqa-mouse-tumor-gene-sets, lab-bench-dbqa-oncogenic-signatures, lab-bench-dbqa-tfbs-gtrd, lab-bench-dbqa-variant-from-sequence, lab-bench-dbqa-variant-multi-sequence, lab-bench-dbqa-vax-response, lab-bench-dbqa-viral-ppi
repository-998a8e0
lab-bench-dbqa-repository-998a8e0
current2025-09-27650 (questions across formal child tasks)lab-bench-dbqa-dga, lab-bench-dbqa-gene-location, lab-bench-dbqa-mirna-targets, lab-bench-dbqa-mouse-tumor-gene-sets, lab-bench-dbqa-oncogenic-signatures, lab-bench-dbqa-tfbs-gtrd, lab-bench-dbqa-variant-from-sequence, lab-bench-dbqa-variant-multi-sequence, lab-bench-dbqa-vax-response, lab-bench-dbqa-viral-ppi

Tracks and subsets

IDCountBasisPartition?Notes
Disease gene associations
lab-bench-dbqa-dga
50questionsExclusive & exhaustive
Gene location
lab-bench-dbqa-gene-location
50questionsExclusive & exhaustive
miRNA targets
lab-bench-dbqa-mirna-targets
50questionsExclusive & exhaustive
Mouse tumor gene sets
lab-bench-dbqa-mouse-tumor-gene-sets
100questionsExclusive & exhaustive
Oncogenic signatures
lab-bench-dbqa-oncogenic-signatures
50questionsExclusive & exhaustive
GTRD transcription-factor binding sites
lab-bench-dbqa-tfbs-gtrd
50questionsExclusive & exhaustive
Protein variant from sequence
lab-bench-dbqa-variant-from-sequence
100questionsExclusive & exhaustive
Protein variant with multiple sequences
lab-bench-dbqa-variant-multi-sequence
100questionsExclusive & exhaustive
Vaccine response gene sets
lab-bench-dbqa-vax-response
50questionsExclusive & exhaustive
Viral protein–protein interactions
lab-bench-dbqa-viral-ppi
50questionsExclusive & exhaustive

Registered child tracks

Scientific Task Atlas

Scientific task classification

complete for repository-998a8e0. Single-purpose formal LAB-Bench track.

Scientific taskCoverageCountMappingEvidence
Scientific database retrievalexplicitly-in-scope650 questions
questions across formal child tasks
official-track
high confidence
lab-bench-dbqa-evidence-repository
Complete formal DbQA category.

Scientific coverage notes

DomainCoverageCountInterpretation
Protein sequenceexplicitly-in-scope200Two ClinVar protein-sequence tasks contain 100 questions each.
Protein-protein bindingexplicitly-in-scope50Viral PPI contains 50 database-retrieval questions.

Evaluation registry

Works and run settings

A setting change—scope, prompt, tools, budget, grader, or repeats—creates a separate run. Charts never cross a comparability group.

Evaluation run

lab-bench-dbqa-dga-creator-mcq

From LAB-Bench: Measuring Capabilities of Language Models for Biology Research

lab-bench-dbqa-dga-paper-v3-zero-shot-cot-no-toolsvpaper-v3

Evaluated models / systems: Claude 3.5 Sonnet (claude-3-5-sonnet-20240620), Claude 3 Haiku (claude-3-haiku-20240307), Claude 3 Opus (claude-3-opus-20240229), Gemini 1.5 Pro (gemini-1.5-pro-001), GPT-4 Turbo (LAB-Bench snapshot not reported), GPT-4o (LAB-Bench snapshot not reported), Meta-Llama-3-70B-Instruct (Anyscale API)

Scopefull · n=50
Shots0
Turnssingle-turn
System prompt publicYes
Reasoning / effortzero-shot chain-of-thought
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNo
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
Repeats3
Graderregex multiple-choice parser with Claude 2 fallback · model: Claude 2 fallback parser · human review: no
Statisticsmean across three runs per model
Contaminationfull pre-release snapshot with 80% later public and 20% private contamination-monitoring splits
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracyabsoluteproportioncorrect / all questions; mean over three runsexact choice
Precision (selective accuracy)absoluteproportioncorrect / attempted questions; mean over three runsinsufficient-information responses excluded from denominator
Coverageabsoluteproportionattempted / all questions; mean over three runsinsufficient-information responses treated as not attempted

Results

ModelMetricValuen
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Accuracy0 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Precision (selective accuracy)0 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Coverage0 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Opus (claude-3-opus-20240229)Accuracy0 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Opus (claude-3-opus-20240229)Precision (selective accuracy)0 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Opus (claude-3-opus-20240229)Coverage0.01 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Gemini 1.5 Pro (gemini-1.5-pro-001)Accuracy0 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Gemini 1.5 Pro (gemini-1.5-pro-001)Precision (selective accuracy)0 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Gemini 1.5 Pro (gemini-1.5-pro-001)Coverage0 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4o (LAB-Bench snapshot not reported)Accuracy0.16 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4o (LAB-Bench snapshot not reported)Precision (selective accuracy)0.17 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4o (LAB-Bench snapshot not reported)Coverage0.95 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4 Turbo (LAB-Bench snapshot not reported)Accuracy0.05 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4 Turbo (LAB-Bench snapshot not reported)Precision (selective accuracy)0.17 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4 Turbo (LAB-Bench snapshot not reported)Coverage0.32 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Haiku (claude-3-haiku-20240307)Accuracy0.15 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Haiku (claude-3-haiku-20240307)Precision (selective accuracy)0.17 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Haiku (claude-3-haiku-20240307)Coverage0.9 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Meta-Llama-3-70B-Instruct (Anyscale API)Accuracy0.25 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Meta-Llama-3-70B-Instruct (Anyscale API)Precision (selective accuracy)0.25 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Meta-Llama-3-70B-Instruct (Anyscale API)Coverage0.99 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50

Evidence

  • section: Appendix D.1–D.4; Table 1 and Appendix Table 6 (Full scope, zero-shot chain-of-thought prompt, no tools, three runs, exact metrics, parser, and model identities.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics
  • table: Tables 2–4, row DbQA_dga_task-v1 (Accuracy, precision, and coverage. Human row excluded.) — supports /results

Evaluation run

lab-bench-dbqa-gene-location-creator-mcq

From LAB-Bench: Measuring Capabilities of Language Models for Biology Research

lab-bench-dbqa-gene-location-paper-v3-zero-shot-cot-no-toolsvpaper-v3

Evaluated models / systems: Claude 3.5 Sonnet (claude-3-5-sonnet-20240620), Claude 3 Haiku (claude-3-haiku-20240307), Claude 3 Opus (claude-3-opus-20240229), Gemini 1.5 Pro (gemini-1.5-pro-001), GPT-4 Turbo (LAB-Bench snapshot not reported), GPT-4o (LAB-Bench snapshot not reported), Meta-Llama-3-70B-Instruct (Anyscale API)

Scopefull · n=50
Shots0
Turnssingle-turn
System prompt publicYes
Reasoning / effortzero-shot chain-of-thought
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNo
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
Repeats3
Graderregex multiple-choice parser with Claude 2 fallback · model: Claude 2 fallback parser · human review: no
Statisticsmean across three runs per model
Contaminationfull pre-release snapshot with 80% later public and 20% private contamination-monitoring splits
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracyabsoluteproportioncorrect / all questions; mean over three runsexact choice
Precision (selective accuracy)absoluteproportioncorrect / attempted questions; mean over three runsinsufficient-information responses excluded from denominator
Coverageabsoluteproportionattempted / all questions; mean over three runsinsufficient-information responses treated as not attempted

Results

ModelMetricValuen
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Accuracy0 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Precision (selective accuracy)0 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Coverage0 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Opus (claude-3-opus-20240229)Accuracy0 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Opus (claude-3-opus-20240229)Precision (selective accuracy)0 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Opus (claude-3-opus-20240229)Coverage0.01 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Gemini 1.5 Pro (gemini-1.5-pro-001)Accuracy0.03 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Gemini 1.5 Pro (gemini-1.5-pro-001)Precision (selective accuracy)0.22 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Gemini 1.5 Pro (gemini-1.5-pro-001)Coverage0.09 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4o (LAB-Bench snapshot not reported)Accuracy0.31 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4o (LAB-Bench snapshot not reported)Precision (selective accuracy)0.38 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4o (LAB-Bench snapshot not reported)Coverage0.81 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4 Turbo (LAB-Bench snapshot not reported)Accuracy0.12 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4 Turbo (LAB-Bench snapshot not reported)Precision (selective accuracy)0.22 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4 Turbo (LAB-Bench snapshot not reported)Coverage0.55 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Haiku (claude-3-haiku-20240307)Accuracy0.18 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Haiku (claude-3-haiku-20240307)Precision (selective accuracy)0.22 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Haiku (claude-3-haiku-20240307)Coverage0.81 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Meta-Llama-3-70B-Instruct (Anyscale API)Accuracy0.2 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Meta-Llama-3-70B-Instruct (Anyscale API)Precision (selective accuracy)0.21 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Meta-Llama-3-70B-Instruct (Anyscale API)Coverage0.97 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50

Evidence

  • section: Appendix D.1–D.4; Table 1 and Appendix Table 6 (Full scope, zero-shot chain-of-thought prompt, no tools, three runs, exact metrics, parser, and model identities.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics
  • table: Tables 2–4, row DbQA_gene_location_task-v1 (Accuracy, precision, and coverage. Human row excluded.) — supports /results

Evaluation run

lab-bench-dbqa-mirna-targets-creator-mcq

From LAB-Bench: Measuring Capabilities of Language Models for Biology Research

lab-bench-dbqa-mirna-targets-paper-v3-zero-shot-cot-no-toolsvpaper-v3

Evaluated models / systems: Claude 3.5 Sonnet (claude-3-5-sonnet-20240620), Claude 3 Haiku (claude-3-haiku-20240307), Claude 3 Opus (claude-3-opus-20240229), Gemini 1.5 Pro (gemini-1.5-pro-001), GPT-4 Turbo (LAB-Bench snapshot not reported), GPT-4o (LAB-Bench snapshot not reported), Meta-Llama-3-70B-Instruct (Anyscale API)

Scopefull · n=50
Shots0
Turnssingle-turn
System prompt publicYes
Reasoning / effortzero-shot chain-of-thought
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNo
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
Repeats3
Graderregex multiple-choice parser with Claude 2 fallback · model: Claude 2 fallback parser · human review: no
Statisticsmean across three runs per model
Contaminationfull pre-release snapshot with 80% later public and 20% private contamination-monitoring splits
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracyabsoluteproportioncorrect / all questions; mean over three runsexact choice
Precision (selective accuracy)absoluteproportioncorrect / attempted questions; mean over three runsinsufficient-information responses excluded from denominator
Coverageabsoluteproportionattempted / all questions; mean over three runsinsufficient-information responses treated as not attempted

Results

ModelMetricValuen
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Accuracy0 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Precision (selective accuracy)0 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Coverage0 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Opus (claude-3-opus-20240229)Accuracy0 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Opus (claude-3-opus-20240229)Precision (selective accuracy)0 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Opus (claude-3-opus-20240229)Coverage0 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Gemini 1.5 Pro (gemini-1.5-pro-001)Accuracy0 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Gemini 1.5 Pro (gemini-1.5-pro-001)Precision (selective accuracy)0 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Gemini 1.5 Pro (gemini-1.5-pro-001)Coverage0 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4o (LAB-Bench snapshot not reported)Accuracy0.1 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4o (LAB-Bench snapshot not reported)Precision (selective accuracy)0.3 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4o (LAB-Bench snapshot not reported)Coverage0.35 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4 Turbo (LAB-Bench snapshot not reported)Accuracy0.01 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4 Turbo (LAB-Bench snapshot not reported)Precision (selective accuracy)0.03 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4 Turbo (LAB-Bench snapshot not reported)Coverage0.15 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Haiku (claude-3-haiku-20240307)Accuracy0.15 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Haiku (claude-3-haiku-20240307)Precision (selective accuracy)0.26 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Haiku (claude-3-haiku-20240307)Coverage0.6 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Meta-Llama-3-70B-Instruct (Anyscale API)Accuracy0.29 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Meta-Llama-3-70B-Instruct (Anyscale API)Precision (selective accuracy)0.29 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Meta-Llama-3-70B-Instruct (Anyscale API)Coverage0.99 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50

Evidence

  • section: Appendix D.1–D.4; Table 1 and Appendix Table 6 (Full scope, zero-shot chain-of-thought prompt, no tools, three runs, exact metrics, parser, and model identities.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics
  • table: Tables 2–4, row DbQA_mirna_targets_task-v1 (Accuracy, precision, and coverage. Human row excluded.) — supports /results

Evaluation run

lab-bench-dbqa-mouse-tumor-gene-sets-creator-mcq

From LAB-Bench: Measuring Capabilities of Language Models for Biology Research

lab-bench-dbqa-mouse-tumor-gene-sets-paper-v3-zero-shot-cot-no-toolsvpaper-v3

Evaluated models / systems: Claude 3.5 Sonnet (claude-3-5-sonnet-20240620), Claude 3 Haiku (claude-3-haiku-20240307), Claude 3 Opus (claude-3-opus-20240229), Gemini 1.5 Pro (gemini-1.5-pro-001), GPT-4 Turbo (LAB-Bench snapshot not reported), GPT-4o (LAB-Bench snapshot not reported), Meta-Llama-3-70B-Instruct (Anyscale API)

Scopefull · n=100
Shots0
Turnssingle-turn
System prompt publicYes
Reasoning / effortzero-shot chain-of-thought
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNo
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
Repeats3
Graderregex multiple-choice parser with Claude 2 fallback · model: Claude 2 fallback parser · human review: no
Statisticsmean across three runs per model
Contaminationfull pre-release snapshot with 80% later public and 20% private contamination-monitoring splits
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracyabsoluteproportioncorrect / all questions; mean over three runsexact choice
Precision (selective accuracy)absoluteproportioncorrect / attempted questions; mean over three runsinsufficient-information responses excluded from denominator
Coverageabsoluteproportionattempted / all questions; mean over three runsinsufficient-information responses treated as not attempted

Results

ModelMetricValuen
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Accuracy0.55 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Precision (selective accuracy)0.73 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Coverage0.76 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
Claude 3 Opus (claude-3-opus-20240229)Accuracy0.36 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
Claude 3 Opus (claude-3-opus-20240229)Precision (selective accuracy)0.71 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
Claude 3 Opus (claude-3-opus-20240229)Coverage0.51 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
Gemini 1.5 Pro (gemini-1.5-pro-001)Accuracy0.1 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
Gemini 1.5 Pro (gemini-1.5-pro-001)Precision (selective accuracy)0.76 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
Gemini 1.5 Pro (gemini-1.5-pro-001)Coverage0.14 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
GPT-4o (LAB-Bench snapshot not reported)Accuracy0.57 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
GPT-4o (LAB-Bench snapshot not reported)Precision (selective accuracy)0.59 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
GPT-4o (LAB-Bench snapshot not reported)Coverage0.96 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
GPT-4 Turbo (LAB-Bench snapshot not reported)Accuracy0.16 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
GPT-4 Turbo (LAB-Bench snapshot not reported)Precision (selective accuracy)0.58 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
GPT-4 Turbo (LAB-Bench snapshot not reported)Coverage0.28 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
Claude 3 Haiku (claude-3-haiku-20240307)Accuracy0.48 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
Claude 3 Haiku (claude-3-haiku-20240307)Precision (selective accuracy)0.55 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
Claude 3 Haiku (claude-3-haiku-20240307)Coverage0.87 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
Meta-Llama-3-70B-Instruct (Anyscale API)Accuracy0.53 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
Meta-Llama-3-70B-Instruct (Anyscale API)Precision (selective accuracy)0.58 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
Meta-Llama-3-70B-Instruct (Anyscale API)Coverage0.93 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100

Evidence

  • section: Appendix D.1–D.4; Table 1 and Appendix Table 6 (Full scope, zero-shot chain-of-thought prompt, no tools, three runs, exact metrics, parser, and model identities.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics
  • table: Tables 2–4, row DbQA_mouse_tumor_gene_sets-v1 (Accuracy, precision, and coverage. Human row excluded.) — supports /results

Evaluation run

lab-bench-dbqa-oncogenic-signatures-creator-mcq

From LAB-Bench: Measuring Capabilities of Language Models for Biology Research

lab-bench-dbqa-oncogenic-signatures-paper-v3-zero-shot-cot-no-toolsvpaper-v3

Evaluated models / systems: Claude 3.5 Sonnet (claude-3-5-sonnet-20240620), Claude 3 Haiku (claude-3-haiku-20240307), Claude 3 Opus (claude-3-opus-20240229), Gemini 1.5 Pro (gemini-1.5-pro-001), GPT-4 Turbo (LAB-Bench snapshot not reported), GPT-4o (LAB-Bench snapshot not reported), Meta-Llama-3-70B-Instruct (Anyscale API)

Scopefull · n=50
Shots0
Turnssingle-turn
System prompt publicYes
Reasoning / effortzero-shot chain-of-thought
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNo
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
Repeats3
Graderregex multiple-choice parser with Claude 2 fallback · model: Claude 2 fallback parser · human review: no
Statisticsmean across three runs per model
Contaminationfull pre-release snapshot with 80% later public and 20% private contamination-monitoring splits
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracyabsoluteproportioncorrect / all questions; mean over three runsexact choice
Precision (selective accuracy)absoluteproportioncorrect / attempted questions; mean over three runsinsufficient-information responses excluded from denominator
Coverageabsoluteproportionattempted / all questions; mean over three runsinsufficient-information responses treated as not attempted

Results

ModelMetricValuen
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Accuracy0.14 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Precision (selective accuracy)0.59 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Coverage0.24 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Opus (claude-3-opus-20240229)Accuracy0.12 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Opus (claude-3-opus-20240229)Precision (selective accuracy)0.57 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Opus (claude-3-opus-20240229)Coverage0.2 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Gemini 1.5 Pro (gemini-1.5-pro-001)Accuracy0.04 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Gemini 1.5 Pro (gemini-1.5-pro-001)Precision (selective accuracy)0.75 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Gemini 1.5 Pro (gemini-1.5-pro-001)Coverage0.05 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4o (LAB-Bench snapshot not reported)Accuracy0.25 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4o (LAB-Bench snapshot not reported)Precision (selective accuracy)0.32 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4o (LAB-Bench snapshot not reported)Coverage0.79 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4 Turbo (LAB-Bench snapshot not reported)Accuracy0.06 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4 Turbo (LAB-Bench snapshot not reported)Precision (selective accuracy)0.61 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4 Turbo (LAB-Bench snapshot not reported)Coverage0.1 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Haiku (claude-3-haiku-20240307)Accuracy0.24 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Haiku (claude-3-haiku-20240307)Precision (selective accuracy)0.3 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Haiku (claude-3-haiku-20240307)Coverage0.8 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Meta-Llama-3-70B-Instruct (Anyscale API)Accuracy0.2 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Meta-Llama-3-70B-Instruct (Anyscale API)Precision (selective accuracy)0.47 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Meta-Llama-3-70B-Instruct (Anyscale API)Coverage0.42 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50

Evidence

  • section: Appendix D.1–D.4; Table 1 and Appendix Table 6 (Full scope, zero-shot chain-of-thought prompt, no tools, three runs, exact metrics, parser, and model identities.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics
  • table: Tables 2–4, row DbQA_oncogenic_signatures_task-v1 (Accuracy, precision, and coverage. Human row excluded.) — supports /results

Evaluation run

lab-bench-dbqa-tfbs-gtrd-creator-mcq

From LAB-Bench: Measuring Capabilities of Language Models for Biology Research

lab-bench-dbqa-tfbs-gtrd-paper-v3-zero-shot-cot-no-toolsvpaper-v3

Evaluated models / systems: Claude 3.5 Sonnet (claude-3-5-sonnet-20240620), Claude 3 Haiku (claude-3-haiku-20240307), Claude 3 Opus (claude-3-opus-20240229), Gemini 1.5 Pro (gemini-1.5-pro-001), GPT-4 Turbo (LAB-Bench snapshot not reported), GPT-4o (LAB-Bench snapshot not reported), Meta-Llama-3-70B-Instruct (Anyscale API)

Scopefull · n=50
Shots0
Turnssingle-turn
System prompt publicYes
Reasoning / effortzero-shot chain-of-thought
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNo
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
Repeats3
Graderregex multiple-choice parser with Claude 2 fallback · model: Claude 2 fallback parser · human review: no
Statisticsmean across three runs per model
Contaminationfull pre-release snapshot with 80% later public and 20% private contamination-monitoring splits
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracyabsoluteproportioncorrect / all questions; mean over three runsexact choice
Precision (selective accuracy)absoluteproportioncorrect / attempted questions; mean over three runsinsufficient-information responses excluded from denominator
Coverageabsoluteproportionattempted / all questions; mean over three runsinsufficient-information responses treated as not attempted

Results

ModelMetricValuen
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Accuracy0 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Precision (selective accuracy)0 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Coverage0 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Opus (claude-3-opus-20240229)Accuracy0 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Opus (claude-3-opus-20240229)Precision (selective accuracy)0 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Opus (claude-3-opus-20240229)Coverage0 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Gemini 1.5 Pro (gemini-1.5-pro-001)Accuracy0 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Gemini 1.5 Pro (gemini-1.5-pro-001)Precision (selective accuracy)0 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Gemini 1.5 Pro (gemini-1.5-pro-001)Coverage0.01 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4o (LAB-Bench snapshot not reported)Accuracy0.01 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4o (LAB-Bench snapshot not reported)Precision (selective accuracy)0.33 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4o (LAB-Bench snapshot not reported)Coverage0.01 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4 Turbo (LAB-Bench snapshot not reported)Accuracy0 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4 Turbo (LAB-Bench snapshot not reported)Precision (selective accuracy)0 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4 Turbo (LAB-Bench snapshot not reported)Coverage0.11 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Haiku (claude-3-haiku-20240307)Accuracy0.11 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Haiku (claude-3-haiku-20240307)Precision (selective accuracy)0.31 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Haiku (claude-3-haiku-20240307)Coverage0.35 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Meta-Llama-3-70B-Instruct (Anyscale API)Accuracy0.19 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Meta-Llama-3-70B-Instruct (Anyscale API)Precision (selective accuracy)0.19 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Meta-Llama-3-70B-Instruct (Anyscale API)Coverage1 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50

Evidence

  • section: Appendix D.1–D.4; Table 1 and Appendix Table 6 (Full scope, zero-shot chain-of-thought prompt, no tools, three runs, exact metrics, parser, and model identities.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics
  • table: Tables 2–4, row DbQA_tfbs_GTRD_task-v1 (Accuracy, precision, and coverage. Human row excluded.) — supports /results

Evaluation run

lab-bench-dbqa-variant-from-sequence-creator-mcq

From LAB-Bench: Measuring Capabilities of Language Models for Biology Research

lab-bench-dbqa-variant-from-sequence-paper-v3-zero-shot-cot-no-toolsvpaper-v3

Evaluated models / systems: Claude 3.5 Sonnet (claude-3-5-sonnet-20240620), Claude 3 Haiku (claude-3-haiku-20240307), Claude 3 Opus (claude-3-opus-20240229), Gemini 1.5 Pro (gemini-1.5-pro-001), GPT-4 Turbo (LAB-Bench snapshot not reported), GPT-4o (LAB-Bench snapshot not reported), Meta-Llama-3-70B-Instruct (Anyscale API)

Scopefull · n=100
Shots0
Turnssingle-turn
System prompt publicYes
Reasoning / effortzero-shot chain-of-thought
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNo
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
Repeats3
Graderregex multiple-choice parser with Claude 2 fallback · model: Claude 2 fallback parser · human review: no
Statisticsmean across three runs per model
Contaminationfull pre-release snapshot with 80% later public and 20% private contamination-monitoring splits
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracyabsoluteproportioncorrect / all questions; mean over three runsexact choice
Precision (selective accuracy)absoluteproportioncorrect / attempted questions; mean over three runsinsufficient-information responses excluded from denominator
Coverageabsoluteproportionattempted / all questions; mean over three runsinsufficient-information responses treated as not attempted

Results

ModelMetricValuen
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Accuracy0.02 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Precision (selective accuracy)0.75 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Coverage0.03 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
Claude 3 Opus (claude-3-opus-20240229)Accuracy0 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
Claude 3 Opus (claude-3-opus-20240229)Precision (selective accuracy)0.11 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
Claude 3 Opus (claude-3-opus-20240229)Coverage0.03 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
Gemini 1.5 Pro (gemini-1.5-pro-001)Accuracy0 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
Gemini 1.5 Pro (gemini-1.5-pro-001)Precision (selective accuracy)0.17 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
Gemini 1.5 Pro (gemini-1.5-pro-001)Coverage0.01 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
GPT-4o (LAB-Bench snapshot not reported)Accuracy0.42 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
GPT-4o (LAB-Bench snapshot not reported)Precision (selective accuracy)0.44 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
GPT-4o (LAB-Bench snapshot not reported)Coverage0.95 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
GPT-4 Turbo (LAB-Bench snapshot not reported)Accuracy0 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
GPT-4 Turbo (LAB-Bench snapshot not reported)Precision (selective accuracy)0 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
GPT-4 Turbo (LAB-Bench snapshot not reported)Coverage0.01 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
Claude 3 Haiku (claude-3-haiku-20240307)Accuracy0.22 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
Claude 3 Haiku (claude-3-haiku-20240307)Precision (selective accuracy)0.36 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
Claude 3 Haiku (claude-3-haiku-20240307)Coverage0.62 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
Meta-Llama-3-70B-Instruct (Anyscale API)Accuracy0.3 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
Meta-Llama-3-70B-Instruct (Anyscale API)Precision (selective accuracy)0.35 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
Meta-Llama-3-70B-Instruct (Anyscale API)Coverage0.85 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100

Evidence

  • section: Appendix D.1–D.4; Table 1 and Appendix Table 6 (Full scope, zero-shot chain-of-thought prompt, no tools, three runs, exact metrics, parser, and model identities.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics
  • table: Tables 2–4, row DbQA_variant_from_sequence_task-v1 (Accuracy, precision, and coverage. Human row excluded.) — supports /results

Evaluation run

lab-bench-dbqa-variant-multi-sequence-creator-mcq

From LAB-Bench: Measuring Capabilities of Language Models for Biology Research

lab-bench-dbqa-variant-multi-sequence-paper-v3-zero-shot-cot-no-toolsvpaper-v3

Evaluated models / systems: Claude 3.5 Sonnet (claude-3-5-sonnet-20240620), Claude 3 Haiku (claude-3-haiku-20240307), Claude 3 Opus (claude-3-opus-20240229), Gemini 1.5 Pro (gemini-1.5-pro-001), GPT-4 Turbo (LAB-Bench snapshot not reported), GPT-4o (LAB-Bench snapshot not reported), Meta-Llama-3-70B-Instruct (Anyscale API)

Scopefull · n=100
Shots0
Turnssingle-turn
System prompt publicYes
Reasoning / effortzero-shot chain-of-thought
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNo
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
Repeats3
Graderregex multiple-choice parser with Claude 2 fallback · model: Claude 2 fallback parser · human review: no
Statisticsmean across three runs per model
Contaminationfull pre-release snapshot with 80% later public and 20% private contamination-monitoring splits
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracyabsoluteproportioncorrect / all questions; mean over three runsexact choice
Precision (selective accuracy)absoluteproportioncorrect / attempted questions; mean over three runsinsufficient-information responses excluded from denominator
Coverageabsoluteproportionattempted / all questions; mean over three runsinsufficient-information responses treated as not attempted

Results

ModelMetricValuen
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Accuracy0.09 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Precision (selective accuracy)0.16 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Coverage0.56 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
Claude 3 Opus (claude-3-opus-20240229)Accuracy0.05 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
Claude 3 Opus (claude-3-opus-20240229)Precision (selective accuracy)0.18 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
Claude 3 Opus (claude-3-opus-20240229)Coverage0.26 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
Gemini 1.5 Pro (gemini-1.5-pro-001)Accuracy0 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
Gemini 1.5 Pro (gemini-1.5-pro-001)Precision (selective accuracy)0.33 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
Gemini 1.5 Pro (gemini-1.5-pro-001)Coverage0.04 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
GPT-4o (LAB-Bench snapshot not reported)Accuracy0.06 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
GPT-4o (LAB-Bench snapshot not reported)Precision (selective accuracy)0.12 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
GPT-4o (LAB-Bench snapshot not reported)Coverage0.49 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
GPT-4 Turbo (LAB-Bench snapshot not reported)Accuracy0 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
GPT-4 Turbo (LAB-Bench snapshot not reported)Precision (selective accuracy)0 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
GPT-4 Turbo (LAB-Bench snapshot not reported)Coverage0.02 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
Claude 3 Haiku (claude-3-haiku-20240307)Accuracy0.14 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
Claude 3 Haiku (claude-3-haiku-20240307)Precision (selective accuracy)0.16 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
Claude 3 Haiku (claude-3-haiku-20240307)Coverage0.84 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
Meta-Llama-3-70B-Instruct (Anyscale API)Accuracy0.07 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
Meta-Llama-3-70B-Instruct (Anyscale API)Precision (selective accuracy)0.12 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100
Meta-Llama-3-70B-Instruct (Anyscale API)Coverage0.63 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
100

Evidence

  • section: Appendix D.1–D.4; Table 1 and Appendix Table 6 (Full scope, zero-shot chain-of-thought prompt, no tools, three runs, exact metrics, parser, and model identities.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics
  • table: Tables 2–4, row DbQA_variant_multi_sequence_task-v1 (Accuracy, precision, and coverage. Human row excluded.) — supports /results

Evaluation run

lab-bench-dbqa-vax-response-creator-mcq

From LAB-Bench: Measuring Capabilities of Language Models for Biology Research

lab-bench-dbqa-vax-response-paper-v3-zero-shot-cot-no-toolsvpaper-v3

Evaluated models / systems: Claude 3.5 Sonnet (claude-3-5-sonnet-20240620), Claude 3 Haiku (claude-3-haiku-20240307), Claude 3 Opus (claude-3-opus-20240229), Gemini 1.5 Pro (gemini-1.5-pro-001), GPT-4 Turbo (LAB-Bench snapshot not reported), GPT-4o (LAB-Bench snapshot not reported), Meta-Llama-3-70B-Instruct (Anyscale API)

Scopefull · n=50
Shots0
Turnssingle-turn
System prompt publicYes
Reasoning / effortzero-shot chain-of-thought
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNo
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
Repeats3
Graderregex multiple-choice parser with Claude 2 fallback · model: Claude 2 fallback parser · human review: no
Statisticsmean across three runs per model
Contaminationfull pre-release snapshot with 80% later public and 20% private contamination-monitoring splits
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracyabsoluteproportioncorrect / all questions; mean over three runsexact choice
Precision (selective accuracy)absoluteproportioncorrect / attempted questions; mean over three runsinsufficient-information responses excluded from denominator
Coverageabsoluteproportionattempted / all questions; mean over three runsinsufficient-information responses treated as not attempted

Results

ModelMetricValuen
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Accuracy0.21 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Precision (selective accuracy)0.57 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Coverage0.36 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Opus (claude-3-opus-20240229)Accuracy0.14 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Opus (claude-3-opus-20240229)Precision (selective accuracy)0.65 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Opus (claude-3-opus-20240229)Coverage0.21 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Gemini 1.5 Pro (gemini-1.5-pro-001)Accuracy0.13 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Gemini 1.5 Pro (gemini-1.5-pro-001)Precision (selective accuracy)0.61 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Gemini 1.5 Pro (gemini-1.5-pro-001)Coverage0.21 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4o (LAB-Bench snapshot not reported)Accuracy0.33 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4o (LAB-Bench snapshot not reported)Precision (selective accuracy)0.45 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4o (LAB-Bench snapshot not reported)Coverage0.73 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4 Turbo (LAB-Bench snapshot not reported)Accuracy0.06 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4 Turbo (LAB-Bench snapshot not reported)Precision (selective accuracy)0.65 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4 Turbo (LAB-Bench snapshot not reported)Coverage0.09 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Haiku (claude-3-haiku-20240307)Accuracy0.32 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Haiku (claude-3-haiku-20240307)Precision (selective accuracy)0.38 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Haiku (claude-3-haiku-20240307)Coverage0.83 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Meta-Llama-3-70B-Instruct (Anyscale API)Accuracy0.15 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Meta-Llama-3-70B-Instruct (Anyscale API)Precision (selective accuracy)0.54 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Meta-Llama-3-70B-Instruct (Anyscale API)Coverage0.29 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50

Evidence

  • section: Appendix D.1–D.4; Table 1 and Appendix Table 6 (Full scope, zero-shot chain-of-thought prompt, no tools, three runs, exact metrics, parser, and model identities.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics
  • table: Tables 2–4, row DbQA_vax_response_task-v1 (Accuracy, precision, and coverage. Human row excluded.) — supports /results

Evaluation run

lab-bench-dbqa-viral-ppi-creator-mcq

From LAB-Bench: Measuring Capabilities of Language Models for Biology Research

lab-bench-dbqa-viral-ppi-paper-v3-zero-shot-cot-no-toolsvpaper-v3

Evaluated models / systems: Claude 3.5 Sonnet (claude-3-5-sonnet-20240620), Claude 3 Haiku (claude-3-haiku-20240307), Claude 3 Opus (claude-3-opus-20240229), Gemini 1.5 Pro (gemini-1.5-pro-001), GPT-4 Turbo (LAB-Bench snapshot not reported), GPT-4o (LAB-Bench snapshot not reported), Meta-Llama-3-70B-Instruct (Anyscale API)

Scopefull · n=50
Shots0
Turnssingle-turn
System prompt publicYes
Reasoning / effortzero-shot chain-of-thought
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNo
External toolsNo
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
Repeats3
Graderregex multiple-choice parser with Claude 2 fallback · model: Claude 2 fallback parser · human review: no
Statisticsmean across three runs per model
Contaminationfull pre-release snapshot with 80% later public and 20% private contamination-monitoring splits
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Accuracyabsoluteproportioncorrect / all questions; mean over three runsexact choice
Precision (selective accuracy)absoluteproportioncorrect / attempted questions; mean over three runsinsufficient-information responses excluded from denominator
Coverageabsoluteproportionattempted / all questions; mean over three runsinsufficient-information responses treated as not attempted

Results

ModelMetricValuen
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Accuracy0 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Precision (selective accuracy)0 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)Coverage0 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Opus (claude-3-opus-20240229)Accuracy0 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Opus (claude-3-opus-20240229)Precision (selective accuracy)0 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Opus (claude-3-opus-20240229)Coverage0 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Gemini 1.5 Pro (gemini-1.5-pro-001)Accuracy0 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Gemini 1.5 Pro (gemini-1.5-pro-001)Precision (selective accuracy)0 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Gemini 1.5 Pro (gemini-1.5-pro-001)Coverage0 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4o (LAB-Bench snapshot not reported)Accuracy0.26 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4o (LAB-Bench snapshot not reported)Precision (selective accuracy)0.69 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4o (LAB-Bench snapshot not reported)Coverage0.38 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4 Turbo (LAB-Bench snapshot not reported)Accuracy0.02 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4 Turbo (LAB-Bench snapshot not reported)Precision (selective accuracy)0.28 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
GPT-4 Turbo (LAB-Bench snapshot not reported)Coverage0.06 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Haiku (claude-3-haiku-20240307)Accuracy0.27 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Haiku (claude-3-haiku-20240307)Precision (selective accuracy)0.49 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Claude 3 Haiku (claude-3-haiku-20240307)Coverage0.55 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Meta-Llama-3-70B-Instruct (Anyscale API)Accuracy0.4 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Meta-Llama-3-70B-Instruct (Anyscale API)Precision (selective accuracy)0.4 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50
Meta-Llama-3-70B-Instruct (Anyscale API)Coverage1 proportion
Creator-paper full-snapshot result; human-baseline row is not encoded as a model result.
50

Evidence

  • section: Appendix D.1–D.4; Table 1 and Appendix Table 6 (Full scope, zero-shot chain-of-thought prompt, no tools, three runs, exact metrics, parser, and model identities.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics
  • table: Tables 2–4, row DbQA_viral_ppi_task-v1 (Accuracy, precision, and coverage. Human row excluded.) — supports /results

Comparable result views

Accuracy

lab-bench-dbqa-dga-creator-mcq · lab-bench-dbqa-dga-paper-v3-zero-shot-cot-no-tools

CSV ↓
Accessible data table
ModelValueComparability group
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)0lab-bench-dbqa-dga-paper-v3-zero-shot-cot-no-tools
Claude 3 Opus (claude-3-opus-20240229)0lab-bench-dbqa-dga-paper-v3-zero-shot-cot-no-tools
Gemini 1.5 Pro (gemini-1.5-pro-001)0lab-bench-dbqa-dga-paper-v3-zero-shot-cot-no-tools
GPT-4o (LAB-Bench snapshot not reported)0.16lab-bench-dbqa-dga-paper-v3-zero-shot-cot-no-tools
GPT-4 Turbo (LAB-Bench snapshot not reported)0.05lab-bench-dbqa-dga-paper-v3-zero-shot-cot-no-tools
Claude 3 Haiku (claude-3-haiku-20240307)0.15lab-bench-dbqa-dga-paper-v3-zero-shot-cot-no-tools
Meta-Llama-3-70B-Instruct (Anyscale API)0.25lab-bench-dbqa-dga-paper-v3-zero-shot-cot-no-tools

Precision (selective accuracy)

lab-bench-dbqa-dga-creator-mcq · lab-bench-dbqa-dga-paper-v3-zero-shot-cot-no-tools

CSV ↓
Accessible data table
ModelValueComparability group
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)0lab-bench-dbqa-dga-paper-v3-zero-shot-cot-no-tools
Claude 3 Opus (claude-3-opus-20240229)0lab-bench-dbqa-dga-paper-v3-zero-shot-cot-no-tools
Gemini 1.5 Pro (gemini-1.5-pro-001)0lab-bench-dbqa-dga-paper-v3-zero-shot-cot-no-tools
GPT-4o (LAB-Bench snapshot not reported)0.17lab-bench-dbqa-dga-paper-v3-zero-shot-cot-no-tools
GPT-4 Turbo (LAB-Bench snapshot not reported)0.17lab-bench-dbqa-dga-paper-v3-zero-shot-cot-no-tools
Claude 3 Haiku (claude-3-haiku-20240307)0.17lab-bench-dbqa-dga-paper-v3-zero-shot-cot-no-tools
Meta-Llama-3-70B-Instruct (Anyscale API)0.25lab-bench-dbqa-dga-paper-v3-zero-shot-cot-no-tools

Coverage

lab-bench-dbqa-dga-creator-mcq · lab-bench-dbqa-dga-paper-v3-zero-shot-cot-no-tools

CSV ↓
Accessible data table
ModelValueComparability group
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)0lab-bench-dbqa-dga-paper-v3-zero-shot-cot-no-tools
Claude 3 Opus (claude-3-opus-20240229)0.01lab-bench-dbqa-dga-paper-v3-zero-shot-cot-no-tools
Gemini 1.5 Pro (gemini-1.5-pro-001)0lab-bench-dbqa-dga-paper-v3-zero-shot-cot-no-tools
GPT-4o (LAB-Bench snapshot not reported)0.95lab-bench-dbqa-dga-paper-v3-zero-shot-cot-no-tools
GPT-4 Turbo (LAB-Bench snapshot not reported)0.32lab-bench-dbqa-dga-paper-v3-zero-shot-cot-no-tools
Claude 3 Haiku (claude-3-haiku-20240307)0.9lab-bench-dbqa-dga-paper-v3-zero-shot-cot-no-tools
Meta-Llama-3-70B-Instruct (Anyscale API)0.99lab-bench-dbqa-dga-paper-v3-zero-shot-cot-no-tools

Accuracy

lab-bench-dbqa-gene-location-creator-mcq · lab-bench-dbqa-gene-location-paper-v3-zero-shot-cot-no-tools

CSV ↓
Accessible data table
ModelValueComparability group
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)0lab-bench-dbqa-gene-location-paper-v3-zero-shot-cot-no-tools
Claude 3 Opus (claude-3-opus-20240229)0lab-bench-dbqa-gene-location-paper-v3-zero-shot-cot-no-tools
Gemini 1.5 Pro (gemini-1.5-pro-001)0.03lab-bench-dbqa-gene-location-paper-v3-zero-shot-cot-no-tools
GPT-4o (LAB-Bench snapshot not reported)0.31lab-bench-dbqa-gene-location-paper-v3-zero-shot-cot-no-tools
GPT-4 Turbo (LAB-Bench snapshot not reported)0.12lab-bench-dbqa-gene-location-paper-v3-zero-shot-cot-no-tools
Claude 3 Haiku (claude-3-haiku-20240307)0.18lab-bench-dbqa-gene-location-paper-v3-zero-shot-cot-no-tools
Meta-Llama-3-70B-Instruct (Anyscale API)0.2lab-bench-dbqa-gene-location-paper-v3-zero-shot-cot-no-tools

Precision (selective accuracy)

lab-bench-dbqa-gene-location-creator-mcq · lab-bench-dbqa-gene-location-paper-v3-zero-shot-cot-no-tools

CSV ↓
Accessible data table
ModelValueComparability group
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)0lab-bench-dbqa-gene-location-paper-v3-zero-shot-cot-no-tools
Claude 3 Opus (claude-3-opus-20240229)0lab-bench-dbqa-gene-location-paper-v3-zero-shot-cot-no-tools
Gemini 1.5 Pro (gemini-1.5-pro-001)0.22lab-bench-dbqa-gene-location-paper-v3-zero-shot-cot-no-tools
GPT-4o (LAB-Bench snapshot not reported)0.38lab-bench-dbqa-gene-location-paper-v3-zero-shot-cot-no-tools
GPT-4 Turbo (LAB-Bench snapshot not reported)0.22lab-bench-dbqa-gene-location-paper-v3-zero-shot-cot-no-tools
Claude 3 Haiku (claude-3-haiku-20240307)0.22lab-bench-dbqa-gene-location-paper-v3-zero-shot-cot-no-tools
Meta-Llama-3-70B-Instruct (Anyscale API)0.21lab-bench-dbqa-gene-location-paper-v3-zero-shot-cot-no-tools

Coverage

lab-bench-dbqa-gene-location-creator-mcq · lab-bench-dbqa-gene-location-paper-v3-zero-shot-cot-no-tools

CSV ↓
Accessible data table
ModelValueComparability group
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)0lab-bench-dbqa-gene-location-paper-v3-zero-shot-cot-no-tools
Claude 3 Opus (claude-3-opus-20240229)0.01lab-bench-dbqa-gene-location-paper-v3-zero-shot-cot-no-tools
Gemini 1.5 Pro (gemini-1.5-pro-001)0.09lab-bench-dbqa-gene-location-paper-v3-zero-shot-cot-no-tools
GPT-4o (LAB-Bench snapshot not reported)0.81lab-bench-dbqa-gene-location-paper-v3-zero-shot-cot-no-tools
GPT-4 Turbo (LAB-Bench snapshot not reported)0.55lab-bench-dbqa-gene-location-paper-v3-zero-shot-cot-no-tools
Claude 3 Haiku (claude-3-haiku-20240307)0.81lab-bench-dbqa-gene-location-paper-v3-zero-shot-cot-no-tools
Meta-Llama-3-70B-Instruct (Anyscale API)0.97lab-bench-dbqa-gene-location-paper-v3-zero-shot-cot-no-tools

Accuracy

lab-bench-dbqa-mirna-targets-creator-mcq · lab-bench-dbqa-mirna-targets-paper-v3-zero-shot-cot-no-tools

CSV ↓
Accessible data table
ModelValueComparability group
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)0lab-bench-dbqa-mirna-targets-paper-v3-zero-shot-cot-no-tools
Claude 3 Opus (claude-3-opus-20240229)0lab-bench-dbqa-mirna-targets-paper-v3-zero-shot-cot-no-tools
Gemini 1.5 Pro (gemini-1.5-pro-001)0lab-bench-dbqa-mirna-targets-paper-v3-zero-shot-cot-no-tools
GPT-4o (LAB-Bench snapshot not reported)0.1lab-bench-dbqa-mirna-targets-paper-v3-zero-shot-cot-no-tools
GPT-4 Turbo (LAB-Bench snapshot not reported)0.01lab-bench-dbqa-mirna-targets-paper-v3-zero-shot-cot-no-tools
Claude 3 Haiku (claude-3-haiku-20240307)0.15lab-bench-dbqa-mirna-targets-paper-v3-zero-shot-cot-no-tools
Meta-Llama-3-70B-Instruct (Anyscale API)0.29lab-bench-dbqa-mirna-targets-paper-v3-zero-shot-cot-no-tools

Precision (selective accuracy)

lab-bench-dbqa-mirna-targets-creator-mcq · lab-bench-dbqa-mirna-targets-paper-v3-zero-shot-cot-no-tools

CSV ↓
Accessible data table
ModelValueComparability group
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)0lab-bench-dbqa-mirna-targets-paper-v3-zero-shot-cot-no-tools
Claude 3 Opus (claude-3-opus-20240229)0lab-bench-dbqa-mirna-targets-paper-v3-zero-shot-cot-no-tools
Gemini 1.5 Pro (gemini-1.5-pro-001)0lab-bench-dbqa-mirna-targets-paper-v3-zero-shot-cot-no-tools
GPT-4o (LAB-Bench snapshot not reported)0.3lab-bench-dbqa-mirna-targets-paper-v3-zero-shot-cot-no-tools
GPT-4 Turbo (LAB-Bench snapshot not reported)0.03lab-bench-dbqa-mirna-targets-paper-v3-zero-shot-cot-no-tools
Claude 3 Haiku (claude-3-haiku-20240307)0.26lab-bench-dbqa-mirna-targets-paper-v3-zero-shot-cot-no-tools
Meta-Llama-3-70B-Instruct (Anyscale API)0.29lab-bench-dbqa-mirna-targets-paper-v3-zero-shot-cot-no-tools

Coverage

lab-bench-dbqa-mirna-targets-creator-mcq · lab-bench-dbqa-mirna-targets-paper-v3-zero-shot-cot-no-tools

CSV ↓
Accessible data table
ModelValueComparability group
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)0lab-bench-dbqa-mirna-targets-paper-v3-zero-shot-cot-no-tools
Claude 3 Opus (claude-3-opus-20240229)0lab-bench-dbqa-mirna-targets-paper-v3-zero-shot-cot-no-tools
Gemini 1.5 Pro (gemini-1.5-pro-001)0lab-bench-dbqa-mirna-targets-paper-v3-zero-shot-cot-no-tools
GPT-4o (LAB-Bench snapshot not reported)0.35lab-bench-dbqa-mirna-targets-paper-v3-zero-shot-cot-no-tools
GPT-4 Turbo (LAB-Bench snapshot not reported)0.15lab-bench-dbqa-mirna-targets-paper-v3-zero-shot-cot-no-tools
Claude 3 Haiku (claude-3-haiku-20240307)0.6lab-bench-dbqa-mirna-targets-paper-v3-zero-shot-cot-no-tools
Meta-Llama-3-70B-Instruct (Anyscale API)0.99lab-bench-dbqa-mirna-targets-paper-v3-zero-shot-cot-no-tools

Accuracy

lab-bench-dbqa-mouse-tumor-gene-sets-creator-mcq · lab-bench-dbqa-mouse-tumor-gene-sets-paper-v3-zero-shot-cot-no-tools

CSV ↓
Accessible data table
ModelValueComparability group
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)0.55lab-bench-dbqa-mouse-tumor-gene-sets-paper-v3-zero-shot-cot-no-tools
Claude 3 Opus (claude-3-opus-20240229)0.36lab-bench-dbqa-mouse-tumor-gene-sets-paper-v3-zero-shot-cot-no-tools
Gemini 1.5 Pro (gemini-1.5-pro-001)0.1lab-bench-dbqa-mouse-tumor-gene-sets-paper-v3-zero-shot-cot-no-tools
GPT-4o (LAB-Bench snapshot not reported)0.57lab-bench-dbqa-mouse-tumor-gene-sets-paper-v3-zero-shot-cot-no-tools
GPT-4 Turbo (LAB-Bench snapshot not reported)0.16lab-bench-dbqa-mouse-tumor-gene-sets-paper-v3-zero-shot-cot-no-tools
Claude 3 Haiku (claude-3-haiku-20240307)0.48lab-bench-dbqa-mouse-tumor-gene-sets-paper-v3-zero-shot-cot-no-tools
Meta-Llama-3-70B-Instruct (Anyscale API)0.53lab-bench-dbqa-mouse-tumor-gene-sets-paper-v3-zero-shot-cot-no-tools

Precision (selective accuracy)

lab-bench-dbqa-mouse-tumor-gene-sets-creator-mcq · lab-bench-dbqa-mouse-tumor-gene-sets-paper-v3-zero-shot-cot-no-tools

CSV ↓
Accessible data table
ModelValueComparability group
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)0.73lab-bench-dbqa-mouse-tumor-gene-sets-paper-v3-zero-shot-cot-no-tools
Claude 3 Opus (claude-3-opus-20240229)0.71lab-bench-dbqa-mouse-tumor-gene-sets-paper-v3-zero-shot-cot-no-tools
Gemini 1.5 Pro (gemini-1.5-pro-001)0.76lab-bench-dbqa-mouse-tumor-gene-sets-paper-v3-zero-shot-cot-no-tools
GPT-4o (LAB-Bench snapshot not reported)0.59lab-bench-dbqa-mouse-tumor-gene-sets-paper-v3-zero-shot-cot-no-tools
GPT-4 Turbo (LAB-Bench snapshot not reported)0.58lab-bench-dbqa-mouse-tumor-gene-sets-paper-v3-zero-shot-cot-no-tools
Claude 3 Haiku (claude-3-haiku-20240307)0.55lab-bench-dbqa-mouse-tumor-gene-sets-paper-v3-zero-shot-cot-no-tools
Meta-Llama-3-70B-Instruct (Anyscale API)0.58lab-bench-dbqa-mouse-tumor-gene-sets-paper-v3-zero-shot-cot-no-tools

Coverage

lab-bench-dbqa-mouse-tumor-gene-sets-creator-mcq · lab-bench-dbqa-mouse-tumor-gene-sets-paper-v3-zero-shot-cot-no-tools

CSV ↓
Accessible data table
ModelValueComparability group
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)0.76lab-bench-dbqa-mouse-tumor-gene-sets-paper-v3-zero-shot-cot-no-tools
Claude 3 Opus (claude-3-opus-20240229)0.51lab-bench-dbqa-mouse-tumor-gene-sets-paper-v3-zero-shot-cot-no-tools
Gemini 1.5 Pro (gemini-1.5-pro-001)0.14lab-bench-dbqa-mouse-tumor-gene-sets-paper-v3-zero-shot-cot-no-tools
GPT-4o (LAB-Bench snapshot not reported)0.96lab-bench-dbqa-mouse-tumor-gene-sets-paper-v3-zero-shot-cot-no-tools
GPT-4 Turbo (LAB-Bench snapshot not reported)0.28lab-bench-dbqa-mouse-tumor-gene-sets-paper-v3-zero-shot-cot-no-tools
Claude 3 Haiku (claude-3-haiku-20240307)0.87lab-bench-dbqa-mouse-tumor-gene-sets-paper-v3-zero-shot-cot-no-tools
Meta-Llama-3-70B-Instruct (Anyscale API)0.93lab-bench-dbqa-mouse-tumor-gene-sets-paper-v3-zero-shot-cot-no-tools

Accuracy

lab-bench-dbqa-oncogenic-signatures-creator-mcq · lab-bench-dbqa-oncogenic-signatures-paper-v3-zero-shot-cot-no-tools

CSV ↓
Accessible data table
ModelValueComparability group
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)0.14lab-bench-dbqa-oncogenic-signatures-paper-v3-zero-shot-cot-no-tools
Claude 3 Opus (claude-3-opus-20240229)0.12lab-bench-dbqa-oncogenic-signatures-paper-v3-zero-shot-cot-no-tools
Gemini 1.5 Pro (gemini-1.5-pro-001)0.04lab-bench-dbqa-oncogenic-signatures-paper-v3-zero-shot-cot-no-tools
GPT-4o (LAB-Bench snapshot not reported)0.25lab-bench-dbqa-oncogenic-signatures-paper-v3-zero-shot-cot-no-tools
GPT-4 Turbo (LAB-Bench snapshot not reported)0.06lab-bench-dbqa-oncogenic-signatures-paper-v3-zero-shot-cot-no-tools
Claude 3 Haiku (claude-3-haiku-20240307)0.24lab-bench-dbqa-oncogenic-signatures-paper-v3-zero-shot-cot-no-tools
Meta-Llama-3-70B-Instruct (Anyscale API)0.2lab-bench-dbqa-oncogenic-signatures-paper-v3-zero-shot-cot-no-tools

Precision (selective accuracy)

lab-bench-dbqa-oncogenic-signatures-creator-mcq · lab-bench-dbqa-oncogenic-signatures-paper-v3-zero-shot-cot-no-tools

CSV ↓
Accessible data table
ModelValueComparability group
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)0.59lab-bench-dbqa-oncogenic-signatures-paper-v3-zero-shot-cot-no-tools
Claude 3 Opus (claude-3-opus-20240229)0.57lab-bench-dbqa-oncogenic-signatures-paper-v3-zero-shot-cot-no-tools
Gemini 1.5 Pro (gemini-1.5-pro-001)0.75lab-bench-dbqa-oncogenic-signatures-paper-v3-zero-shot-cot-no-tools
GPT-4o (LAB-Bench snapshot not reported)0.32lab-bench-dbqa-oncogenic-signatures-paper-v3-zero-shot-cot-no-tools
GPT-4 Turbo (LAB-Bench snapshot not reported)0.61lab-bench-dbqa-oncogenic-signatures-paper-v3-zero-shot-cot-no-tools
Claude 3 Haiku (claude-3-haiku-20240307)0.3lab-bench-dbqa-oncogenic-signatures-paper-v3-zero-shot-cot-no-tools
Meta-Llama-3-70B-Instruct (Anyscale API)0.47lab-bench-dbqa-oncogenic-signatures-paper-v3-zero-shot-cot-no-tools

Coverage

lab-bench-dbqa-oncogenic-signatures-creator-mcq · lab-bench-dbqa-oncogenic-signatures-paper-v3-zero-shot-cot-no-tools

CSV ↓
Accessible data table
ModelValueComparability group
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)0.24lab-bench-dbqa-oncogenic-signatures-paper-v3-zero-shot-cot-no-tools
Claude 3 Opus (claude-3-opus-20240229)0.2lab-bench-dbqa-oncogenic-signatures-paper-v3-zero-shot-cot-no-tools
Gemini 1.5 Pro (gemini-1.5-pro-001)0.05lab-bench-dbqa-oncogenic-signatures-paper-v3-zero-shot-cot-no-tools
GPT-4o (LAB-Bench snapshot not reported)0.79lab-bench-dbqa-oncogenic-signatures-paper-v3-zero-shot-cot-no-tools
GPT-4 Turbo (LAB-Bench snapshot not reported)0.1lab-bench-dbqa-oncogenic-signatures-paper-v3-zero-shot-cot-no-tools
Claude 3 Haiku (claude-3-haiku-20240307)0.8lab-bench-dbqa-oncogenic-signatures-paper-v3-zero-shot-cot-no-tools
Meta-Llama-3-70B-Instruct (Anyscale API)0.42lab-bench-dbqa-oncogenic-signatures-paper-v3-zero-shot-cot-no-tools

Accuracy

lab-bench-dbqa-tfbs-gtrd-creator-mcq · lab-bench-dbqa-tfbs-gtrd-paper-v3-zero-shot-cot-no-tools

CSV ↓
Accessible data table
ModelValueComparability group
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)0lab-bench-dbqa-tfbs-gtrd-paper-v3-zero-shot-cot-no-tools
Claude 3 Opus (claude-3-opus-20240229)0lab-bench-dbqa-tfbs-gtrd-paper-v3-zero-shot-cot-no-tools
Gemini 1.5 Pro (gemini-1.5-pro-001)0lab-bench-dbqa-tfbs-gtrd-paper-v3-zero-shot-cot-no-tools
GPT-4o (LAB-Bench snapshot not reported)0.01lab-bench-dbqa-tfbs-gtrd-paper-v3-zero-shot-cot-no-tools
GPT-4 Turbo (LAB-Bench snapshot not reported)0lab-bench-dbqa-tfbs-gtrd-paper-v3-zero-shot-cot-no-tools
Claude 3 Haiku (claude-3-haiku-20240307)0.11lab-bench-dbqa-tfbs-gtrd-paper-v3-zero-shot-cot-no-tools
Meta-Llama-3-70B-Instruct (Anyscale API)0.19lab-bench-dbqa-tfbs-gtrd-paper-v3-zero-shot-cot-no-tools

Precision (selective accuracy)

lab-bench-dbqa-tfbs-gtrd-creator-mcq · lab-bench-dbqa-tfbs-gtrd-paper-v3-zero-shot-cot-no-tools

CSV ↓
Accessible data table
ModelValueComparability group
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)0lab-bench-dbqa-tfbs-gtrd-paper-v3-zero-shot-cot-no-tools
Claude 3 Opus (claude-3-opus-20240229)0lab-bench-dbqa-tfbs-gtrd-paper-v3-zero-shot-cot-no-tools
Gemini 1.5 Pro (gemini-1.5-pro-001)0lab-bench-dbqa-tfbs-gtrd-paper-v3-zero-shot-cot-no-tools
GPT-4o (LAB-Bench snapshot not reported)0.33lab-bench-dbqa-tfbs-gtrd-paper-v3-zero-shot-cot-no-tools
GPT-4 Turbo (LAB-Bench snapshot not reported)0lab-bench-dbqa-tfbs-gtrd-paper-v3-zero-shot-cot-no-tools
Claude 3 Haiku (claude-3-haiku-20240307)0.31lab-bench-dbqa-tfbs-gtrd-paper-v3-zero-shot-cot-no-tools
Meta-Llama-3-70B-Instruct (Anyscale API)0.19lab-bench-dbqa-tfbs-gtrd-paper-v3-zero-shot-cot-no-tools

Coverage

lab-bench-dbqa-tfbs-gtrd-creator-mcq · lab-bench-dbqa-tfbs-gtrd-paper-v3-zero-shot-cot-no-tools

CSV ↓
Accessible data table
ModelValueComparability group
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)0lab-bench-dbqa-tfbs-gtrd-paper-v3-zero-shot-cot-no-tools
Claude 3 Opus (claude-3-opus-20240229)0lab-bench-dbqa-tfbs-gtrd-paper-v3-zero-shot-cot-no-tools
Gemini 1.5 Pro (gemini-1.5-pro-001)0.01lab-bench-dbqa-tfbs-gtrd-paper-v3-zero-shot-cot-no-tools
GPT-4o (LAB-Bench snapshot not reported)0.01lab-bench-dbqa-tfbs-gtrd-paper-v3-zero-shot-cot-no-tools
GPT-4 Turbo (LAB-Bench snapshot not reported)0.11lab-bench-dbqa-tfbs-gtrd-paper-v3-zero-shot-cot-no-tools
Claude 3 Haiku (claude-3-haiku-20240307)0.35lab-bench-dbqa-tfbs-gtrd-paper-v3-zero-shot-cot-no-tools
Meta-Llama-3-70B-Instruct (Anyscale API)1lab-bench-dbqa-tfbs-gtrd-paper-v3-zero-shot-cot-no-tools

Accuracy

lab-bench-dbqa-variant-from-sequence-creator-mcq · lab-bench-dbqa-variant-from-sequence-paper-v3-zero-shot-cot-no-tools

CSV ↓
Accessible data table
ModelValueComparability group
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)0.02lab-bench-dbqa-variant-from-sequence-paper-v3-zero-shot-cot-no-tools
Claude 3 Opus (claude-3-opus-20240229)0lab-bench-dbqa-variant-from-sequence-paper-v3-zero-shot-cot-no-tools
Gemini 1.5 Pro (gemini-1.5-pro-001)0lab-bench-dbqa-variant-from-sequence-paper-v3-zero-shot-cot-no-tools
GPT-4o (LAB-Bench snapshot not reported)0.42lab-bench-dbqa-variant-from-sequence-paper-v3-zero-shot-cot-no-tools
GPT-4 Turbo (LAB-Bench snapshot not reported)0lab-bench-dbqa-variant-from-sequence-paper-v3-zero-shot-cot-no-tools
Claude 3 Haiku (claude-3-haiku-20240307)0.22lab-bench-dbqa-variant-from-sequence-paper-v3-zero-shot-cot-no-tools
Meta-Llama-3-70B-Instruct (Anyscale API)0.3lab-bench-dbqa-variant-from-sequence-paper-v3-zero-shot-cot-no-tools

Precision (selective accuracy)

lab-bench-dbqa-variant-from-sequence-creator-mcq · lab-bench-dbqa-variant-from-sequence-paper-v3-zero-shot-cot-no-tools

CSV ↓
Accessible data table
ModelValueComparability group
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)0.75lab-bench-dbqa-variant-from-sequence-paper-v3-zero-shot-cot-no-tools
Claude 3 Opus (claude-3-opus-20240229)0.11lab-bench-dbqa-variant-from-sequence-paper-v3-zero-shot-cot-no-tools
Gemini 1.5 Pro (gemini-1.5-pro-001)0.17lab-bench-dbqa-variant-from-sequence-paper-v3-zero-shot-cot-no-tools
GPT-4o (LAB-Bench snapshot not reported)0.44lab-bench-dbqa-variant-from-sequence-paper-v3-zero-shot-cot-no-tools
GPT-4 Turbo (LAB-Bench snapshot not reported)0lab-bench-dbqa-variant-from-sequence-paper-v3-zero-shot-cot-no-tools
Claude 3 Haiku (claude-3-haiku-20240307)0.36lab-bench-dbqa-variant-from-sequence-paper-v3-zero-shot-cot-no-tools
Meta-Llama-3-70B-Instruct (Anyscale API)0.35lab-bench-dbqa-variant-from-sequence-paper-v3-zero-shot-cot-no-tools

Coverage

lab-bench-dbqa-variant-from-sequence-creator-mcq · lab-bench-dbqa-variant-from-sequence-paper-v3-zero-shot-cot-no-tools

CSV ↓
Accessible data table
ModelValueComparability group
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)0.03lab-bench-dbqa-variant-from-sequence-paper-v3-zero-shot-cot-no-tools
Claude 3 Opus (claude-3-opus-20240229)0.03lab-bench-dbqa-variant-from-sequence-paper-v3-zero-shot-cot-no-tools
Gemini 1.5 Pro (gemini-1.5-pro-001)0.01lab-bench-dbqa-variant-from-sequence-paper-v3-zero-shot-cot-no-tools
GPT-4o (LAB-Bench snapshot not reported)0.95lab-bench-dbqa-variant-from-sequence-paper-v3-zero-shot-cot-no-tools
GPT-4 Turbo (LAB-Bench snapshot not reported)0.01lab-bench-dbqa-variant-from-sequence-paper-v3-zero-shot-cot-no-tools
Claude 3 Haiku (claude-3-haiku-20240307)0.62lab-bench-dbqa-variant-from-sequence-paper-v3-zero-shot-cot-no-tools
Meta-Llama-3-70B-Instruct (Anyscale API)0.85lab-bench-dbqa-variant-from-sequence-paper-v3-zero-shot-cot-no-tools

Accuracy

lab-bench-dbqa-variant-multi-sequence-creator-mcq · lab-bench-dbqa-variant-multi-sequence-paper-v3-zero-shot-cot-no-tools

CSV ↓
Accessible data table
ModelValueComparability group
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)0.09lab-bench-dbqa-variant-multi-sequence-paper-v3-zero-shot-cot-no-tools
Claude 3 Opus (claude-3-opus-20240229)0.05lab-bench-dbqa-variant-multi-sequence-paper-v3-zero-shot-cot-no-tools
Gemini 1.5 Pro (gemini-1.5-pro-001)0lab-bench-dbqa-variant-multi-sequence-paper-v3-zero-shot-cot-no-tools
GPT-4o (LAB-Bench snapshot not reported)0.06lab-bench-dbqa-variant-multi-sequence-paper-v3-zero-shot-cot-no-tools
GPT-4 Turbo (LAB-Bench snapshot not reported)0lab-bench-dbqa-variant-multi-sequence-paper-v3-zero-shot-cot-no-tools
Claude 3 Haiku (claude-3-haiku-20240307)0.14lab-bench-dbqa-variant-multi-sequence-paper-v3-zero-shot-cot-no-tools
Meta-Llama-3-70B-Instruct (Anyscale API)0.07lab-bench-dbqa-variant-multi-sequence-paper-v3-zero-shot-cot-no-tools

Precision (selective accuracy)

lab-bench-dbqa-variant-multi-sequence-creator-mcq · lab-bench-dbqa-variant-multi-sequence-paper-v3-zero-shot-cot-no-tools

CSV ↓
Accessible data table
ModelValueComparability group
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)0.16lab-bench-dbqa-variant-multi-sequence-paper-v3-zero-shot-cot-no-tools
Claude 3 Opus (claude-3-opus-20240229)0.18lab-bench-dbqa-variant-multi-sequence-paper-v3-zero-shot-cot-no-tools
Gemini 1.5 Pro (gemini-1.5-pro-001)0.33lab-bench-dbqa-variant-multi-sequence-paper-v3-zero-shot-cot-no-tools
GPT-4o (LAB-Bench snapshot not reported)0.12lab-bench-dbqa-variant-multi-sequence-paper-v3-zero-shot-cot-no-tools
GPT-4 Turbo (LAB-Bench snapshot not reported)0lab-bench-dbqa-variant-multi-sequence-paper-v3-zero-shot-cot-no-tools
Claude 3 Haiku (claude-3-haiku-20240307)0.16lab-bench-dbqa-variant-multi-sequence-paper-v3-zero-shot-cot-no-tools
Meta-Llama-3-70B-Instruct (Anyscale API)0.12lab-bench-dbqa-variant-multi-sequence-paper-v3-zero-shot-cot-no-tools

Coverage

lab-bench-dbqa-variant-multi-sequence-creator-mcq · lab-bench-dbqa-variant-multi-sequence-paper-v3-zero-shot-cot-no-tools

CSV ↓
Accessible data table
ModelValueComparability group
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)0.56lab-bench-dbqa-variant-multi-sequence-paper-v3-zero-shot-cot-no-tools
Claude 3 Opus (claude-3-opus-20240229)0.26lab-bench-dbqa-variant-multi-sequence-paper-v3-zero-shot-cot-no-tools
Gemini 1.5 Pro (gemini-1.5-pro-001)0.04lab-bench-dbqa-variant-multi-sequence-paper-v3-zero-shot-cot-no-tools
GPT-4o (LAB-Bench snapshot not reported)0.49lab-bench-dbqa-variant-multi-sequence-paper-v3-zero-shot-cot-no-tools
GPT-4 Turbo (LAB-Bench snapshot not reported)0.02lab-bench-dbqa-variant-multi-sequence-paper-v3-zero-shot-cot-no-tools
Claude 3 Haiku (claude-3-haiku-20240307)0.84lab-bench-dbqa-variant-multi-sequence-paper-v3-zero-shot-cot-no-tools
Meta-Llama-3-70B-Instruct (Anyscale API)0.63lab-bench-dbqa-variant-multi-sequence-paper-v3-zero-shot-cot-no-tools

Accuracy

lab-bench-dbqa-vax-response-creator-mcq · lab-bench-dbqa-vax-response-paper-v3-zero-shot-cot-no-tools

CSV ↓
Accessible data table
ModelValueComparability group
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)0.21lab-bench-dbqa-vax-response-paper-v3-zero-shot-cot-no-tools
Claude 3 Opus (claude-3-opus-20240229)0.14lab-bench-dbqa-vax-response-paper-v3-zero-shot-cot-no-tools
Gemini 1.5 Pro (gemini-1.5-pro-001)0.13lab-bench-dbqa-vax-response-paper-v3-zero-shot-cot-no-tools
GPT-4o (LAB-Bench snapshot not reported)0.33lab-bench-dbqa-vax-response-paper-v3-zero-shot-cot-no-tools
GPT-4 Turbo (LAB-Bench snapshot not reported)0.06lab-bench-dbqa-vax-response-paper-v3-zero-shot-cot-no-tools
Claude 3 Haiku (claude-3-haiku-20240307)0.32lab-bench-dbqa-vax-response-paper-v3-zero-shot-cot-no-tools
Meta-Llama-3-70B-Instruct (Anyscale API)0.15lab-bench-dbqa-vax-response-paper-v3-zero-shot-cot-no-tools

Precision (selective accuracy)

lab-bench-dbqa-vax-response-creator-mcq · lab-bench-dbqa-vax-response-paper-v3-zero-shot-cot-no-tools

CSV ↓
Accessible data table
ModelValueComparability group
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)0.57lab-bench-dbqa-vax-response-paper-v3-zero-shot-cot-no-tools
Claude 3 Opus (claude-3-opus-20240229)0.65lab-bench-dbqa-vax-response-paper-v3-zero-shot-cot-no-tools
Gemini 1.5 Pro (gemini-1.5-pro-001)0.61lab-bench-dbqa-vax-response-paper-v3-zero-shot-cot-no-tools
GPT-4o (LAB-Bench snapshot not reported)0.45lab-bench-dbqa-vax-response-paper-v3-zero-shot-cot-no-tools
GPT-4 Turbo (LAB-Bench snapshot not reported)0.65lab-bench-dbqa-vax-response-paper-v3-zero-shot-cot-no-tools
Claude 3 Haiku (claude-3-haiku-20240307)0.38lab-bench-dbqa-vax-response-paper-v3-zero-shot-cot-no-tools
Meta-Llama-3-70B-Instruct (Anyscale API)0.54lab-bench-dbqa-vax-response-paper-v3-zero-shot-cot-no-tools

Coverage

lab-bench-dbqa-vax-response-creator-mcq · lab-bench-dbqa-vax-response-paper-v3-zero-shot-cot-no-tools

CSV ↓
Accessible data table
ModelValueComparability group
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)0.36lab-bench-dbqa-vax-response-paper-v3-zero-shot-cot-no-tools
Claude 3 Opus (claude-3-opus-20240229)0.21lab-bench-dbqa-vax-response-paper-v3-zero-shot-cot-no-tools
Gemini 1.5 Pro (gemini-1.5-pro-001)0.21lab-bench-dbqa-vax-response-paper-v3-zero-shot-cot-no-tools
GPT-4o (LAB-Bench snapshot not reported)0.73lab-bench-dbqa-vax-response-paper-v3-zero-shot-cot-no-tools
GPT-4 Turbo (LAB-Bench snapshot not reported)0.09lab-bench-dbqa-vax-response-paper-v3-zero-shot-cot-no-tools
Claude 3 Haiku (claude-3-haiku-20240307)0.83lab-bench-dbqa-vax-response-paper-v3-zero-shot-cot-no-tools
Meta-Llama-3-70B-Instruct (Anyscale API)0.29lab-bench-dbqa-vax-response-paper-v3-zero-shot-cot-no-tools

Accuracy

lab-bench-dbqa-viral-ppi-creator-mcq · lab-bench-dbqa-viral-ppi-paper-v3-zero-shot-cot-no-tools

CSV ↓
Accessible data table
ModelValueComparability group
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)0lab-bench-dbqa-viral-ppi-paper-v3-zero-shot-cot-no-tools
Claude 3 Opus (claude-3-opus-20240229)0lab-bench-dbqa-viral-ppi-paper-v3-zero-shot-cot-no-tools
Gemini 1.5 Pro (gemini-1.5-pro-001)0lab-bench-dbqa-viral-ppi-paper-v3-zero-shot-cot-no-tools
GPT-4o (LAB-Bench snapshot not reported)0.26lab-bench-dbqa-viral-ppi-paper-v3-zero-shot-cot-no-tools
GPT-4 Turbo (LAB-Bench snapshot not reported)0.02lab-bench-dbqa-viral-ppi-paper-v3-zero-shot-cot-no-tools
Claude 3 Haiku (claude-3-haiku-20240307)0.27lab-bench-dbqa-viral-ppi-paper-v3-zero-shot-cot-no-tools
Meta-Llama-3-70B-Instruct (Anyscale API)0.4lab-bench-dbqa-viral-ppi-paper-v3-zero-shot-cot-no-tools

Precision (selective accuracy)

lab-bench-dbqa-viral-ppi-creator-mcq · lab-bench-dbqa-viral-ppi-paper-v3-zero-shot-cot-no-tools

CSV ↓
Accessible data table
ModelValueComparability group
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)0lab-bench-dbqa-viral-ppi-paper-v3-zero-shot-cot-no-tools
Claude 3 Opus (claude-3-opus-20240229)0lab-bench-dbqa-viral-ppi-paper-v3-zero-shot-cot-no-tools
Gemini 1.5 Pro (gemini-1.5-pro-001)0lab-bench-dbqa-viral-ppi-paper-v3-zero-shot-cot-no-tools
GPT-4o (LAB-Bench snapshot not reported)0.69lab-bench-dbqa-viral-ppi-paper-v3-zero-shot-cot-no-tools
GPT-4 Turbo (LAB-Bench snapshot not reported)0.28lab-bench-dbqa-viral-ppi-paper-v3-zero-shot-cot-no-tools
Claude 3 Haiku (claude-3-haiku-20240307)0.49lab-bench-dbqa-viral-ppi-paper-v3-zero-shot-cot-no-tools
Meta-Llama-3-70B-Instruct (Anyscale API)0.4lab-bench-dbqa-viral-ppi-paper-v3-zero-shot-cot-no-tools

Coverage

lab-bench-dbqa-viral-ppi-creator-mcq · lab-bench-dbqa-viral-ppi-paper-v3-zero-shot-cot-no-tools

CSV ↓
Accessible data table
ModelValueComparability group
Claude 3.5 Sonnet (claude-3-5-sonnet-20240620)0lab-bench-dbqa-viral-ppi-paper-v3-zero-shot-cot-no-tools
Claude 3 Opus (claude-3-opus-20240229)0lab-bench-dbqa-viral-ppi-paper-v3-zero-shot-cot-no-tools
Gemini 1.5 Pro (gemini-1.5-pro-001)0lab-bench-dbqa-viral-ppi-paper-v3-zero-shot-cot-no-tools
GPT-4o (LAB-Bench snapshot not reported)0.38lab-bench-dbqa-viral-ppi-paper-v3-zero-shot-cot-no-tools
GPT-4 Turbo (LAB-Bench snapshot not reported)0.06lab-bench-dbqa-viral-ppi-paper-v3-zero-shot-cot-no-tools
Claude 3 Haiku (claude-3-haiku-20240307)0.55lab-bench-dbqa-viral-ppi-paper-v3-zero-shot-cot-no-tools
Meta-Llama-3-70B-Instruct (Anyscale API)1lab-bench-dbqa-viral-ppi-paper-v3-zero-shot-cot-no-tools

Evidence and change history

Source locators remain visible; expand an item to inspect the exact Registry fields it supports.

LAB-Bench: Measuring Capabilities of Language Models for Biology Research · table: Table 1 and Appendix Table 6; task-specific Appendix C sections (Category total, task definitions, domains, modalities, and creator evaluation.) · Supports 15 fields

Open source →

  • /name
  • /organizations
  • /release_date
  • /kind
  • /domains
  • /capabilities
  • /modalities
  • /task_counts/total
  • /task_counts/basis
  • /task_counts/subsets
  • /access/level
  • /access/license
  • /versions/0/task_counts/total
  • /versions/0/task_counts/basis
  • /versions/0/task_counts/subsets
lab-bench-dbqa-repository-resource · repository-path: DbQA/*-splits.json and task.py at 998a8e0a40cf116c80e1b0e7a805ebb5fb9fa838 (Public/private/total 520/130/650; exact child files and evaluator.) · Supports 14 fields

Open source →

  • /latest_version
  • /task_counts/total
  • /task_counts/basis
  • /task_counts/subsets
  • /access/level
  • /access/tasks
  • /access/artifacts
  • /access/grader
  • /access/license
  • /resources
  • /versions/1/task_counts/total
  • /versions/1/task_counts/basis
  • /versions/1/task_counts/subsets
  • /scientific_task_classification/entries/0

View source-level modification history on GitHub →