official-release · benchmark creator

BioSecBench-Surveillance repository result snapshot

LatchBio · 2026-07-09

Relationship layer

Benchmark usage

This table records what the work did with each benchmark before attempting to normalize a run. Partial claims remain visible without being treated as comparable evaluations.

evaluation

biosecbench-8d53fd8-claude-code-use

Normalizedfull · n=100

Benchmark: BioSecBench-Surveillance · version initial-release

Selection
not applicable · All 100 registered evaluations
Metrics
endpoint pass rate

Official commit-pinned Claude Code evaluation result snapshot.

Evidence
  • repository-path: METHODS.md and results/config_results.csv at commit 8d53fd8517cc74202eb18b618e8b39b4ffaf0c87 (Full-scope Claude Code evaluation relation, models, metric, and linked run.)
    Supports: /benchmark_version, /relation_type, /status, /model_ids, /scope, /metric_labels, /evaluation_run_ids, /notes

evaluation

biosecbench-8d53fd8-openai-codex-use

Normalizedfull · n=100

Benchmark: BioSecBench-Surveillance · version initial-release

Selection
not applicable · All 100 registered evaluations
Metrics
endpoint pass rate

Official commit-pinned OpenAI Codex evaluation result snapshot.

Evidence
  • repository-path: METHODS.md and results/config_results.csv at commit 8d53fd8517cc74202eb18b618e8b39b4ffaf0c87 (Full-scope OpenAI Codex evaluation relation, models, metric, and linked run.)
    Supports: /benchmark_version, /relation_type, /status, /model_ids, /scope, /metric_labels, /evaluation_run_ids, /notes

evaluation

biosecbench-8d53fd8-pi-use

Normalizedfull · n=100

Benchmark: BioSecBench-Surveillance · version initial-release

Selection
not applicable · All 100 registered evaluations
Metrics
endpoint pass rate

Official commit-pinned PI evaluation result snapshot.

Evidence
  • repository-path: METHODS.md and results/config_results.csv at commit 8d53fd8517cc74202eb18b618e8b39b4ffaf0c87 (Full-scope PI evaluation relation, models, metric, and linked run.)
    Supports: /benchmark_version, /relation_type, /status, /model_ids, /scope, /metric_labels, /evaluation_run_ids, /notes

Normalized evaluation runs

BioSecBench-Surveillance3 runs

Open benchmark record →

Evaluation run

biosecbench-8d53fd8-claude-code

From BioSecBench-Surveillance repository result snapshot

biosecbench-8d53fd8-claude-codevinitial-release

Evaluated models / systems: Opus 4.6, Opus 4.7, Opus 4.8, Sonnet 4.6

Scopefull · n=100
ShotsNot reported
TurnsNot reported
System prompt publicNot reported
Reasoning / effortNot reported
BrowserNot reported
InternetYes
Databasesresistance databases, virulence databases
Code executionNot reported
Containeridentical containerized sandbox
External toolsClaude Code, assemblers, aligners, taxonomic classifiers, variant callers
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
Repeats3
Graderdeterministic typed-field checks · human review: not reported
Statistics95% Student-t confidence interval over evaluations
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
endpoint pass rateabsolute%mean of per-evaluation pass rates; evaluations with no gradeable run are droppedNot reported

Results

ModelMetricValuen
Opus 4.6endpoint pass rate46.7 %
Claude Code; n is the number of gradable evaluations after evaluations with no gradable run were dropped; the configured benchmark scope was all 100 evaluations.
80
Opus 4.7endpoint pass rate44.2 %
Claude Code; n is the number of gradable evaluations after evaluations with no gradable run were dropped; the configured benchmark scope was all 100 evaluations.
77
Sonnet 4.6endpoint pass rate43.6 %
Claude Code; n is the number of gradable evaluations after evaluations with no gradable run were dropped; the configured benchmark scope was all 100 evaluations.
83
Opus 4.8endpoint pass rate39.5 %
Claude Code; n is the number of gradable evaluations after evaluations with no gradable run were dropped; the configured benchmark scope was all 100 evaluations.
76

Evidence

  • repository-path: METHODS.md at commit 8d53fd8517cc74202eb18b618e8b39b4ffaf0c87 (Full 100-evaluation scope, Claude Code harness, tools, three repeats, deterministic grading, aggregation, and confidence intervals.) — supports /benchmark_version, /scope, /protocol, /metrics
  • repository-path: results/config_results.csv; harness=claude-code at commit 8d53fd8517cc74202eb18b618e8b39b4ffaf0c87 (Exact model labels, endpoint pass rates, 95% confidence intervals, and gradable evaluation counts.) — supports /results

Evaluation run

biosecbench-8d53fd8-openai-codex

From BioSecBench-Surveillance repository result snapshot

biosecbench-8d53fd8-openai-codexvinitial-release

Evaluated models / systems: GPT-5.4, GPT-5.5

Scopefull · n=100
ShotsNot reported
TurnsNot reported
System prompt publicNot reported
Reasoning / effortNot reported
BrowserNot reported
InternetYes
Databasesresistance databases, virulence databases
Code executionNot reported
Containeridentical containerized sandbox
External toolsOpenAI Codex, assemblers, aligners, taxonomic classifiers, variant callers
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
Repeats3
Graderdeterministic typed-field checks · human review: not reported
Statistics95% Student-t confidence interval over evaluations
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
endpoint pass rateabsolute%mean of per-evaluation pass rates; evaluations with no gradeable run are droppedNot reported

Results

ModelMetricValuen
GPT-5.5endpoint pass rate50.2 %
OpenAI Codex; n is the number of gradable evaluations after evaluations with no gradable run were dropped; the configured benchmark scope was all 100 evaluations.
93
GPT-5.4endpoint pass rate41.3 %
OpenAI Codex; n is the number of gradable evaluations after evaluations with no gradable run were dropped; the configured benchmark scope was all 100 evaluations.
92

Evidence

  • repository-path: METHODS.md at commit 8d53fd8517cc74202eb18b618e8b39b4ffaf0c87 (Full 100-evaluation scope, OpenAI Codex harness, tools, three repeats, deterministic grading, aggregation, and confidence intervals.) — supports /benchmark_version, /scope, /protocol, /metrics
  • repository-path: results/config_results.csv; harness=openai-codex at commit 8d53fd8517cc74202eb18b618e8b39b4ffaf0c87 (Exact model labels, endpoint pass rates, 95% confidence intervals, and gradable evaluation counts.) — supports /results

Evaluation run

biosecbench-8d53fd8-pi

From BioSecBench-Surveillance repository result snapshot

biosecbench-8d53fd8-pivinitial-release

Evaluated models / systems: Opus 4.6, Opus 4.7, Opus 4.8, Sonnet 4.6, gemini-3.1-pro-preview, Gemini 3.5 Flash, GPT-5.4, GPT-5.5, grok-4.20-beta-0309-reasoning, Grok 4.3

Scopefull · n=100
ShotsNot reported
TurnsNot reported
System prompt publicNot reported
Reasoning / effortNot reported
BrowserNot reported
InternetYes
Databasesresistance databases, virulence databases
Code executionNot reported
Containeridentical containerized sandbox
External toolsPI, assemblers, aligners, taxonomic classifiers, variant callers
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
Repeats3
Graderdeterministic typed-field checks · human review: not reported
Statistics95% Student-t confidence interval over evaluations
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
endpoint pass rateabsolute%mean of per-evaluation pass rates; evaluations with no gradeable run are droppedNot reported

Results

ModelMetricValuen
Opus 4.8endpoint pass rate50.2 %
PI; n is the number of gradable evaluations after evaluations with no gradable run were dropped; the configured benchmark scope was all 100 evaluations.
83
GPT-5.5endpoint pass rate44.8 %
PI; n is the number of gradable evaluations after evaluations with no gradable run were dropped; the configured benchmark scope was all 100 evaluations.
77
Opus 4.7endpoint pass rate49.6 %
PI; n is the number of gradable evaluations after evaluations with no gradable run were dropped; the configured benchmark scope was all 100 evaluations.
78
Sonnet 4.6endpoint pass rate48.6 %
PI; n is the number of gradable evaluations after evaluations with no gradable run were dropped; the configured benchmark scope was all 100 evaluations.
83
Gemini 3.5 Flashendpoint pass rate47.1 %
PI; n is the number of gradable evaluations after evaluations with no gradable run were dropped; the configured benchmark scope was all 100 evaluations.
92
Opus 4.6endpoint pass rate45.7 %
PI; n is the number of gradable evaluations after evaluations with no gradable run were dropped; the configured benchmark scope was all 100 evaluations.
81
gemini-3.1-pro-previewendpoint pass rate45.3 %
PI; n is the number of gradable evaluations after evaluations with no gradable run were dropped; the configured benchmark scope was all 100 evaluations.
95
GPT-5.4endpoint pass rate38.5 %
PI; n is the number of gradable evaluations after evaluations with no gradable run were dropped; the configured benchmark scope was all 100 evaluations.
81
Grok 4.3endpoint pass rate16 %
PI; n is the number of gradable evaluations after evaluations with no gradable run were dropped; the configured benchmark scope was all 100 evaluations.
100
grok-4.20-beta-0309-reasoningendpoint pass rate13.7 %
PI; n is the number of gradable evaluations after evaluations with no gradable run were dropped; the configured benchmark scope was all 100 evaluations.
100

Evidence

  • repository-path: METHODS.md at commit 8d53fd8517cc74202eb18b618e8b39b4ffaf0c87 (Full 100-evaluation scope, PI harness, tools, three repeats, deterministic grading, aggregation, and confidence intervals.) — supports /benchmark_version, /scope, /protocol, /metrics
  • repository-path: results/config_results.csv; harness=pi at commit 8d53fd8517cc74202eb18b618e8b39b4ffaf0c87 (Exact model labels, endpoint pass rates, 95% confidence intervals, and gradable evaluation counts.) — supports /results