preprint · benchmark creator

BioSecBench-Surveillance: A Verifiable Benchmark for AI Agents in Pathogen Genomic Surveillance

LatchBio · Aclid · 2026-07-21

Relationship layer

Benchmark usage

This table records what the work did with each benchmark before attempting to normalize a run. Partial claims remain visible without being treated as comparable evaluations.

benchmark creation

biosecbench-surveillance-a-verifiable-benchmark-for-ai-biosecbench-surveillance-1-use

Non-evaluationunknown

Benchmark: BioSecBench-Surveillance · version Not reported

Selection
not applicable
Models
Not reported / not applicable
Metrics
Not reported / not applicable
Linked runs
None

Not reported / unresolved: Benchmark version is not reported.; The exact public-subset size is not reported.; The repository license is not reported.

AI-assisted double-pass extraction; values are limited to independently supported claims.

Evidence
  • section: Abstract
    Supports: /relation_type
  • section: Abstract
    Supports: /benchmark_id

evaluation

biosecbench-surveillance-a-verifiable-benchmark-for-ai-biosecbench-surveillance-2-use

Partialunknown · n=100

Benchmark: BioSecBench-Surveillance · version Not reported

Selection
not reported
Metrics
Endpoint pass rate
Linked runs
None

Not reported / unresolved: Benchmark version is not reported.; Exact deployment snapshots and model release dates are not reported.; Prompt text, shots, token budget, and random seed are not reported.; Gradable evaluation n is reported only for Opus 4.8 / PI.; Exact confidence intervals are not numerically printed for seven PI results.; benchmark version; numeric result

AI-assisted double-pass extraction; values are limited to independently supported claims.

Evidence
  • figure: Figure 2
    Supports: /relation_type
  • figure: Figure 2
    Supports: /benchmark_id
  • section: Methods — Agent runs and execution
    Supports: /scope
  • section: Methods — Agent runs and execution
    Supports: /scope
  • section: Methods — Outcome classification and aggregation
    Supports: /metric_labels
  • figure: Figure 2
    Supports: /model_ids
  • figure: Figure 2
    Supports: /model_ids
  • figure: Figure 2
    Supports: /model_ids
  • figure: Figure 2
    Supports: /model_ids
  • figure: Figure 2
    Supports: /model_ids
  • figure: Figure 2
    Supports: /model_ids
  • figure: Figure 2
    Supports: /model_ids
  • figure: Figure 2
    Supports: /model_ids
  • figure: Figure 2
    Supports: /model_ids
  • figure: Figure 2
    Supports: /model_ids

evaluation

biosecbench-surveillance-a-verifiable-benchmark-for-ai-biosecbench-surveillance-3-use

Partialunknown · n=100

Benchmark: BioSecBench-Surveillance · version Not reported

Selection
not reported
Metrics
Endpoint pass rate
Linked runs
None

Not reported / unresolved: Benchmark version is not reported.; Exact deployment snapshots and model release dates are not reported.; Prompt text, shots, token budget, and random seed are not reported.; Per-configuration gradable n and numerical confidence intervals are not reported.; benchmark version; numeric result

AI-assisted double-pass extraction; values are limited to independently supported claims.

Evidence
  • figure: Figure 2
    Supports: /relation_type
  • figure: Figure 2
    Supports: /benchmark_id
  • section: Methods — Agent runs and execution
    Supports: /scope
  • section: Methods — Agent runs and execution
    Supports: /scope
  • section: Methods — Outcome classification and aggregation
    Supports: /metric_labels
  • figure: Figure 2
    Supports: /model_ids
  • figure: Figure 2
    Supports: /model_ids
  • figure: Figure 2
    Supports: /model_ids
  • figure: Figure 2
    Supports: /model_ids

evaluation

biosecbench-surveillance-a-verifiable-benchmark-for-ai-biosecbench-surveillance-4-use

Partialunknown · n=100

Benchmark: BioSecBench-Surveillance · version Not reported

Selection
not reported
Metrics
Endpoint pass rate
Linked runs
None

Not reported / unresolved: Benchmark version is not reported.; Exact deployment snapshots and model release dates are not reported.; Prompt text, shots, token budget, and random seed are not reported.; Gradable evaluation n is not reported for either Codex configuration.; GPT-5.4 / Codex has no numerically printed confidence interval.; benchmark version; numeric result

AI-assisted double-pass extraction; values are limited to independently supported claims.

Evidence
  • figure: Figure 2
    Supports: /relation_type
  • figure: Figure 2
    Supports: /benchmark_id
  • section: Methods — Agent runs and execution
    Supports: /scope
  • section: Methods — Agent runs and execution
    Supports: /scope
  • section: Methods — Outcome classification and aggregation
    Supports: /metric_labels
  • figure: Figure 2
    Supports: /model_ids
  • figure: Figure 2
    Supports: /model_ids

Normalized evaluation runs

This source has no normalized model run. It may be a creator-only source or a partial/non-evaluation benchmark use.