agentic-eval · audited-with-caveats · verified 2026-07-30

BioSecBench-Surveillance

Agentic evaluation of pathogen genomic-surveillance workflow selection and analysis from raw or near-raw sequencing data.

Audited with caveats: 1 field(s) are marked provisional or conflicted. Warnings are shown next to affected values and these claims are excluded from unqualified comparisons.

Benchmark definition

What is counted

Version
initial-release
Total
100 (The source explicitly defines the benchmark as 100 evaluations.)
Task formats
agentic sequencing-data analysis with structured JSON answers
Capabilities
Data analysisTool useScientific reasoning
Modalities
TextDNA or RNA sequenceRaw omicsDatabase

Version history

VersionStatusRelease / as-ofTotalFormal tracks
initial-release
biosecbench-surveillance-initial-release-version
current2026-07-21100 (The source explicitly defines the benchmark as 100 evaluations.)None registered

Tracks and subsets

IDCountBasisPartition?Notes
Task category — Variant detection evaluations
task-category-variant-detection
23Printed Figure 1A bar label.Exclusive & exhaustive
Task category — Taxonomic classification evaluations
task-category-taxonomic-classification
16Printed Figure 1A bar label.Exclusive & exhaustive
Task category — AMR characterization evaluations
task-category-amr-characterization
15Printed Figure 1A bar label.Exclusive & exhaustive
Task category — Toxin and virulence characterization evaluations
task-category-toxin-and-virulence-characterization
12Printed Figure 1A bar label.Exclusive & exhaustive
Task category — Genetic-engineering characterization evaluations
task-category-genetic-engineering-characterization
12Printed Figure 1A bar label.Exclusive & exhaustive
Task category — Source tracking evaluations
task-category-source-tracking
12Printed Figure 1A bar label.Exclusive & exhaustive
Task category — Anomaly detection evaluations
task-category-anomaly-detection
10Printed Figure 1A bar label.Exclusive & exhaustive
Sample type — Isolate evaluations
sample-type-isolate
34Printed Figure 1B bar label.Exclusive & exhaustive
Sample type — Wastewater evaluations
sample-type-wastewater
31Printed Figure 1B bar label.Exclusive & exhaustive
Sample type — Clinical evaluations
sample-type-clinical
23Printed Figure 1B bar label.Exclusive & exhaustive
Sample type — Agricultural evaluations
sample-type-agricultural
5Printed Figure 1B bar label.Exclusive & exhaustive
Sample type — Air evaluations
sample-type-air
5Printed Figure 1B bar label.Exclusive & exhaustive
Sample type — Water evaluations
sample-type-water
2Printed Figure 1B bar label.Exclusive & exhaustive
Sequencing technology — Short-read evaluations
sequencing-technology-short-read
80Printed Figure 1C bar label.Exclusive & exhaustive
Sequencing technology — Long-read evaluations
sequencing-technology-long-read
15Printed Figure 1C bar label.Exclusive & exhaustive
Sequencing technology — Hybrid evaluations
sequencing-technology-hybrid
5Printed Figure 1C bar label.Exclusive & exhaustive
Nucleic-acid target — DNA evaluations
nucleic-acid-target-dna
64Printed Figure 1D bar label.Exclusive & exhaustive
Nucleic-acid target — RNA evaluations
nucleic-acid-target-rna
33Printed Figure 1D bar label.Exclusive & exhaustive
Nucleic-acid target — Total NA evaluations
nucleic-acid-target-total-na
3Printed Figure 1D bar label.Exclusive & exhaustive
Assay type — Shotgun evaluations
assay-type-shotgun
85Printed Figure 1E bar label.Exclusive & exhaustive
Assay type — Targeted/amplicon/capture evaluations
assay-type-targeted-amplicon-capture
15Figure 1 defines Targeted as targeted/amplicon/capture.Exclusive & exhaustive

Scientific Task Atlas

Scientific task classification

partial for initial-release. No Scientific Task claim passed independent high-confidence verification; task mapping remains pending a targeted official-source audit.

The source names only a broad direction; no more specific leaf task can be assigned without inference.

Relationship registry

How works use this benchmark

Partial claims, non-evaluation uses, and third-party summaries stay visible without entering model comparisons.

Normalized evaluations

evaluation

biosecbench-8d53fd8-claude-code-use

Normalizedfull · n=100

Work: BioSecBench-Surveillance repository result snapshot · source version biosecbench-surveillance-repository-result-snapshot-8d53fd8

Selection
not applicable · All 100 registered evaluations
Metrics
endpoint pass rate

Official commit-pinned Claude Code evaluation result snapshot.

Evidence
  • repository-path: METHODS.md and results/config_results.csv at commit 8d53fd8517cc74202eb18b618e8b39b4ffaf0c87 (Full-scope Claude Code evaluation relation, models, metric, and linked run.)
    Supports: /benchmark_version, /relation_type, /status, /model_ids, /scope, /metric_labels, /evaluation_run_ids, /notes

evaluation

biosecbench-8d53fd8-openai-codex-use

Normalizedfull · n=100

Work: BioSecBench-Surveillance repository result snapshot · source version biosecbench-surveillance-repository-result-snapshot-8d53fd8

Selection
not applicable · All 100 registered evaluations
Metrics
endpoint pass rate

Official commit-pinned OpenAI Codex evaluation result snapshot.

Evidence
  • repository-path: METHODS.md and results/config_results.csv at commit 8d53fd8517cc74202eb18b618e8b39b4ffaf0c87 (Full-scope OpenAI Codex evaluation relation, models, metric, and linked run.)
    Supports: /benchmark_version, /relation_type, /status, /model_ids, /scope, /metric_labels, /evaluation_run_ids, /notes

evaluation

biosecbench-8d53fd8-pi-use

Normalizedfull · n=100

Work: BioSecBench-Surveillance repository result snapshot · source version biosecbench-surveillance-repository-result-snapshot-8d53fd8

Selection
not applicable · All 100 registered evaluations
Metrics
endpoint pass rate

Official commit-pinned PI evaluation result snapshot.

Evidence
  • repository-path: METHODS.md and results/config_results.csv at commit 8d53fd8517cc74202eb18b618e8b39b4ffaf0c87 (Full-scope PI evaluation relation, models, metric, and linked run.)
    Supports: /benchmark_version, /relation_type, /status, /model_ids, /scope, /metric_labels, /evaluation_run_ids, /notes

Partial evaluation claims

evaluation

biosecbench-surveillance-a-verifiable-benchmark-for-ai-biosecbench-surveillance-2-use

Partialunknown · n=100

Work: BioSecBench-Surveillance: A Verifiable Benchmark for AI Agents in Pathogen Genomic Surveillance · source version biosecbench-surveillance-a-verifiable-benchmark-for-ai-arxiv-v1

Selection
not reported
Metrics
Endpoint pass rate
Linked runs
None

Not reported / unresolved: Benchmark version is not reported.; Exact deployment snapshots and model release dates are not reported.; Prompt text, shots, token budget, and random seed are not reported.; Gradable evaluation n is reported only for Opus 4.8 / PI.; Exact confidence intervals are not numerically printed for seven PI results.; benchmark version; numeric result

AI-assisted double-pass extraction; values are limited to independently supported claims.

Evidence
  • figure: Figure 2
    Supports: /relation_type
  • figure: Figure 2
    Supports: /benchmark_id
  • section: Methods — Agent runs and execution
    Supports: /scope
  • section: Methods — Agent runs and execution
    Supports: /scope
  • section: Methods — Outcome classification and aggregation
    Supports: /metric_labels
  • figure: Figure 2
    Supports: /model_ids
  • figure: Figure 2
    Supports: /model_ids
  • figure: Figure 2
    Supports: /model_ids
  • figure: Figure 2
    Supports: /model_ids
  • figure: Figure 2
    Supports: /model_ids
  • figure: Figure 2
    Supports: /model_ids
  • figure: Figure 2
    Supports: /model_ids
  • figure: Figure 2
    Supports: /model_ids
  • figure: Figure 2
    Supports: /model_ids
  • figure: Figure 2
    Supports: /model_ids

evaluation

biosecbench-surveillance-a-verifiable-benchmark-for-ai-biosecbench-surveillance-3-use

Partialunknown · n=100

Work: BioSecBench-Surveillance: A Verifiable Benchmark for AI Agents in Pathogen Genomic Surveillance · source version biosecbench-surveillance-a-verifiable-benchmark-for-ai-arxiv-v1

Selection
not reported
Metrics
Endpoint pass rate
Linked runs
None

Not reported / unresolved: Benchmark version is not reported.; Exact deployment snapshots and model release dates are not reported.; Prompt text, shots, token budget, and random seed are not reported.; Per-configuration gradable n and numerical confidence intervals are not reported.; benchmark version; numeric result

AI-assisted double-pass extraction; values are limited to independently supported claims.

Evidence
  • figure: Figure 2
    Supports: /relation_type
  • figure: Figure 2
    Supports: /benchmark_id
  • section: Methods — Agent runs and execution
    Supports: /scope
  • section: Methods — Agent runs and execution
    Supports: /scope
  • section: Methods — Outcome classification and aggregation
    Supports: /metric_labels
  • figure: Figure 2
    Supports: /model_ids
  • figure: Figure 2
    Supports: /model_ids
  • figure: Figure 2
    Supports: /model_ids
  • figure: Figure 2
    Supports: /model_ids

evaluation

biosecbench-surveillance-a-verifiable-benchmark-for-ai-biosecbench-surveillance-4-use

Partialunknown · n=100

Work: BioSecBench-Surveillance: A Verifiable Benchmark for AI Agents in Pathogen Genomic Surveillance · source version biosecbench-surveillance-a-verifiable-benchmark-for-ai-arxiv-v1

Selection
not reported
Metrics
Endpoint pass rate
Linked runs
None

Not reported / unresolved: Benchmark version is not reported.; Exact deployment snapshots and model release dates are not reported.; Prompt text, shots, token budget, and random seed are not reported.; Gradable evaluation n is not reported for either Codex configuration.; GPT-5.4 / Codex has no numerically printed confidence interval.; benchmark version; numeric result

AI-assisted double-pass extraction; values are limited to independently supported claims.

Evidence
  • figure: Figure 2
    Supports: /relation_type
  • figure: Figure 2
    Supports: /benchmark_id
  • section: Methods — Agent runs and execution
    Supports: /scope
  • section: Methods — Agent runs and execution
    Supports: /scope
  • section: Methods — Outcome classification and aggregation
    Supports: /metric_labels
  • figure: Figure 2
    Supports: /model_ids
  • figure: Figure 2
    Supports: /model_ids

Creation, training, validation, or model-selection uses

benchmark creation

biosecbench-surveillance-a-verifiable-benchmark-for-ai-biosecbench-surveillance-1-use

Non-evaluationunknown

Work: BioSecBench-Surveillance: A Verifiable Benchmark for AI Agents in Pathogen Genomic Surveillance · source version biosecbench-surveillance-a-verifiable-benchmark-for-ai-arxiv-v1

Selection
not applicable
Models
Not reported / not applicable
Metrics
Not reported / not applicable
Linked runs
None

Not reported / unresolved: Benchmark version is not reported.; The exact public-subset size is not reported.; The repository license is not reported.

AI-assisted double-pass extraction; values are limited to independently supported claims.

Evidence
  • section: Abstract
    Supports: /relation_type
  • section: Abstract
    Supports: /benchmark_id

Evaluation registry

Works and run settings

A setting change—scope, prompt, tools, budget, grader, or repeats—creates a separate run. Charts never cross a comparability group.

Evaluation run

biosecbench-8d53fd8-claude-code

From BioSecBench-Surveillance repository result snapshot

biosecbench-8d53fd8-claude-codevinitial-release

Evaluated models / systems: Opus 4.6, Opus 4.7, Opus 4.8, Sonnet 4.6

Scopefull · n=100
ShotsNot reported
TurnsNot reported
System prompt publicNot reported
Reasoning / effortNot reported
BrowserNot reported
InternetYes
Databasesresistance databases, virulence databases
Code executionNot reported
Containeridentical containerized sandbox
External toolsClaude Code, assemblers, aligners, taxonomic classifiers, variant callers
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
Repeats3
Graderdeterministic typed-field checks · human review: not reported
Statistics95% Student-t confidence interval over evaluations
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
endpoint pass rateabsolute%mean of per-evaluation pass rates; evaluations with no gradeable run are droppedNot reported

Results

ModelMetricValuen
Opus 4.6endpoint pass rate46.7 %
Claude Code; n is the number of gradable evaluations after evaluations with no gradable run were dropped; the configured benchmark scope was all 100 evaluations.
80
Opus 4.7endpoint pass rate44.2 %
Claude Code; n is the number of gradable evaluations after evaluations with no gradable run were dropped; the configured benchmark scope was all 100 evaluations.
77
Sonnet 4.6endpoint pass rate43.6 %
Claude Code; n is the number of gradable evaluations after evaluations with no gradable run were dropped; the configured benchmark scope was all 100 evaluations.
83
Opus 4.8endpoint pass rate39.5 %
Claude Code; n is the number of gradable evaluations after evaluations with no gradable run were dropped; the configured benchmark scope was all 100 evaluations.
76

Evidence

  • repository-path: METHODS.md at commit 8d53fd8517cc74202eb18b618e8b39b4ffaf0c87 (Full 100-evaluation scope, Claude Code harness, tools, three repeats, deterministic grading, aggregation, and confidence intervals.) — supports /benchmark_version, /scope, /protocol, /metrics
  • repository-path: results/config_results.csv; harness=claude-code at commit 8d53fd8517cc74202eb18b618e8b39b4ffaf0c87 (Exact model labels, endpoint pass rates, 95% confidence intervals, and gradable evaluation counts.) — supports /results

Evaluation run

biosecbench-8d53fd8-openai-codex

From BioSecBench-Surveillance repository result snapshot

biosecbench-8d53fd8-openai-codexvinitial-release

Evaluated models / systems: GPT-5.4, GPT-5.5

Scopefull · n=100
ShotsNot reported
TurnsNot reported
System prompt publicNot reported
Reasoning / effortNot reported
BrowserNot reported
InternetYes
Databasesresistance databases, virulence databases
Code executionNot reported
Containeridentical containerized sandbox
External toolsOpenAI Codex, assemblers, aligners, taxonomic classifiers, variant callers
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
Repeats3
Graderdeterministic typed-field checks · human review: not reported
Statistics95% Student-t confidence interval over evaluations
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
endpoint pass rateabsolute%mean of per-evaluation pass rates; evaluations with no gradeable run are droppedNot reported

Results

ModelMetricValuen
GPT-5.5endpoint pass rate50.2 %
OpenAI Codex; n is the number of gradable evaluations after evaluations with no gradable run were dropped; the configured benchmark scope was all 100 evaluations.
93
GPT-5.4endpoint pass rate41.3 %
OpenAI Codex; n is the number of gradable evaluations after evaluations with no gradable run were dropped; the configured benchmark scope was all 100 evaluations.
92

Evidence

  • repository-path: METHODS.md at commit 8d53fd8517cc74202eb18b618e8b39b4ffaf0c87 (Full 100-evaluation scope, OpenAI Codex harness, tools, three repeats, deterministic grading, aggregation, and confidence intervals.) — supports /benchmark_version, /scope, /protocol, /metrics
  • repository-path: results/config_results.csv; harness=openai-codex at commit 8d53fd8517cc74202eb18b618e8b39b4ffaf0c87 (Exact model labels, endpoint pass rates, 95% confidence intervals, and gradable evaluation counts.) — supports /results

Evaluation run

biosecbench-8d53fd8-pi

From BioSecBench-Surveillance repository result snapshot

biosecbench-8d53fd8-pivinitial-release

Evaluated models / systems: Opus 4.6, Opus 4.7, Opus 4.8, Sonnet 4.6, gemini-3.1-pro-preview, Gemini 3.5 Flash, GPT-5.4, GPT-5.5, grok-4.20-beta-0309-reasoning, Grok 4.3

Scopefull · n=100
ShotsNot reported
TurnsNot reported
System prompt publicNot reported
Reasoning / effortNot reported
BrowserNot reported
InternetYes
Databasesresistance databases, virulence databases
Code executionNot reported
Containeridentical containerized sandbox
External toolsPI, assemblers, aligners, taxonomic classifiers, variant callers
Token budgetNot reported
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
Repeats3
Graderdeterministic typed-field checks · human review: not reported
Statistics95% Student-t confidence interval over evaluations
ContaminationNot reported
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
endpoint pass rateabsolute%mean of per-evaluation pass rates; evaluations with no gradeable run are droppedNot reported

Results

ModelMetricValuen
Opus 4.8endpoint pass rate50.2 %
PI; n is the number of gradable evaluations after evaluations with no gradable run were dropped; the configured benchmark scope was all 100 evaluations.
83
GPT-5.5endpoint pass rate44.8 %
PI; n is the number of gradable evaluations after evaluations with no gradable run were dropped; the configured benchmark scope was all 100 evaluations.
77
Opus 4.7endpoint pass rate49.6 %
PI; n is the number of gradable evaluations after evaluations with no gradable run were dropped; the configured benchmark scope was all 100 evaluations.
78
Sonnet 4.6endpoint pass rate48.6 %
PI; n is the number of gradable evaluations after evaluations with no gradable run were dropped; the configured benchmark scope was all 100 evaluations.
83
Gemini 3.5 Flashendpoint pass rate47.1 %
PI; n is the number of gradable evaluations after evaluations with no gradable run were dropped; the configured benchmark scope was all 100 evaluations.
92
Opus 4.6endpoint pass rate45.7 %
PI; n is the number of gradable evaluations after evaluations with no gradable run were dropped; the configured benchmark scope was all 100 evaluations.
81
gemini-3.1-pro-previewendpoint pass rate45.3 %
PI; n is the number of gradable evaluations after evaluations with no gradable run were dropped; the configured benchmark scope was all 100 evaluations.
95
GPT-5.4endpoint pass rate38.5 %
PI; n is the number of gradable evaluations after evaluations with no gradable run were dropped; the configured benchmark scope was all 100 evaluations.
81
Grok 4.3endpoint pass rate16 %
PI; n is the number of gradable evaluations after evaluations with no gradable run were dropped; the configured benchmark scope was all 100 evaluations.
100
grok-4.20-beta-0309-reasoningendpoint pass rate13.7 %
PI; n is the number of gradable evaluations after evaluations with no gradable run were dropped; the configured benchmark scope was all 100 evaluations.
100

Evidence

  • repository-path: METHODS.md at commit 8d53fd8517cc74202eb18b618e8b39b4ffaf0c87 (Full 100-evaluation scope, PI harness, tools, three repeats, deterministic grading, aggregation, and confidence intervals.) — supports /benchmark_version, /scope, /protocol, /metrics
  • repository-path: results/config_results.csv; harness=pi at commit 8d53fd8517cc74202eb18b618e8b39b4ffaf0c87 (Exact model labels, endpoint pass rates, 95% confidence intervals, and gradable evaluation counts.) — supports /results

Comparable result views

endpoint pass rate

biosecbench-8d53fd8-claude-code · biosecbench-8d53fd8-claude-code

CSV ↓
Accessible data table
ModelValueComparability group
Opus 4.646.7biosecbench-8d53fd8-claude-code
Opus 4.744.2biosecbench-8d53fd8-claude-code
Sonnet 4.643.6biosecbench-8d53fd8-claude-code
Opus 4.839.5biosecbench-8d53fd8-claude-code

endpoint pass rate

biosecbench-8d53fd8-openai-codex · biosecbench-8d53fd8-openai-codex

CSV ↓
Accessible data table
ModelValueComparability group
GPT-5.550.2biosecbench-8d53fd8-openai-codex
GPT-5.441.3biosecbench-8d53fd8-openai-codex

endpoint pass rate

biosecbench-8d53fd8-pi · biosecbench-8d53fd8-pi

CSV ↓
Accessible data table
ModelValueComparability group
Opus 4.850.2biosecbench-8d53fd8-pi
GPT-5.544.8biosecbench-8d53fd8-pi
Opus 4.749.6biosecbench-8d53fd8-pi
Sonnet 4.648.6biosecbench-8d53fd8-pi
Gemini 3.5 Flash47.1biosecbench-8d53fd8-pi
Opus 4.645.7biosecbench-8d53fd8-pi
gemini-3.1-pro-preview45.3biosecbench-8d53fd8-pi
GPT-5.438.5biosecbench-8d53fd8-pi
Grok 4.316biosecbench-8d53fd8-pi
grok-4.20-beta-0309-reasoning13.7biosecbench-8d53fd8-pi

Evidence and change history

Source locators remain visible; expand an item to inspect the exact Registry fields it supports.

BioSecBench-Surveillance: A Verifiable Benchmark for AI Agents in Pathogen Genomic Surveillance · section: Data availability · Supports 1 field

Open source →

  • /access/artifacts
BioSecBench-Surveillance: A Verifiable Benchmark for AI Agents in Pathogen Genomic Surveillance · section: Data availability · Supports 1 field

Open source →

  • /access/biosafety_notes
BioSecBench-Surveillance: A Verifiable Benchmark for AI Agents in Pathogen Genomic Surveillance · section: Data availability · Supports 1 field

Open source →

  • /access/level
BioSecBench-Surveillance: A Verifiable Benchmark for AI Agents in Pathogen Genomic Surveillance · section: Data availability · Supports 1 field

Open source →

  • /access/tasks
BioSecBench-Surveillance: A Verifiable Benchmark for AI Agents in Pathogen Genomic Surveillance · section: Introduction · Supports 1 field

Open source →

  • /capabilities
BioSecBench-Surveillance: A Verifiable Benchmark for AI Agents in Pathogen Genomic Surveillance · section: Methods — Benchmark composition and data; Agent runs and execution · Supports 1 field

Open source →

  • /domains
BioSecBench-Surveillance: A Verifiable Benchmark for AI Agents in Pathogen Genomic Surveillance · section: Introduction · Supports 1 field

Open source →

  • /kind
BioSecBench-Surveillance: A Verifiable Benchmark for AI Agents in Pathogen Genomic Surveillance · section: Methods — Benchmark composition and data; Agent runs and execution · Supports 1 field

Open source →

  • /modalities
BioSecBench-Surveillance: A Verifiable Benchmark for AI Agents in Pathogen Genomic Surveillance · section: Abstract · Supports 1 field

Open source →

  • /name
BioSecBench-Surveillance: A Verifiable Benchmark for AI Agents in Pathogen Genomic Surveillance · page: Author-affiliation mapping · Supports 1 field

Open source →

  • /organizations
BioSecBench-Surveillance: A Verifiable Benchmark for AI Agents in Pathogen Genomic Surveillance · page: arXiv dateline · Supports 1 field

Open source →

  • /release_date
BioSecBench-Surveillance: A Verifiable Benchmark for AI Agents in Pathogen Genomic Surveillance · section: Introduction and Benchmark construction · Supports 1 field

Open source →

  • /summary
BioSecBench-Surveillance: A Verifiable Benchmark for AI Agents in Pathogen Genomic Surveillance · section: Methods — Task format and deterministic grading · Supports 1 field

Open source →

  • /task_formats
BioSecBench-Surveillance: A Verifiable Benchmark for AI Agents in Pathogen Genomic Surveillance · section: Abstract · Supports 4 fields

Open source →

  • /task_counts/total
  • /task_counts/basis
  • /versions/0/task_counts/total
  • /versions/0/task_counts/basis
BioSecBench-Surveillance: A Verifiable Benchmark for AI Agents in Pathogen Genomic Surveillance · figure: Figure 1A · Supports 4 fields

Open source →

  • /task_counts/subsets
  • /versions/0/task_counts/subsets
  • /task_counts/subsets/0
  • /versions/0/task_counts/subsets/0
BioSecBench-Surveillance: A Verifiable Benchmark for AI Agents in Pathogen Genomic Surveillance · figure: Figure 1A · Supports 2 fields

Open source →

  • /task_counts/subsets/1
  • /versions/0/task_counts/subsets/1
BioSecBench-Surveillance: A Verifiable Benchmark for AI Agents in Pathogen Genomic Surveillance · figure: Figure 1A · Supports 2 fields

Open source →

  • /task_counts/subsets/2
  • /versions/0/task_counts/subsets/2
BioSecBench-Surveillance: A Verifiable Benchmark for AI Agents in Pathogen Genomic Surveillance · figure: Figure 1A · Supports 2 fields

Open source →

  • /task_counts/subsets/3
  • /versions/0/task_counts/subsets/3
BioSecBench-Surveillance: A Verifiable Benchmark for AI Agents in Pathogen Genomic Surveillance · figure: Figure 1A · Supports 2 fields

Open source →

  • /task_counts/subsets/4
  • /versions/0/task_counts/subsets/4
BioSecBench-Surveillance: A Verifiable Benchmark for AI Agents in Pathogen Genomic Surveillance · figure: Figure 1A · Supports 2 fields

Open source →

  • /task_counts/subsets/5
  • /versions/0/task_counts/subsets/5
BioSecBench-Surveillance: A Verifiable Benchmark for AI Agents in Pathogen Genomic Surveillance · figure: Figure 1A · Supports 2 fields

Open source →

  • /task_counts/subsets/6
  • /versions/0/task_counts/subsets/6
BioSecBench-Surveillance: A Verifiable Benchmark for AI Agents in Pathogen Genomic Surveillance · figure: Figure 1B · Supports 2 fields

Open source →

  • /task_counts/subsets/7
  • /versions/0/task_counts/subsets/7
BioSecBench-Surveillance: A Verifiable Benchmark for AI Agents in Pathogen Genomic Surveillance · figure: Figure 1B · Supports 2 fields

Open source →

  • /task_counts/subsets/8
  • /versions/0/task_counts/subsets/8
BioSecBench-Surveillance: A Verifiable Benchmark for AI Agents in Pathogen Genomic Surveillance · figure: Figure 1B · Supports 2 fields

Open source →

  • /task_counts/subsets/9
  • /versions/0/task_counts/subsets/9
BioSecBench-Surveillance: A Verifiable Benchmark for AI Agents in Pathogen Genomic Surveillance · figure: Figure 1B · Supports 2 fields

Open source →

  • /task_counts/subsets/10
  • /versions/0/task_counts/subsets/10
BioSecBench-Surveillance: A Verifiable Benchmark for AI Agents in Pathogen Genomic Surveillance · figure: Figure 1B · Supports 2 fields

Open source →

  • /task_counts/subsets/11
  • /versions/0/task_counts/subsets/11
BioSecBench-Surveillance: A Verifiable Benchmark for AI Agents in Pathogen Genomic Surveillance · figure: Figure 1B · Supports 2 fields

Open source →

  • /task_counts/subsets/12
  • /versions/0/task_counts/subsets/12
BioSecBench-Surveillance: A Verifiable Benchmark for AI Agents in Pathogen Genomic Surveillance · figure: Figure 1C · Supports 2 fields

Open source →

  • /task_counts/subsets/13
  • /versions/0/task_counts/subsets/13
BioSecBench-Surveillance: A Verifiable Benchmark for AI Agents in Pathogen Genomic Surveillance · figure: Figure 1C · Supports 2 fields

Open source →

  • /task_counts/subsets/14
  • /versions/0/task_counts/subsets/14
BioSecBench-Surveillance: A Verifiable Benchmark for AI Agents in Pathogen Genomic Surveillance · figure: Figure 1C · Supports 2 fields

Open source →

  • /task_counts/subsets/15
  • /versions/0/task_counts/subsets/15
BioSecBench-Surveillance: A Verifiable Benchmark for AI Agents in Pathogen Genomic Surveillance · figure: Figure 1D · Supports 2 fields

Open source →

  • /task_counts/subsets/16
  • /versions/0/task_counts/subsets/16
BioSecBench-Surveillance: A Verifiable Benchmark for AI Agents in Pathogen Genomic Surveillance · figure: Figure 1D · Supports 2 fields

Open source →

  • /task_counts/subsets/17
  • /versions/0/task_counts/subsets/17
BioSecBench-Surveillance: A Verifiable Benchmark for AI Agents in Pathogen Genomic Surveillance · figure: Figure 1D · Supports 2 fields

Open source →

  • /task_counts/subsets/18
  • /versions/0/task_counts/subsets/18
BioSecBench-Surveillance: A Verifiable Benchmark for AI Agents in Pathogen Genomic Surveillance · figure: Figure 1E · Supports 2 fields

Open source →

  • /task_counts/subsets/19
  • /versions/0/task_counts/subsets/19
BioSecBench-Surveillance: A Verifiable Benchmark for AI Agents in Pathogen Genomic Surveillance · figure: Figure 1E · Supports 2 fields

Open source →

  • /task_counts/subsets/20
  • /versions/0/task_counts/subsets/20
BioSecBench-Surveillance: A Verifiable Benchmark for AI Agents in Pathogen Genomic Surveillance · section: Data availability · Supports 3 fields

Open source →

  • /resources
  • /implementations
  • /access/license
BioSecBench-Surveillance: A Verifiable Benchmark for AI Agents in Pathogen Genomic Surveillance · page: arXiv dateline · Supports 1 field

Open source →

  • /resources/0
BioSecBench-Surveillance: A Verifiable Benchmark for AI Agents in Pathogen Genomic Surveillance · page: arXiv dateline · Supports 2 fields

Open source →

  • /latest_version
  • /versions/0

Unresolved field claims

  • /access/licenseProvisional · high — The double-pass review verified the official resource identity but did not establish a redistributable benchmark license; the value remains null pending source-level license verification.
    Evidence: biosecbench-surveillance-automated-resource-evidence

View source-level modification history on GitHub →