Relationship layer
Benchmark usage
This table records what the work did with each benchmark before attempting to normalize a run. Partial claims remain visible without being treated as comparable evaluations.
benchmark creation
biosecbench-surveillance-a-verifiable-benchmark-for-ai-biosecbench-surveillance-1-use
Non-evaluationunknown
Benchmark: BioSecBench-Surveillance · version Not reported
- Selection
- not applicable
- Models
- Not reported / not applicable
- Metrics
- Not reported / not applicable
- Linked runs
- None
Not reported / unresolved: Benchmark version is not reported.; The exact public-subset size is not reported.; The repository license is not reported.
AI-assisted double-pass extraction; values are limited to independently supported claims.
Evidence
- section: Abstract
Supports: /relation_type - section: Abstract
Supports: /benchmark_id
evaluation
biosecbench-surveillance-a-verifiable-benchmark-for-ai-biosecbench-surveillance-2-use
Partialunknown · n=100
Benchmark: BioSecBench-Surveillance · version Not reported
- Selection
- not reported
- Metrics
- Endpoint pass rate
- Linked runs
- None
Not reported / unresolved: Benchmark version is not reported.; Exact deployment snapshots and model release dates are not reported.; Prompt text, shots, token budget, and random seed are not reported.; Gradable evaluation n is reported only for Opus 4.8 / PI.; Exact confidence intervals are not numerically printed for seven PI results.; benchmark version; numeric result
AI-assisted double-pass extraction; values are limited to independently supported claims.
Evidence
- figure: Figure 2
Supports: /relation_type - figure: Figure 2
Supports: /benchmark_id - section: Methods — Agent runs and execution
Supports: /scope - section: Methods — Agent runs and execution
Supports: /scope - section: Methods — Outcome classification and aggregation
Supports: /metric_labels - figure: Figure 2
Supports: /model_ids - figure: Figure 2
Supports: /model_ids - figure: Figure 2
Supports: /model_ids - figure: Figure 2
Supports: /model_ids - figure: Figure 2
Supports: /model_ids - figure: Figure 2
Supports: /model_ids - figure: Figure 2
Supports: /model_ids - figure: Figure 2
Supports: /model_ids - figure: Figure 2
Supports: /model_ids - figure: Figure 2
Supports: /model_ids
evaluation
biosecbench-surveillance-a-verifiable-benchmark-for-ai-biosecbench-surveillance-3-use
Partialunknown · n=100
Benchmark: BioSecBench-Surveillance · version Not reported
- Selection
- not reported
- Metrics
- Endpoint pass rate
- Linked runs
- None
Not reported / unresolved: Benchmark version is not reported.; Exact deployment snapshots and model release dates are not reported.; Prompt text, shots, token budget, and random seed are not reported.; Per-configuration gradable n and numerical confidence intervals are not reported.; benchmark version; numeric result
AI-assisted double-pass extraction; values are limited to independently supported claims.
Evidence
- figure: Figure 2
Supports: /relation_type - figure: Figure 2
Supports: /benchmark_id - section: Methods — Agent runs and execution
Supports: /scope - section: Methods — Agent runs and execution
Supports: /scope - section: Methods — Outcome classification and aggregation
Supports: /metric_labels - figure: Figure 2
Supports: /model_ids - figure: Figure 2
Supports: /model_ids - figure: Figure 2
Supports: /model_ids - figure: Figure 2
Supports: /model_ids
evaluation
biosecbench-surveillance-a-verifiable-benchmark-for-ai-biosecbench-surveillance-4-use
Partialunknown · n=100
Benchmark: BioSecBench-Surveillance · version Not reported
- Selection
- not reported
- Metrics
- Endpoint pass rate
- Linked runs
- None
Not reported / unresolved: Benchmark version is not reported.; Exact deployment snapshots and model release dates are not reported.; Prompt text, shots, token budget, and random seed are not reported.; Gradable evaluation n is not reported for either Codex configuration.; GPT-5.4 / Codex has no numerically printed confidence interval.; benchmark version; numeric result
AI-assisted double-pass extraction; values are limited to independently supported claims.
Evidence
- figure: Figure 2
Supports: /relation_type - figure: Figure 2
Supports: /benchmark_id - section: Methods — Agent runs and execution
Supports: /scope - section: Methods — Agent runs and execution
Supports: /scope - section: Methods — Outcome classification and aggregation
Supports: /metric_labels - figure: Figure 2
Supports: /model_ids - figure: Figure 2
Supports: /model_ids