Evaluation run
bixbench-v1-5-agentic-mcq-no-refusal-images
Evaluated models / systems: Claude 3.5 Sonnet 20241022, GPT-4o (BixBench version not reported)
Metrics, results, and full protocol
Metrics
| Metric | Kind / baseline | Unit | Aggregation | Threshold / tolerance |
|---|---|---|---|---|
| Accuracy | absolute | percent | question-replica-weighted mean | exact selected option |
No numeric result rows are published yet; the verified protocol remains useful.
Evidence
- repository-path: image model configs, v1.5_paper_results.yaml, README.md, and scripts/run_agentic.sh at commit 49311180bdacb324c596f2e07596c126f2004008 (Defines images allowed, forced-choice condition, full dataset, agent settings, metric, and majority-vote postprocessing.) — supports /benchmark_version, /scope, /protocol, /metrics