Evaluation run
bixbench-creator-paper
From BixBench: a Comprehensive Benchmark for LLM-based Agents in Computational Biology
Evaluated models / systems: Claude 3.5 Sonnet (BixBench version not reported), GPT-4o (BixBench version not reported)
Metrics, results, and full protocol
Metrics
| Metric | Kind / baseline | Unit | Aggregation | Threshold / tolerance |
|---|---|---|---|---|
| Accuracy | absolute | percent | question-trajectory-weighted mean over ten runs per question | binary LLM judge |
Results
| Model | Metric | Value | n |
|---|---|---|---|
| GPT-4o (BixBench version not reported) | Accuracy | 9 percent Rounded open-answer headline in the paper text; ten trajectories per capsule, with plots/images allowed. | 296 |
| Claude 3.5 Sonnet (BixBench version not reported) | Accuracy | 17 percent Rounded open-answer headline in the paper text; ten trajectories per capsule, with plots/images allowed. | 296 |
Evidence
- page: PDF pp. 4–7, §§3.2.2–3.2.5 and §4.1 (Reports 296 questions, Docker environment, three tools, notebook execution, ten parallel analyses, plots/images condition, Claude judge, and aggregation.) — supports /benchmark_version, /scope, /protocol, /metrics
- repository-path: bixbench/config.yaml at commit 6c28217959d5d7dd6f48c59894534fced7c6c040 (Records ReActAgent, temperature 1.0, maximum 25 steps, public prompt key, and total_questions 296.) — supports /protocol/reasoning, /protocol/system_prompt_public, /protocol/time_budget, /protocol/temperature
- page: PDF pp. 1 and 6, abstract and §4.1; Figure 4 (Text explicitly reports Claude 3.5 Sonnet at 17% and GPT-4o at 9% in open-answer evaluation.) — supports /results