Evaluation run
blade-creator-decision-mcq
From BLADE: Benchmarking Language Model Agents for Data-Driven Science
Evaluated models / systems: Claude 3.5 Sonnet (claude-3-5-sonnet-20240620; BLADE), CodeLlama Instruct 7B (BLADE), DeepSeek-Coder Instruct 6.7B (BLADE), Gemini 1.5 Pro (BLADE snapshot not reported), GPT-3.5 Turbo (BLADE snapshot not reported), GPT-4o (BLADE snapshot not reported), Llama 3 70B (BLADE snapshot not reported), Llama 3 8B (BLADE snapshot not reported), Mixtral 8x22B (BLADE snapshot not reported)
Metrics, results, and full protocol
Metrics
| Metric | Kind / baseline | Unit | Aggregation | Threshold / tolerance |
|---|---|---|---|---|
| Accuracy | absolute | percent | correct choices divided by all 188 MCQs | exact option identifier |
No numeric result rows are published yet; the verified protocol remains useful.
Evidence
- section: arXiv v3 §§4.1 and 6, Figure 3, and Appendix A.6 Figure 19 (Reports 188 MCQs, nine evaluated model families, temperature zero, accuracy with 95% intervals, and the public prompt.) — supports /benchmark_version, /model_ids, /scope, /protocol/temperature, /protocol/statistical, /metrics, /comparability_group
- repository-path: run_mcq.py, run_scripts/sh_run_mcq.sh, blade_bench/baselines/lm/mcq.py, blade_bench/eval/datamodel/run_mcq.py, and blade_bench/conf/llm_config.yml at commit 6118fa8d5007b91aa8c91c518182db82446a4547 (Confirms one direct prompt per item, no tool loop, exact-choice grading, complete public prompt, 20/168 question components, and available exact model strings.) — supports /scope, /protocol/shots, /protocol/turns, /protocol/system_prompt_public, /protocol/reasoning, /protocol/tools, /protocol/repeats, /protocol/grader, /model_ids