official-release · benchmark creator
Single-cell Omics Arena repository result snapshot
University of California, Irvine · 2025-08-03
Relationship layer
Benchmark usage
This table records what the work did with each benchmark before attempting to normalize a run. Partial claims remain visible without being treated as comparable evaluations.
evaluation
soar-e5d2b3e-rna-zero-shot-cot-use
Normalizedsubset · n=1191
Benchmark: Single-cell Omics Arena · version initial-release
- Selection
- formal-subset · All 1,191 records in the pinned SOAR-RNA JSON artifact, iterated in repository order with shuffle disabled.
- Metrics
- R-1, R-2, R-L, MET., B-1, B-2, BLEU
Not reported / unresolved: repeat count; confidence intervals; contamination or decontamination analysis
Official commit-pinned SOAR-RNA two-call zero-shot chain-of-thought result snapshot; the run covers the formal SOAR-RNA subset, not the full SOAR suite.
Evidence
- repository-path: readme.md; soar_benchmark/configs/cell_type_annotation/experiment_soar_rna.py; soar_benchmark/task.py; soar_benchmark/datasets/soar_rna.json at commit e5d2b3e2619cb56fece5fba78fae989a67fd0c13 (SOAR-RNA zero-shot CoT relation, exact models, formal subset scope, metrics, and linked run.)
Supports: /benchmark_version, /relation_type, /status, /model_ids, /scope, /metric_labels, /evaluation_run_ids, /reporting_gaps, /notes
evaluation
soar-e5d2b3e-rna-zero-shot-use
Normalizedsubset · n=1191
Benchmark: Single-cell Omics Arena · version initial-release
- Selection
- formal-subset · All 1,191 records in the pinned SOAR-RNA JSON artifact, iterated in repository order with shuffle disabled.
- Metrics
- R-1, R-2, R-L, MET., B-1, B-2, BLEU
Not reported / unresolved: repeat count; confidence intervals; contamination or decontamination analysis
Official commit-pinned SOAR-RNA zero-shot result snapshot; the run covers the formal SOAR-RNA subset, not the full SOAR suite.
Evidence
- repository-path: readme.md; soar_benchmark/configs/cell_type_annotation/experiment_soar_rna.py; soar_benchmark/datasets/soar_rna.json at commit e5d2b3e2619cb56fece5fba78fae989a67fd0c13 (SOAR-RNA zero-shot relation, exact models, formal subset scope, metrics, and linked run.)
Supports: /benchmark_version, /relation_type, /status, /model_ids, /scope, /metric_labels, /evaluation_run_ids, /reporting_gaps, /notes
Normalized evaluation runs
Single-cell Omics Arena2 runs
Open benchmark record →
soar-e5d2b3e-rna-zero-shotvinitial-release
Evaluated models / systems: gpt-4o-2024-05-13, gpt-4o-mini-2024-07-18
Scopesubset · n=1191
Shots0
Turnssingle-turn
System prompt publicYes
Reasoning / effortnone
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNo
External tools
Token budget1024 maximum output tokens
Time / cost budgetNot reported
Temperature0
SeedNot reported
RepeatsNot reported
Graderautomatic evaluate-library text-overlap metrics · human review: not reported
Statisticsmetrics computed over all predictions and normalized references
ContaminationNot reported
Metrics, results, and full protocol
Metrics
| Metric | Kind / baseline | Unit | Aggregation | Threshold / tolerance |
|---|
| R-1 | absolute | source-reported score | all predictions and normalized references | Not reported |
| R-2 | absolute | source-reported score | all predictions and normalized references | Not reported |
| R-L | absolute | source-reported score | all predictions and normalized references | Not reported |
| MET. | absolute | source-reported score | all predictions and normalized references | Not reported |
| B-1 | absolute | source-reported score | all predictions and normalized references; maximum order two | Not reported |
| B-2 | absolute | source-reported score | all predictions and normalized references; maximum order two | Not reported |
| BLEU | absolute | source-reported score | geometric average through order two over all predictions and normalized references | Not reported |
Results
Evidence
- repository-path: soar_benchmark/datasets/soar_rna.json; soar_benchmark/configs/cell_type_annotation/experiment_soar_rna.py; soar_benchmark/task.py; soar_benchmark/pipeline.py; soar_benchmark/prompt_templates/factory.py; analysis/cell_type_annotation/eval_multiple.py at commit e5d2b3e2619cb56fece5fba78fae989a67fd0c13 (Scope, exact models, zero-shot prompt, model call, budgets, tools, and automatic metric implementation.) — supports /benchmark_version, /scope, /model_ids, /protocol, /metrics
- repository-path: readme.md, SOAR-RNA Benchmark, Zero-shot Cell Type Annotation table at commit e5d2b3e2619cb56fece5fba78fae989a67fd0c13 (Exact GPT-4o and GPT-4o mini values for the seven printed metrics.) — supports /results
soar-e5d2b3e-rna-zero-shot-cotvinitial-release
Evaluated models / systems: gpt-4o-2024-05-13, gpt-4o-mini-2024-07-18
Scopesubset · n=1191
Shots0
Turnstwo model calls
System prompt publicYes
Reasoning / effortzero-shot chain-of-thought
BrowserNo
InternetNo
DatabasesNo
Code executionNo
ContainerNo
External tools
Token budget1024 maximum output tokens per call; two calls per example
Time / cost budgetNot reported
Temperature0
SeedNot reported
RepeatsNot reported
Graderautomatic evaluate-library text-overlap metrics · human review: not reported
Statisticsmetrics computed over all predictions and normalized references
ContaminationNot reported
Metrics, results, and full protocol
Metrics
| Metric | Kind / baseline | Unit | Aggregation | Threshold / tolerance |
|---|
| R-1 | absolute | source-reported score | all predictions and normalized references | Not reported |
| R-2 | absolute | source-reported score | all predictions and normalized references | Not reported |
| R-L | absolute | source-reported score | all predictions and normalized references | Not reported |
| MET. | absolute | source-reported score | all predictions and normalized references | Not reported |
| B-1 | absolute | source-reported score | all predictions and normalized references; maximum order two | Not reported |
| B-2 | absolute | source-reported score | all predictions and normalized references; maximum order two | Not reported |
| BLEU | absolute | source-reported score | geometric average through order two over all predictions and normalized references | Not reported |
Results
Evidence
- repository-path: soar_benchmark/datasets/soar_rna.json; soar_benchmark/configs/cell_type_annotation/experiment_soar_rna.py; soar_benchmark/task.py; soar_benchmark/pipeline.py; soar_benchmark/prompt_templates/factory.py; analysis/cell_type_annotation/eval_multiple.py at commit e5d2b3e2619cb56fece5fba78fae989a67fd0c13 (Scope, exact models, two-call CoT prompt, model calls, budgets, tools, and automatic metric implementation.) — supports /benchmark_version, /scope, /model_ids, /protocol, /metrics
- repository-path: readme.md, SOAR-RNA Benchmark, Zero-shot Chain-of-thought Cell Type Annotation table at commit e5d2b3e2619cb56fece5fba78fae989a67fd0c13 (Exact GPT-4o and GPT-4o mini values for the seven printed metrics.) — supports /results