paper · benchmark creator

Beyond Single Chains: Benchmarking Macromolecular Complex Prediction Methods With the Continuous Automated Model EvaluatiOn (CAMEO)

SIB Swiss Institute of Bioinformatics · Biozentrum University of Basel · 2025-09-28

Relationship layer

Benchmark usage

This table records what the work did with each benchmark before attempting to normalize a run. Partial claims remain visible without being treated as comparable evaluations.

No BenchmarkUse relation is normalized for this legacy work yet. Existing EvaluationRuns remain available below.

Normalized evaluation runs

CAMEO3 runs

Open benchmark record →

cameo-2024-antibody-three-server-commonv2024-complex-study

Evaluated models / systems: AlphaFold 3 v3.0.1 (CAMEO baseline), MultiFOLD (CAMEO 2024 participant), SWISS-MODEL (CAMEO 2024 participant)

Scopesubset · n=83
ShotsNot applicable
TurnsNot applicable
System prompt publicNot applicable
Reasoning / effortNot applicable
Browsermethod-specific
Internetallowed before each weekly deadline
Databasesmethod-specific public structure and sequence data
Code executionallowed
ContainerNot reported
External toolsmethod-specific server pipeline
Token budgetNot applicable
Time / cost budgetapproximately 3.5 days per weekly target
TemperatureNot applicable
SeedNot reported
Repeatsup to five submitted models per target
Graderfully automated OpenStructure complex LDDT scoring · human review: no
Statisticsmedian Complex LDDT across 83 common-subset antibody targets
Contaminationexperimental antibody-complex structures withheld until PDB release
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Median Complex LDDTabsoluteproportionmedian across 83 antibody common-subset targetsmissing chains are penalized

Results

ModelMetricValuen
AlphaFold 3 v3.0.1 (CAMEO baseline)Median Complex LDDT0.83 proportion
Median reported in creator-paper Section 2.4.2.
83
MultiFOLD (CAMEO 2024 participant)Median Complex LDDT0.76 proportion
Median reported in creator-paper Section 2.4.2.
83

Evidence

  • section: Section 2.4.2 and Figure 2A (Reports 83 antibody common-subset targets, AF3 median LDDT 0.83, MultiFOLD median LDDT 0.76, and the model-1 comparison protocol.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics, /results
cameo-2024-ligand-four-baseline-commonv2024-complex-study

Evaluated models / systems: AlphaFold 3 v3.0.1 (CAMEO baseline), SWISS-MODEL + Schrödinger Glide, SWISS-MODEL + AutoDock Vina (AutoDock4 scoring), SWISS-MODEL + AutoDock Vina (vina scoring)

Scopesubset · n=2584
ShotsNot applicable
TurnsNot applicable
System prompt publicNot applicable
Reasoning / effortNot applicable
Browsermethod-specific
Internetallowed before each weekly deadline
Databasesmethod-specific public structure and sequence data
Code executionallowed
ContainerNot reported
External toolsAlphaFold 3 v3.0.1 or SWISS-MODEL followed by Glide or AutoDock Vina 1.2.5
Token budgetNot applicable
Time / cost budgetapproximately 3.5 days per weekly target
TemperatureNot applicable
SeedNot reported
Repeatsup to five models or ligand poses per target
Graderfully automated OpenStructure complex and ligand scoring · human review: no
Statisticscommon-subset aggregation across targets predicted by all four baselines
ContaminationPDB pre-release sequences and ligand identities available while experimental structures and poses remain withheld
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Ligand success rateabsolutepercent of ligand entitiespercentage across common-subset ligand entitiesA success is a symmetry-corrected BiSyRMSD below 2 Å after binding-site superposition.
LDDT-PLIabsoluteproportionweighted aggregation across relevant ligand entitiesNot reported
BiSyRMSDabsoluteangstrombest-scored submitted pose per ligand entity in the creator analysis2

No numeric result rows are published yet; the verified protocol remains useful.

Evidence

  • section: Sections 2.3, 2.4.1, 2.5.1-2.5.2, and 3.2; Figure 1 (Reports the four baselines, 2,584-target common subset, 6,152 entities, top-ranked model handling, ligand metrics, and exact pipeline versions.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics
cameo-2024-ppi-three-server-commonv2024-complex-study

Evaluated models / systems: AlphaFold 3 v3.0.1 (CAMEO baseline), MultiFOLD (CAMEO 2024 participant), SWISS-MODEL (CAMEO 2024 participant)

Scopesubset · n=392
ShotsNot applicable
TurnsNot applicable
System prompt publicNot applicable
Reasoning / effortNot applicable
Browsermethod-specific
Internetallowed before each weekly deadline
Databasesmethod-specific public structure and sequence data
Code executionallowed
ContainerNot reported
External toolsmethod-specific server pipeline
Token budgetNot applicable
Time / cost budgetapproximately 3.5 days per weekly target
TemperatureNot applicable
SeedNot reported
Repeatsup to five submitted models per target
Graderfully automated OpenStructure whole-complex and interface scoring · human review: no
Statisticsscore distributions on the three-server common subset
Contaminationexperimental complex structures withheld until the weekly PDB release
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Complex LDDTabsoluteproportionper target distribution with penalties for missing chainsNot reported
Mapped complex LDDTabsoluteproportionper target distribution over mapped chainsremoves the missing-chain stoichiometry penalty
Complex iLDDTabsoluteproportionper target inter-chain-contact distribution with stoichiometry penaltiesNot reported
Mapped complex iLDDTabsoluteproportionper target inter-chain contacts over mapped chainsremoves the missing-chain stoichiometry penalty

No numeric result rows are published yet; the verified protocol remains useful.

Evidence

  • section: Sections 2.2, 2.4.2, and 2.5.1; Figure 2 (Reports the 392-target common subset, three systems, blind stoichiometry setting, model-1 analysis, and four metrics.) — supports /benchmark_version, /model_ids, /scope, /protocol, /metrics