paper · benchmark creator

Assessment of Protein Complex Predictions in CASP16: Are We Making Progress?

University of Texas Southwestern Medical Center · University of California Davis · Stanford University · 2025-10-31

Relationship layer

Benchmark usage

This table records what the work did with each benchmark before attempting to normalize a run. Partial claims remain visible without being treated as comparable evaluations.

No BenchmarkUse relation is normalized for this legacy work yet. Existing EvaluationRuns remain available below.

Normalized evaluation runs

CASP Protein Multimers1 run

Open benchmark record →

casp16-multimer-phase1-regularvCASP16
Scopefull · n=40
ShotsNot applicable
TurnsNot applicable
System prompt publicNot applicable
Reasoning / effortNot applicable
Browsermethod-specific
Internetmethod-specific
Databasesmethod-specific
Code executionallowed
ContainerNot reported
External toolsmethod-specific
Token budgetNot applicable
Time / cost budgetapproximately three weeks per target for regular groups
TemperatureNot applicable
SeedNot reported
Repeatsup to five standard models per target
Graderofficial CASP/OpenStructure and CAPRI-compatible scoring plus independent complex assessor team · human review: yes
Statisticscumulative per-target z-scores and head-to-head comparisons; rankings reported for first and best models
Contaminationexperimental complex structures withheld until prediction closes
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
DockQabsoluteproportionper target followed by cumulative z-score ranking0.5
TM-scoreabsoluteproportionper target followed by cumulative z-score rankingNot reported
lDDTabsoluteproportionper target followed by cumulative z-score rankingNot reported
ICSabsoluteF1 scoreper target followed by cumulative z-score rankingNot reported
IPSabsoluteJaccard scoreper target followed by cumulative z-score rankingNot reported
QSbestabsoluteJaccard scoreper target followed by cumulative z-score rankingNot reported

No numeric result rows are published yet; the verified protocol remains useful.

Evidence

  • section: Overview of targets; Phase 1 Ranking; Performance Evaluation and Ranking (40 Phase-1 targets, supplied stoichiometry, overall/interface metrics, first/best model analyses, and group ranking aggregation.) — supports /benchmark_version, /scope, /protocol/grader, /protocol/statistical, /metrics
  • section: Protein Complexes; Registration; Model submission/format; Assessment (Blind regular-group protocol, participant method freedom, deadlines, and assessors.) — supports /protocol/shots, /protocol/turns, /protocol/system_prompt_public, /protocol/reasoning, /protocol/tools, /protocol/token_budget, /protocol/time_budget, /protocol/temperature, /protocol/seed, /protocol/repeats, /protocol/contamination