paper · benchmark creator

CASP16 Protein Monomer Structure Prediction Assessment

University of Texas Southwestern Medical Center · University of California Davis · 2025-08-17

Relationship layer

Benchmark usage

This table records what the work did with each benchmark before attempting to normalize a run. Partial claims remain visible without being treated as comparable evaluations.

No BenchmarkUse relation is normalized for this legacy work yet. Existing EvaluationRuns remain available below.

Normalized evaluation runs

CASP Protein Monomers1 run

Open benchmark record →

Evaluation run

casp16-monomer-regular-official

From CASP16 Protein Monomer Structure Prediction Assessment

casp16-monomer-regular-officialvCASP16
Scopefull · n=54
ShotsNot applicable
TurnsNot applicable
System prompt publicNot applicable
Reasoning / effortNot applicable
Browsermethod-specific
Internetmethod-specific
Databasesmethod-specific
Code executionallowed
ContainerNot reported
External toolsmethod-specific
Token budgetNot applicable
Time / cost budgetapproximately three weeks per target for regular groups
TemperatureNot applicable
SeedNot reported
Repeatsup to five submitted models per target
Graderofficial CASP structure-comparison pipeline plus independent monomer assessor team · human review: yes
Statisticsper-evaluation-unit scores converted to group ranking summaries including cumulative clipped z-scores
Contaminationexperimental structures withheld until prediction closes
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
GDT_TSabsolutescorecomputed per evaluation unit; official group rankings use first/best model views and z-score summariesNot reported
GDT_HAabsolutescoreper evaluation unitNot reported
lDDTabsoluteproportionper evaluation unitNot reported
TM-scoreabsoluteproportionper evaluation unitNot reported
SUM Z-score (> -2.0)absolutez-score sumsum across eligible evaluation units after clipping values below -2.0Not reported

No numeric result rows are published yet; the verified protocol remains useful.

Evidence

  • section: Results and Discussion; performance evaluation and ranking (54 evaluation units, model-1/best-model analyses, metrics, and assessor interpretation.) — supports /benchmark_version, /scope, /protocol/grader, /protocol/statistical, /metrics
  • section: Registration; Targets; Model submission/format; Assessment (Blind targets, regular versus server settings, deadlines, and independent assessment.) — supports /protocol/shots, /protocol/turns, /protocol/system_prompt_public, /protocol/reasoning, /protocol/tools, /protocol/token_budget, /protocol/time_budget, /protocol/temperature, /protocol/seed, /protocol/repeats, /protocol/contamination