paper · benchmark creator

GuacaMol: Benchmarking Models for de Novo Molecular Design

BenevolentAI · 2019-03-19

Relationship layer

Benchmark usage

This table records what the work did with each benchmark before attempting to normalize a run. Partial claims remain visible without being treated as comparable evaluations.

No BenchmarkUse relation is normalized for this legacy work yet. Existing EvaluationRuns remain available below.

Normalized evaluation runs

GuacaMol1 run

Open benchmark record →

Evaluation run

guacamol-creator-full

From GuacaMol: Benchmarking Models for de Novo Molecular Design

guacamol-v2-task-nativevsuite-v2
Scopefull · n=25
ShotsNot applicable
TurnsNot applicable
System prompt publicNot applicable
Reasoning / effortNot applicable
BrowserNot applicable
InternetNot applicable
DatabasesNot applicable
Code executionNot applicable
Containerofficial Dockerfile
External toolsRDKit 2018.09.1 or newer and FCD 1.1
Token budgetNot applicable
Time / cost budgetNot reported
TemperatureNot reported
SeedNot reported
RepeatsNot reported
Graderdeterministic chemistry scoring functions · human review: no
StatisticsEach formal benchmark returns a normalized score; no registry-wide sum is created.
ContaminationStandardized ChEMBL training data exclude a designated holdout set and highly similar molecules.
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Validityabsoluteproportiongenerated sampleNot reported
Uniquenessabsoluteproportiongenerated sampleNot reported
Noveltyabsoluteproportiongenerated sampleNot reported
KL divergence scoreabsolutenormalized scorephysicochemical descriptor distributionsNot reported
Fréchet ChemNet Distance scoreabsolutenormalized scoregenerated versus reference distributionNot reported
Goal-directed benchmark scoreabsolutenormalized scorebenchmark-specific top generated moleculesNot reported

No numeric result rows are published yet; the verified protocol remains useful.

Evidence

  • section: Sections 2-4 and Tables 1-2; official benchmark_suites.py v2 (Defines both modes, data standardization, baseline generators, and scoring. The pinned implementation establishes the current twenty-problem v2 list.) — supports /scope, /protocol, /metrics