paper · benchmark creator

MoleculeNet: a benchmark for molecular machine learning

Stanford University · DeepChem · 2017-10-31

Relationship layer

Benchmark usage

This table records what the work did with each benchmark before attempting to normalize a run. Partial claims remain visible without being treated as comparable evaluations.

No BenchmarkUse relation is normalized for this legacy work yet. Existing EvaluationRuns remain available below.

Normalized evaluation runs

MoleculeNet1 run

Open benchmark record →

Evaluation run

moleculenet-creator-full

From MoleculeNet: a benchmark for molecular machine learning

moleculenet-original-task-nativevoriginal-2017
Scopefull · n=17
ShotsNot applicable
TurnsNot applicable
System prompt publicNot applicable
Reasoning / effortNot applicable
BrowserNot applicable
InternetNot applicable
DatabasesNot applicable
Code executionNot applicable
ContainerNot reported
External toolsDeepChem featurizers, splitters, conventional ML and graph-based models
Token budgetNot applicable
Time / cost budgetNot reported
TemperatureNot applicable
SeedNot reported
Repeats3
Graderdeterministic dataset-specific scorer · human review: no
StatisticsMean and standard deviation over three independent runs for each dataset-model setting; no cross-dataset normalized total.
Contamination80/10/10 train-validation-test partitions with dataset-specific random, stratified, scaffold, or time splits.
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
MAEabsolutedataset-specific property unitsheld-out examples and endpointsNot reported
RMSEabsolutedataset-specific property unitsheld-out examplesNot reported
ROC-AUCabsoluteareadataset-specific macro endpoint average where applicableNot reported
PRC-AUCabsoluteareadataset-specific endpoint average where applicableNot reported

No numeric result rows are published yet; the verified protocol remains useful.

Evidence

  • table: Methods Sections 3.1-3.5; Results and Discussion; Appendix Performances; Tables 1-3 (Defines all datasets, 80/10/10 partitions, recommended split/metric per collection, featurizers, models, three-run mean/standard-deviation aggregation, and creator results.) — supports /scope, /protocol, /metrics