Scopefull · n=17
ShotsNot applicable
TurnsNot applicable
System prompt publicNot applicable
Reasoning / effortNot applicable
BrowserNot applicable
InternetNot applicable
DatabasesNot applicable
Code executionNot applicable
ContainerNot reported
External toolsDeepChem featurizers, splitters, conventional ML and graph-based models
Token budgetNot applicable
Time / cost budgetNot reported
TemperatureNot applicable
SeedNot reported
Repeats3
Graderdeterministic dataset-specific scorer · human review: no
StatisticsMean and standard deviation over three independent runs for each dataset-model setting; no cross-dataset normalized total.
Contamination80/10/10 train-validation-test partitions with dataset-specific random, stratified, scaffold, or time splits.
Metrics, results, and full protocol
Metrics
| Metric | Kind / baseline | Unit | Aggregation | Threshold / tolerance |
|---|---|---|---|---|
| MAE | absolute | dataset-specific property units | held-out examples and endpoints | Not reported |
| RMSE | absolute | dataset-specific property units | held-out examples | Not reported |
| ROC-AUC | absolute | area | dataset-specific macro endpoint average where applicable | Not reported |
| PRC-AUC | absolute | area | dataset-specific endpoint average where applicable | Not reported |
No numeric result rows are published yet; the verified protocol remains useful.
Evidence
- table: Methods Sections 3.1-3.5; Results and Discussion; Appendix Performances; Tables 1-3 (Defines all datasets, 80/10/10 partitions, recommended split/metric per collection, featurizers, models, three-run mean/standard-deviation aggregation, and creator results.) — supports /scope, /protocol, /metrics