Evaluation registry
Works and run settings
A setting change—scope, prompt, tools, budget, grader, or repeats—creates a separate run. Charts never cross a comparability group.
flip-aav-des-mut-spearmanvoriginal-2021
Evaluated models / systems: Convolutional network (FLIP), ESM-untrained (mean), ESM-untrained (mut mean), ESM-1b (mean; FLIP head), ESM-1b (mut mean; FLIP head), ESM-1v (mean; FLIP head), ESM-1v (mut mean; FLIP head), Levenshtein distance baseline (FLIP), Ridge regression (FLIP)
Scopesubset · n=82583
ShotsNot applicable
TurnsNot applicable
System prompt publicNot applicable
Reasoning / effortNot applicable
BrowserNot applicable
InternetNot applicable
DatabasesNot applicable
Code executionNot applicable
ContainerNot reported
External toolsmodel-specific frozen embeddings or one-hot sequence encoding
Token budgetNot applicable
Time / cost budgetNot reported
TemperatureNot applicable
SeedNot reported
RepeatsNot reported
Graderdeterministic regression scorer · human review: no
StatisticsSpearman correlation over held-out test examples; the main table reports point values without confidence intervals.
ContaminationPool shift: model-designed variants are training data and sampled variants are held out.
Metrics, results, and full protocol
Metrics
| Metric | Kind / baseline | Unit | Aggregation | Threshold / tolerance |
|---|
| Spearman correlation | absolute | correlation | computed across all examples in this held-out test split | Not reported |
| Mean squared error | absolute | squared fitness units | mean across all examples in this held-out test split | Not reported |
Results
Evidence
- table: pp. 4–7, Table 2 and Sections 3–4 (Exact train/test counts, split rule, baseline identities, pooling, optimizer, batch sizes, early stopping, and hardware.) — supports /scope, /protocol, /model_ids
- table: p. 9, Table 5 (Des-Mut column) (Published held-out Spearman values; em dashes and NA entries are omitted rather than encoded as zeros.) — supports /metrics, /results
flip-aav-low-vs-high-spearmanvoriginal-2021
Evaluated models / systems: Convolutional network (FLIP), ESM-untrained (mean), ESM-untrained (mut mean), ESM-untrained (per AA), ESM-1b (mean; FLIP head), ESM-1b (mut mean; FLIP head), ESM-1b (per AA; FLIP head), ESM-1v (mean; FLIP head), ESM-1v (mut mean; FLIP head), ESM-1v (per AA; FLIP head), Levenshtein distance baseline (FLIP), Ridge regression (FLIP)
Scopesubset · n=35037
ShotsNot applicable
TurnsNot applicable
System prompt publicNot applicable
Reasoning / effortNot applicable
BrowserNot applicable
InternetNot applicable
DatabasesNot applicable
Code executionNot applicable
ContainerNot reported
External toolsmodel-specific frozen embeddings or one-hot sequence encoding
Token budgetNot applicable
Time / cost budgetNot reported
TemperatureNot applicable
SeedNot reported
RepeatsNot reported
Graderdeterministic regression scorer · human review: no
StatisticsSpearman correlation over held-out test examples; the main table reports point values without confidence intervals.
ContaminationFitness extrapolation above wild-type fitness.
Metrics, results, and full protocol
Metrics
| Metric | Kind / baseline | Unit | Aggregation | Threshold / tolerance |
|---|
| Spearman correlation | absolute | correlation | computed across all examples in this held-out test split | Not reported |
| Mean squared error | absolute | squared fitness units | mean across all examples in this held-out test split | Not reported |
Results
Evidence
- table: pp. 4–7, Table 2 and Sections 3–4 (Exact train/test counts, split rule, baseline identities, pooling, optimizer, batch sizes, early stopping, and hardware.) — supports /scope, /protocol, /model_ids
- table: p. 9, Table 5 (low-vs-high column) (Published held-out Spearman values; em dashes and NA entries are omitted rather than encoded as zeros.) — supports /metrics, /results
flip-aav-mut-des-spearmanvoriginal-2021
Evaluated models / systems: Convolutional network (FLIP), ESM-untrained (mean), ESM-untrained (mut mean), ESM-untrained (per AA), ESM-1b (mean; FLIP head), ESM-1b (mut mean; FLIP head), ESM-1b (per AA; FLIP head), ESM-1v (mean; FLIP head), ESM-1v (mut mean; FLIP head), ESM-1v (per AA; FLIP head), Levenshtein distance baseline (FLIP), Ridge regression (FLIP)
Scopesubset · n=201426
ShotsNot applicable
TurnsNot applicable
System prompt publicNot applicable
Reasoning / effortNot applicable
BrowserNot applicable
InternetNot applicable
DatabasesNot applicable
Code executionNot applicable
ContainerNot reported
External toolsmodel-specific frozen embeddings or one-hot sequence encoding
Token budgetNot applicable
Time / cost budgetNot reported
TemperatureNot applicable
SeedNot reported
RepeatsNot reported
Graderdeterministic regression scorer · human review: no
StatisticsSpearman correlation over held-out test examples; the main table reports point values without confidence intervals.
ContaminationPool shift: sampled variants are training data and model-designed variants are held out.
Metrics, results, and full protocol
Metrics
| Metric | Kind / baseline | Unit | Aggregation | Threshold / tolerance |
|---|
| Spearman correlation | absolute | correlation | computed across all examples in this held-out test split | Not reported |
| Mean squared error | absolute | squared fitness units | mean across all examples in this held-out test split | Not reported |
Results
Evidence
- table: pp. 4–7, Table 2 and Sections 3–4 (Exact train/test counts, split rule, baseline identities, pooling, optimizer, batch sizes, early stopping, and hardware.) — supports /scope, /protocol, /model_ids
- table: p. 9, Table 5 (Mut-Des column) (Published held-out Spearman values; em dashes and NA entries are omitted rather than encoded as zeros.) — supports /metrics, /results
flip-aav-one-vs-rest-spearmanvoriginal-2021
Evaluated models / systems: Convolutional network (FLIP), ESM-untrained (mean), ESM-untrained (mut mean), ESM-untrained (per AA), ESM-1b (mean; FLIP head), ESM-1b (mut mean; FLIP head), ESM-1b (per AA; FLIP head), ESM-1v (mean; FLIP head), ESM-1v (mut mean; FLIP head), ESM-1v (per AA; FLIP head), Levenshtein distance baseline (FLIP), Ridge regression (FLIP)
Scopesubset · n=81413
ShotsNot applicable
TurnsNot applicable
System prompt publicNot applicable
Reasoning / effortNot applicable
BrowserNot applicable
InternetNot applicable
DatabasesNot applicable
Code executionNot applicable
ContainerNot reported
External toolsmodel-specific frozen embeddings or one-hot sequence encoding
Token budgetNot applicable
Time / cost budgetNot reported
TemperatureNot applicable
SeedNot reported
RepeatsNot reported
Graderdeterministic regression scorer · human review: no
StatisticsSpearman correlation over held-out test examples; the main table reports point values without confidence intervals.
ContaminationMutation-depth extrapolation; the held-out set excludes wild type and single-mutant training examples.
Metrics, results, and full protocol
Metrics
| Metric | Kind / baseline | Unit | Aggregation | Threshold / tolerance |
|---|
| Spearman correlation | absolute | correlation | computed across all examples in this held-out test split | Not reported |
| Mean squared error | absolute | squared fitness units | mean across all examples in this held-out test split | Not reported |
Results
Evidence
- table: pp. 4–7, Table 2 and Sections 3–4 (Exact train/test counts, split rule, baseline identities, pooling, optimizer, batch sizes, early stopping, and hardware.) — supports /scope, /protocol, /model_ids
- table: p. 9, Table 5 (1-vs-rest column) (Published held-out Spearman values; em dashes and NA entries are omitted rather than encoded as zeros.) — supports /metrics, /results
flip-aav-sampled-spearmanvoriginal-2021
Evaluated models / systems: Convolutional network (FLIP), ESM-untrained (per AA), ESM-1b (per AA; FLIP head), ESM-1v (per AA; FLIP head), Ridge regression (FLIP)
Scopesubset · n=16517
ShotsNot applicable
TurnsNot applicable
System prompt publicNot applicable
Reasoning / effortNot applicable
BrowserNot applicable
InternetNot applicable
DatabasesNot applicable
Code executionNot applicable
ContainerNot reported
External toolsmodel-specific frozen embeddings or one-hot sequence encoding
Token budgetNot applicable
Time / cost budgetNot reported
TemperatureNot applicable
SeedNot reported
RepeatsNot reported
Graderdeterministic regression scorer · human review: no
StatisticsSpearman correlation over held-out test examples; the main table reports point values without confidence intervals.
ContaminationRandom split can place closely related variants across train/test and is explicitly described as optimistic.
Metrics, results, and full protocol
Metrics
| Metric | Kind / baseline | Unit | Aggregation | Threshold / tolerance |
|---|
| Spearman correlation | absolute | correlation | computed across all examples in this held-out test split | Not reported |
| Mean squared error | absolute | squared fitness units | mean across all examples in this held-out test split | Not reported |
Results
Evidence
- table: pp. 4–7, Table 2 and Sections 3–4 (Exact train/test counts, split rule, baseline identities, pooling, optimizer, batch sizes, early stopping, and hardware.) — supports /scope, /protocol, /model_ids
- table: p. 10, Table 7 (AAV column) (Published held-out Spearman values; em dashes and NA entries are omitted rather than encoded as zeros.) — supports /metrics, /results
flip-aav-seven-vs-rest-spearmanvoriginal-2021
Evaluated models / systems: Convolutional network (FLIP), ESM-untrained (mean), ESM-untrained (mut mean), ESM-untrained (per AA), ESM-1b (mean; FLIP head), ESM-1b (mut mean; FLIP head), ESM-1b (per AA; FLIP head), ESM-1v (mean; FLIP head), ESM-1v (mut mean; FLIP head), ESM-1v (per AA; FLIP head), Levenshtein distance baseline (FLIP), Ridge regression (FLIP)
Scopesubset · n=12581
ShotsNot applicable
TurnsNot applicable
System prompt publicNot applicable
Reasoning / effortNot applicable
BrowserNot applicable
InternetNot applicable
DatabasesNot applicable
Code executionNot applicable
ContainerNot reported
External toolsmodel-specific frozen embeddings or one-hot sequence encoding
Token budgetNot applicable
Time / cost budgetNot reported
TemperatureNot applicable
SeedNot reported
RepeatsNot reported
Graderdeterministic regression scorer · human review: no
StatisticsSpearman correlation over held-out test examples; the main table reports point values without confidence intervals.
ContaminationMutation-depth extrapolation above seven changes.
Metrics, results, and full protocol
Metrics
| Metric | Kind / baseline | Unit | Aggregation | Threshold / tolerance |
|---|
| Spearman correlation | absolute | correlation | computed across all examples in this held-out test split | Not reported |
| Mean squared error | absolute | squared fitness units | mean across all examples in this held-out test split | Not reported |
Results
Evidence
- table: pp. 4–7, Table 2 and Sections 3–4 (Exact train/test counts, split rule, baseline identities, pooling, optimizer, batch sizes, early stopping, and hardware.) — supports /scope, /protocol, /model_ids
- table: p. 9, Table 5 (7-vs-rest column) (Published held-out Spearman values; em dashes and NA entries are omitted rather than encoded as zeros.) — supports /metrics, /results
flip-aav-two-vs-rest-spearmanvoriginal-2021
Evaluated models / systems: Convolutional network (FLIP), ESM-untrained (mean), ESM-untrained (mut mean), ESM-untrained (per AA), ESM-1b (mean; FLIP head), ESM-1b (mut mean; FLIP head), ESM-1b (per AA; FLIP head), ESM-1v (mean; FLIP head), ESM-1v (mut mean; FLIP head), ESM-1v (per AA; FLIP head), Levenshtein distance baseline (FLIP), Ridge regression (FLIP)
Scopesubset · n=50776
ShotsNot applicable
TurnsNot applicable
System prompt publicNot applicable
Reasoning / effortNot applicable
BrowserNot applicable
InternetNot applicable
DatabasesNot applicable
Code executionNot applicable
ContainerNot reported
External toolsmodel-specific frozen embeddings or one-hot sequence encoding
Token budgetNot applicable
Time / cost budgetNot reported
TemperatureNot applicable
SeedNot reported
RepeatsNot reported
Graderdeterministic regression scorer · human review: no
StatisticsSpearman correlation over held-out test examples; the main table reports point values without confidence intervals.
ContaminationMutation-depth extrapolation above two changes.
Metrics, results, and full protocol
Metrics
| Metric | Kind / baseline | Unit | Aggregation | Threshold / tolerance |
|---|
| Spearman correlation | absolute | correlation | computed across all examples in this held-out test split | Not reported |
| Mean squared error | absolute | squared fitness units | mean across all examples in this held-out test split | Not reported |
Results
Evidence
- table: pp. 4–7, Table 2 and Sections 3–4 (Exact train/test counts, split rule, baseline identities, pooling, optimizer, batch sizes, early stopping, and hardware.) — supports /scope, /protocol, /model_ids
- table: p. 9, Table 5 (2-vs-rest column) (Published held-out Spearman values; em dashes and NA entries are omitted rather than encoded as zeros.) — supports /metrics, /results