Evaluation registry
Works and run settings
A setting change—scope, prompt, tools, budget, grader, or repeats—creates a separate run. Charts never cross a comparability group.
flip-gb1-low-vs-high-spearmanvoriginal-2021
Evaluated models / systems: BLOSUM62 baseline (FLIP), Convolutional network (FLIP), ESM-untrained (mean), ESM-untrained (mut mean), ESM-untrained (per AA), ESM-1b (mean; FLIP head), ESM-1b (mut mean; FLIP head), ESM-1b (per AA; FLIP head), ESM-1v (mean; FLIP head), ESM-1v (mut mean; FLIP head), ESM-1v (per AA; FLIP head), Levenshtein distance baseline (FLIP), Ridge regression (FLIP)
Scopesubset · n=3644
ShotsNot applicable
TurnsNot applicable
System prompt publicNot applicable
Reasoning / effortNot applicable
BrowserNot applicable
InternetNot applicable
DatabasesNot applicable
Code executionNot applicable
ContainerNot reported
External toolsmodel-specific frozen embeddings or one-hot sequence encoding
Token budgetNot applicable
Time / cost budgetNot reported
TemperatureNot applicable
SeedNot reported
RepeatsNot reported
Graderdeterministic regression scorer · human review: no
StatisticsSpearman correlation over held-out test examples; the main table reports point values without confidence intervals.
ContaminationFitness extrapolation above wild-type fitness.
Metrics, results, and full protocol
Metrics
| Metric | Kind / baseline | Unit | Aggregation | Threshold / tolerance |
|---|
| Spearman correlation | absolute | correlation | computed across all examples in this held-out test split | Not reported |
| Mean squared error | absolute | squared fitness units | mean across all examples in this held-out test split | Not reported |
Results
Evidence
- table: pp. 4–7, Table 2 and Sections 3–4 (Exact train/test counts, split rule, baseline identities, pooling, optimizer, batch sizes, early stopping, and hardware.) — supports /scope, /protocol, /model_ids
- table: p. 8, Table 4 (low-vs-high column) (Published held-out Spearman values; em dashes and NA entries are omitted rather than encoded as zeros.) — supports /metrics, /results
flip-gb1-one-vs-rest-spearmanvoriginal-2021
Evaluated models / systems: BLOSUM62 baseline (FLIP), Convolutional network (FLIP), ESM-untrained (mean), ESM-untrained (mut mean), ESM-untrained (per AA), ESM-1b (mean; FLIP head), ESM-1b (mut mean; FLIP head), ESM-1b (per AA; FLIP head), ESM-1v (mean; FLIP head), ESM-1v (mut mean; FLIP head), ESM-1v (per AA; FLIP head), Levenshtein distance baseline (FLIP), Ridge regression (FLIP)
Scopesubset · n=8704
ShotsNot applicable
TurnsNot applicable
System prompt publicNot applicable
Reasoning / effortNot applicable
BrowserNot applicable
InternetNot applicable
DatabasesNot applicable
Code executionNot applicable
ContainerNot reported
External toolsmodel-specific frozen embeddings or one-hot sequence encoding
Token budgetNot applicable
Time / cost budgetNot reported
TemperatureNot applicable
SeedNot reported
RepeatsNot reported
Graderdeterministic regression scorer · human review: no
StatisticsSpearman correlation over held-out test examples; the main table reports point values without confidence intervals.
ContaminationMutation-depth extrapolation from wild type and single mutants.
Metrics, results, and full protocol
Metrics
| Metric | Kind / baseline | Unit | Aggregation | Threshold / tolerance |
|---|
| Spearman correlation | absolute | correlation | computed across all examples in this held-out test split | Not reported |
| Mean squared error | absolute | squared fitness units | mean across all examples in this held-out test split | Not reported |
Results
Evidence
- table: pp. 4–7, Table 2 and Sections 3–4 (Exact train/test counts, split rule, baseline identities, pooling, optimizer, batch sizes, early stopping, and hardware.) — supports /scope, /protocol, /model_ids
- table: p. 8, Table 4 (1-vs-rest column) (Published held-out Spearman values; em dashes and NA entries are omitted rather than encoded as zeros.) — supports /metrics, /results
flip-gb1-sampled-spearmanvoriginal-2021
Evaluated models / systems: Convolutional network (FLIP), ESM-untrained (per AA), ESM-1b (per AA; FLIP head), ESM-1v (per AA; FLIP head), Ridge regression (FLIP)
Scopesubset · n=1772
ShotsNot applicable
TurnsNot applicable
System prompt publicNot applicable
Reasoning / effortNot applicable
BrowserNot applicable
InternetNot applicable
DatabasesNot applicable
Code executionNot applicable
ContainerNot reported
External toolsmodel-specific frozen embeddings or one-hot sequence encoding
Token budgetNot applicable
Time / cost budgetNot reported
TemperatureNot applicable
SeedNot reported
RepeatsNot reported
Graderdeterministic regression scorer · human review: no
StatisticsSpearman correlation over held-out test examples; the main table reports point values without confidence intervals.
ContaminationRandom split can place closely related variants across train/test and is explicitly described as optimistic.
Metrics, results, and full protocol
Metrics
| Metric | Kind / baseline | Unit | Aggregation | Threshold / tolerance |
|---|
| Spearman correlation | absolute | correlation | computed across all examples in this held-out test split | Not reported |
| Mean squared error | absolute | squared fitness units | mean across all examples in this held-out test split | Not reported |
Results
Evidence
- table: pp. 4–7, Table 2 and Sections 3–4 (Exact train/test counts, split rule, baseline identities, pooling, optimizer, batch sizes, early stopping, and hardware.) — supports /scope, /protocol, /model_ids
- table: p. 10, Table 7 (GB1 column) (Published held-out Spearman values; em dashes and NA entries are omitted rather than encoded as zeros.) — supports /metrics, /results
flip-gb1-three-vs-rest-spearmanvoriginal-2021
Evaluated models / systems: BLOSUM62 baseline (FLIP), Convolutional network (FLIP), ESM-untrained (mean), ESM-untrained (mut mean), ESM-untrained (per AA), ESM-1b (mean; FLIP head), ESM-1b (mut mean; FLIP head), ESM-1b (per AA; FLIP head), ESM-1v (mean; FLIP head), ESM-1v (mut mean; FLIP head), ESM-1v (per AA; FLIP head), Levenshtein distance baseline (FLIP), Ridge regression (FLIP)
Scopesubset · n=5765
ShotsNot applicable
TurnsNot applicable
System prompt publicNot applicable
Reasoning / effortNot applicable
BrowserNot applicable
InternetNot applicable
DatabasesNot applicable
Code executionNot applicable
ContainerNot reported
External toolsmodel-specific frozen embeddings or one-hot sequence encoding
Token budgetNot applicable
Time / cost budgetNot reported
TemperatureNot applicable
SeedNot reported
RepeatsNot reported
Graderdeterministic regression scorer · human review: no
StatisticsSpearman correlation over held-out test examples; the main table reports point values without confidence intervals.
ContaminationMutation-depth extrapolation to four simultaneous mutations.
Metrics, results, and full protocol
Metrics
| Metric | Kind / baseline | Unit | Aggregation | Threshold / tolerance |
|---|
| Spearman correlation | absolute | correlation | computed across all examples in this held-out test split | Not reported |
| Mean squared error | absolute | squared fitness units | mean across all examples in this held-out test split | Not reported |
Results
Evidence
- table: pp. 4–7, Table 2 and Sections 3–4 (Exact train/test counts, split rule, baseline identities, pooling, optimizer, batch sizes, early stopping, and hardware.) — supports /scope, /protocol, /model_ids
- table: p. 8, Table 4 (3-vs-rest column) (Published held-out Spearman values; em dashes and NA entries are omitted rather than encoded as zeros.) — supports /metrics, /results
flip-gb1-two-vs-rest-spearmanvoriginal-2021
Evaluated models / systems: BLOSUM62 baseline (FLIP), Convolutional network (FLIP), ESM-untrained (mean), ESM-untrained (mut mean), ESM-untrained (per AA), ESM-1b (mean; FLIP head), ESM-1b (mut mean; FLIP head), ESM-1b (per AA; FLIP head), ESM-1v (mean; FLIP head), ESM-1v (mut mean; FLIP head), ESM-1v (per AA; FLIP head), Levenshtein distance baseline (FLIP), Ridge regression (FLIP)
Scopesubset · n=8306
ShotsNot applicable
TurnsNot applicable
System prompt publicNot applicable
Reasoning / effortNot applicable
BrowserNot applicable
InternetNot applicable
DatabasesNot applicable
Code executionNot applicable
ContainerNot reported
External toolsmodel-specific frozen embeddings or one-hot sequence encoding
Token budgetNot applicable
Time / cost budgetNot reported
TemperatureNot applicable
SeedNot reported
RepeatsNot reported
Graderdeterministic regression scorer · human review: no
StatisticsSpearman correlation over held-out test examples; the main table reports point values without confidence intervals.
ContaminationMutation-depth extrapolation above two mutations.
Metrics, results, and full protocol
Metrics
| Metric | Kind / baseline | Unit | Aggregation | Threshold / tolerance |
|---|
| Spearman correlation | absolute | correlation | computed across all examples in this held-out test split | Not reported |
| Mean squared error | absolute | squared fitness units | mean across all examples in this held-out test split | Not reported |
Results
Evidence
- table: pp. 4–7, Table 2 and Sections 3–4 (Exact train/test counts, split rule, baseline identities, pooling, optimizer, batch sizes, early stopping, and hardware.) — supports /scope, /protocol, /model_ids
- table: p. 8, Table 4 (2-vs-rest column) (Published held-out Spearman values; em dashes and NA entries are omitted rather than encoded as zeros.) — supports /metrics, /results