FLIP AAV
Seven supervised splits over sampled and machine-designed AAV2 VP-1 capsid variants, measuring generalization across mutation depth, fitness, and sampled-versus-designed pools.
7 evaluation run(s)
suite · audited · verified 2026-07-21
A supervised protein sequence-to-fitness benchmark that turns three experimental landscapes into 15 biologically motivated dataset splits for testing generalization in protein engineering.
Benchmark definition
| Version | Status | Release / as-of | Total | Formal tracks |
|---|---|---|---|---|
original-2021flip-original-2021 | current | 2021-10-11 | 15 (dataset-and-split benchmark tasks) | flip-aav, flip-gb1, flip-meltome |
| ID | Count | Basis | Partition? | Notes |
|---|---|---|---|---|
AAV landscape splitsflip-aav-tasks | 7 | dataset-and-split tasks | Exclusive & exhaustive | Six active comparison splits plus one sampled split used mainly for discourse. |
GB1 landscape splitsflip-gb1-tasks | 5 | dataset-and-split tasks | Exclusive & exhaustive | Four active comparison splits plus one sampled split used mainly for discourse. |
Meltome thermostability splitsflip-meltome-tasks | 3 | dataset-and-split tasks | Exclusive & exhaustive | Mixed, Human, and Human-cell are all active. |
Active performance-comparison splitsflip-active-comparison-tasks | 13 | dataset-and-split tasks | No | Current official repository semaphore marks all original splits active except the two sampled splits. |
Sampled discourse splitsflip-discourse-sampled-tasks | 2 | dataset-and-split tasks | No | AAV Sampled and GB1 Sampled are orange: the repository warns against performance comparisons because random sampling can overestimate performance. |
Seven supervised splits over sampled and machine-designed AAV2 VP-1 capsid variants, measuring generalization across mutation depth, fitness, and sampled-versus-designed pools.
7 evaluation run(s)
Five supervised splits over a downsampled, highly epistatic four-site GB1 immunoglobulin-binding landscape, designed to test mutation-depth and low-to-high-fitness generalization.
5 evaluation run(s)
Three supervised sequence-to-melting-temperature splits spanning all species, human proteins, and a single human cell line, with sequence-cluster-aware train/test separation.
3 evaluation run(s)
Scientific Task Atlas
complete for original-2021. FLIP evaluates sequence-to-fitness prediction; it is not a sequence-generation benchmark.
| Scientific task | Coverage | Count | Mapping | Evidence |
|---|---|---|---|---|
| Protein fitness prediction | explicitly-in-scope | 15 tasks Dataset-and-split benchmark tasks. | official-taxonomy high confidence | flip-evidence-paper-definitionAll formal tasks are supervised fitness or phenotype prediction. |
| Protein mutation-effect prediction | explicitly-in-scope | 15 tasks Dataset-and-split benchmark tasks. | official-taxonomy high confidence | flip-evidence-paper-definitionThe same tasks evaluate generalization across mutated sequence landscapes; claims overlap and are never summed. |
| Domain | Coverage | Count | Interpretation |
|---|---|---|---|
| Protein-protein binding | explicitly-in-scope | 5 | All five GB1 dataset-split tasks predict the fitness of an immunoglobulin-binding protein domain; this is task count, not a distinct set of five assays. |
| Protein design | explicitly-in-scope | Not reported | All tasks are motivated by protein engineering, but FLIP evaluates prediction/regression rather than sequence generation or optimization and does not publish a separate design-task count. |
Evaluation registry
A setting change—scope, prompt, tools, budget, grader, or repeats—creates a separate run. Charts never cross a comparability group.
Evaluation run
From FLIP: Benchmark tasks in fitness landscape inference for proteins
Evaluated models / systems: Convolutional network (FLIP), ESM-untrained (mean), ESM-untrained (mut mean), ESM-1b (mean; FLIP head), ESM-1b (mut mean; FLIP head), ESM-1v (mean; FLIP head), ESM-1v (mut mean; FLIP head), Levenshtein distance baseline (FLIP), Ridge regression (FLIP)
| Metric | Kind / baseline | Unit | Aggregation | Threshold / tolerance |
|---|---|---|---|---|
| Spearman correlation | absolute | correlation | computed across all examples in this held-out test split | Not reported |
| Mean squared error | absolute | squared fitness units | mean across all examples in this held-out test split | Not reported |
| Model | Metric | Value | n |
|---|---|---|---|
| ESM-1b (mean; FLIP head) | Spearman correlation | 0.59 correlation Held-out test-set Spearman from the creator paper. | 82583 |
| ESM-1b (mut mean; FLIP head) | Spearman correlation | 0.7 correlation Held-out test-set Spearman from the creator paper. | 82583 |
| ESM-1v (mean; FLIP head) | Spearman correlation | 0.44 correlation Held-out test-set Spearman from the creator paper. | 82583 |
| ESM-1v (mut mean; FLIP head) | Spearman correlation | 0.71 correlation Held-out test-set Spearman from the creator paper. | 82583 |
| ESM-untrained (mean) | Spearman correlation | 0.34 correlation Held-out test-set Spearman from the creator paper. | 82583 |
| ESM-untrained (mut mean) | Spearman correlation | 0.64 correlation Held-out test-set Spearman from the creator paper. | 82583 |
| Ridge regression (FLIP) | Spearman correlation | 0.53 correlation Held-out test-set Spearman from the creator paper. | 82583 |
| Convolutional network (FLIP) | Spearman correlation | 0.75 correlation Held-out test-set Spearman from the creator paper. | 82583 |
| Levenshtein distance baseline (FLIP) | Spearman correlation | -0.07 correlation Held-out test-set Spearman from the creator paper. | 82583 |
Evaluation run
From FLIP: Benchmark tasks in fitness landscape inference for proteins
Evaluated models / systems: Convolutional network (FLIP), ESM-untrained (mean), ESM-untrained (mut mean), ESM-untrained (per AA), ESM-1b (mean; FLIP head), ESM-1b (mut mean; FLIP head), ESM-1b (per AA; FLIP head), ESM-1v (mean; FLIP head), ESM-1v (mut mean; FLIP head), ESM-1v (per AA; FLIP head), Levenshtein distance baseline (FLIP), Ridge regression (FLIP)
| Metric | Kind / baseline | Unit | Aggregation | Threshold / tolerance |
|---|---|---|---|---|
| Spearman correlation | absolute | correlation | computed across all examples in this held-out test split | Not reported |
| Mean squared error | absolute | squared fitness units | mean across all examples in this held-out test split | Not reported |
| Model | Metric | Value | n |
|---|---|---|---|
| ESM-1b (per AA; FLIP head) | Spearman correlation | 0.39 correlation Held-out test-set Spearman from the creator paper. | 35037 |
| ESM-1b (mean; FLIP head) | Spearman correlation | 0.18 correlation Held-out test-set Spearman from the creator paper. | 35037 |
| ESM-1b (mut mean; FLIP head) | Spearman correlation | 0.33 correlation Held-out test-set Spearman from the creator paper. | 35037 |
| ESM-1v (per AA; FLIP head) | Spearman correlation | 0.34 correlation Held-out test-set Spearman from the creator paper. | 35037 |
| ESM-1v (mean; FLIP head) | Spearman correlation | 0.2 correlation Held-out test-set Spearman from the creator paper. | 35037 |
| ESM-1v (mut mean; FLIP head) | Spearman correlation | 0.31 correlation Held-out test-set Spearman from the creator paper. | 35037 |
| ESM-untrained (per AA) | Spearman correlation | 0.08 correlation Held-out test-set Spearman from the creator paper. | 35037 |
| ESM-untrained (mean) | Spearman correlation | 0.22 correlation Held-out test-set Spearman from the creator paper. | 35037 |
| ESM-untrained (mut mean) | Spearman correlation | 0.24 correlation Held-out test-set Spearman from the creator paper. | 35037 |
| Ridge regression (FLIP) | Spearman correlation | 0.12 correlation Held-out test-set Spearman from the creator paper. | 35037 |
| Convolutional network (FLIP) | Spearman correlation | 0.34 correlation Held-out test-set Spearman from the creator paper. | 35037 |
| Levenshtein distance baseline (FLIP) | Spearman correlation | 0.25 correlation Held-out test-set Spearman from the creator paper. | 35037 |
Evaluation run
From FLIP: Benchmark tasks in fitness landscape inference for proteins
Evaluated models / systems: Convolutional network (FLIP), ESM-untrained (mean), ESM-untrained (mut mean), ESM-untrained (per AA), ESM-1b (mean; FLIP head), ESM-1b (mut mean; FLIP head), ESM-1b (per AA; FLIP head), ESM-1v (mean; FLIP head), ESM-1v (mut mean; FLIP head), ESM-1v (per AA; FLIP head), Levenshtein distance baseline (FLIP), Ridge regression (FLIP)
| Metric | Kind / baseline | Unit | Aggregation | Threshold / tolerance |
|---|---|---|---|---|
| Spearman correlation | absolute | correlation | computed across all examples in this held-out test split | Not reported |
| Mean squared error | absolute | squared fitness units | mean across all examples in this held-out test split | Not reported |
| Model | Metric | Value | n |
|---|---|---|---|
| ESM-1b (per AA; FLIP head) | Spearman correlation | 0.76 correlation Held-out test-set Spearman from the creator paper. | 201426 |
| ESM-1b (mean; FLIP head) | Spearman correlation | 0.63 correlation Held-out test-set Spearman from the creator paper. | 201426 |
| ESM-1b (mut mean; FLIP head) | Spearman correlation | 0.7 correlation Held-out test-set Spearman from the creator paper. | 201426 |
| ESM-1v (per AA; FLIP head) | Spearman correlation | 0.79 correlation Held-out test-set Spearman from the creator paper. | 201426 |
| ESM-1v (mean; FLIP head) | Spearman correlation | 0.55 correlation Held-out test-set Spearman from the creator paper. | 201426 |
| ESM-1v (mut mean; FLIP head) | Spearman correlation | 0.7 correlation Held-out test-set Spearman from the creator paper. | 201426 |
| ESM-untrained (per AA) | Spearman correlation | 0.56 correlation Held-out test-set Spearman from the creator paper. | 201426 |
| ESM-untrained (mean) | Spearman correlation | 0.27 correlation Held-out test-set Spearman from the creator paper. | 201426 |
| ESM-untrained (mut mean) | Spearman correlation | 0.62 correlation Held-out test-set Spearman from the creator paper. | 201426 |
| Ridge regression (FLIP) | Spearman correlation | 0.64 correlation Held-out test-set Spearman from the creator paper. | 201426 |
| Convolutional network (FLIP) | Spearman correlation | 0.71 correlation Held-out test-set Spearman from the creator paper. | 201426 |
| Levenshtein distance baseline (FLIP) | Spearman correlation | 0.6 correlation Held-out test-set Spearman from the creator paper. | 201426 |
Evaluation run
From FLIP: Benchmark tasks in fitness landscape inference for proteins
Evaluated models / systems: Convolutional network (FLIP), ESM-untrained (mean), ESM-untrained (mut mean), ESM-untrained (per AA), ESM-1b (mean; FLIP head), ESM-1b (mut mean; FLIP head), ESM-1b (per AA; FLIP head), ESM-1v (mean; FLIP head), ESM-1v (mut mean; FLIP head), ESM-1v (per AA; FLIP head), Levenshtein distance baseline (FLIP), Ridge regression (FLIP)
| Metric | Kind / baseline | Unit | Aggregation | Threshold / tolerance |
|---|---|---|---|---|
| Spearman correlation | absolute | correlation | computed across all examples in this held-out test split | Not reported |
| Mean squared error | absolute | squared fitness units | mean across all examples in this held-out test split | Not reported |
| Model | Metric | Value | n |
|---|---|---|---|
| ESM-1b (per AA; FLIP head) | Spearman correlation | 0.03 correlation Held-out test-set Spearman from the creator paper. | 81413 |
| ESM-1b (mean; FLIP head) | Spearman correlation | 0.04 correlation Held-out test-set Spearman from the creator paper. | 81413 |
| ESM-1b (mut mean; FLIP head) | Spearman correlation | 0.31 correlation Held-out test-set Spearman from the creator paper. | 81413 |
| ESM-1v (per AA; FLIP head) | Spearman correlation | 0.1 correlation Held-out test-set Spearman from the creator paper. | 81413 |
| ESM-1v (mean; FLIP head) | Spearman correlation | 0.18 correlation Held-out test-set Spearman from the creator paper. | 81413 |
| ESM-1v (mut mean; FLIP head) | Spearman correlation | 0.44 correlation Held-out test-set Spearman from the creator paper. | 81413 |
| ESM-untrained (per AA) | Spearman correlation | 0.18 correlation Held-out test-set Spearman from the creator paper. | 81413 |
| ESM-untrained (mean) | Spearman correlation | 0.01 correlation Held-out test-set Spearman from the creator paper. | 81413 |
| ESM-untrained (mut mean) | Spearman correlation | 0.26 correlation Held-out test-set Spearman from the creator paper. | 81413 |
| Ridge regression (FLIP) | Spearman correlation | 0.22 correlation Held-out test-set Spearman from the creator paper. | 81413 |
| Convolutional network (FLIP) | Spearman correlation | 0.48 correlation Held-out test-set Spearman from the creator paper. | 81413 |
| Levenshtein distance baseline (FLIP) | Spearman correlation | -0.11 correlation Held-out test-set Spearman from the creator paper. | 81413 |
Evaluation run
From FLIP: Benchmark tasks in fitness landscape inference for proteins
Evaluated models / systems: Convolutional network (FLIP), ESM-untrained (per AA), ESM-1b (per AA; FLIP head), ESM-1v (per AA; FLIP head), Ridge regression (FLIP)
| Metric | Kind / baseline | Unit | Aggregation | Threshold / tolerance |
|---|---|---|---|---|
| Spearman correlation | absolute | correlation | computed across all examples in this held-out test split | Not reported |
| Mean squared error | absolute | squared fitness units | mean across all examples in this held-out test split | Not reported |
| Model | Metric | Value | n |
|---|---|---|---|
| ESM-1b (per AA; FLIP head) | Spearman correlation | 0.9 correlation Held-out test-set Spearman from the creator paper. | 16517 |
| ESM-1v (per AA; FLIP head) | Spearman correlation | 0.92 correlation Held-out test-set Spearman from the creator paper. | 16517 |
| ESM-untrained (per AA) | Spearman correlation | 0.78 correlation Held-out test-set Spearman from the creator paper. | 16517 |
| Ridge regression (FLIP) | Spearman correlation | 0.83 correlation Held-out test-set Spearman from the creator paper. | 16517 |
| Convolutional network (FLIP) | Spearman correlation | 0.92 correlation Held-out test-set Spearman from the creator paper. | 16517 |
Evaluation run
From FLIP: Benchmark tasks in fitness landscape inference for proteins
Evaluated models / systems: Convolutional network (FLIP), ESM-untrained (mean), ESM-untrained (mut mean), ESM-untrained (per AA), ESM-1b (mean; FLIP head), ESM-1b (mut mean; FLIP head), ESM-1b (per AA; FLIP head), ESM-1v (mean; FLIP head), ESM-1v (mut mean; FLIP head), ESM-1v (per AA; FLIP head), Levenshtein distance baseline (FLIP), Ridge regression (FLIP)
| Metric | Kind / baseline | Unit | Aggregation | Threshold / tolerance |
|---|---|---|---|---|
| Spearman correlation | absolute | correlation | computed across all examples in this held-out test split | Not reported |
| Mean squared error | absolute | squared fitness units | mean across all examples in this held-out test split | Not reported |
| Model | Metric | Value | n |
|---|---|---|---|
| ESM-1b (per AA; FLIP head) | Spearman correlation | 0.65 correlation Held-out test-set Spearman from the creator paper. | 12581 |
| ESM-1b (mean; FLIP head) | Spearman correlation | 0.46 correlation Held-out test-set Spearman from the creator paper. | 12581 |
| ESM-1b (mut mean; FLIP head) | Spearman correlation | 0.61 correlation Held-out test-set Spearman from the creator paper. | 12581 |
| ESM-1v (per AA; FLIP head) | Spearman correlation | 0.7 correlation Held-out test-set Spearman from the creator paper. | 12581 |
| ESM-1v (mean; FLIP head) | Spearman correlation | 0.45 correlation Held-out test-set Spearman from the creator paper. | 12581 |
| ESM-1v (mut mean; FLIP head) | Spearman correlation | 0.64 correlation Held-out test-set Spearman from the creator paper. | 12581 |
| ESM-untrained (per AA) | Spearman correlation | 0.42 correlation Held-out test-set Spearman from the creator paper. | 12581 |
| ESM-untrained (mean) | Spearman correlation | 0.22 correlation Held-out test-set Spearman from the creator paper. | 12581 |
| ESM-untrained (mut mean) | Spearman correlation | 0.56 correlation Held-out test-set Spearman from the creator paper. | 12581 |
| Ridge regression (FLIP) | Spearman correlation | 0.65 correlation Held-out test-set Spearman from the creator paper. | 12581 |
| Convolutional network (FLIP) | Spearman correlation | 0.74 correlation Held-out test-set Spearman from the creator paper. | 12581 |
| Levenshtein distance baseline (FLIP) | Spearman correlation | 0.53 correlation Held-out test-set Spearman from the creator paper. | 12581 |
Evaluation run
From FLIP: Benchmark tasks in fitness landscape inference for proteins
Evaluated models / systems: Convolutional network (FLIP), ESM-untrained (mean), ESM-untrained (mut mean), ESM-untrained (per AA), ESM-1b (mean; FLIP head), ESM-1b (mut mean; FLIP head), ESM-1b (per AA; FLIP head), ESM-1v (mean; FLIP head), ESM-1v (mut mean; FLIP head), ESM-1v (per AA; FLIP head), Levenshtein distance baseline (FLIP), Ridge regression (FLIP)
| Metric | Kind / baseline | Unit | Aggregation | Threshold / tolerance |
|---|---|---|---|---|
| Spearman correlation | absolute | correlation | computed across all examples in this held-out test split | Not reported |
| Mean squared error | absolute | squared fitness units | mean across all examples in this held-out test split | Not reported |
| Model | Metric | Value | n |
|---|---|---|---|
| ESM-1b (per AA; FLIP head) | Spearman correlation | 0.65 correlation Held-out test-set Spearman from the creator paper. | 50776 |
| ESM-1b (mean; FLIP head) | Spearman correlation | 0.26 correlation Held-out test-set Spearman from the creator paper. | 50776 |
| ESM-1b (mut mean; FLIP head) | Spearman correlation | 0.65 correlation Held-out test-set Spearman from the creator paper. | 50776 |
| ESM-1v (per AA; FLIP head) | Spearman correlation | 0.7 correlation Held-out test-set Spearman from the creator paper. | 50776 |
| ESM-1v (mean; FLIP head) | Spearman correlation | 0.16 correlation Held-out test-set Spearman from the creator paper. | 50776 |
| ESM-1v (mut mean; FLIP head) | Spearman correlation | 0.64 correlation Held-out test-set Spearman from the creator paper. | 50776 |
| ESM-untrained (per AA) | Spearman correlation | 0.22 correlation Held-out test-set Spearman from the creator paper. | 50776 |
| ESM-untrained (mean) | Spearman correlation | 0.14 correlation Held-out test-set Spearman from the creator paper. | 50776 |
| ESM-untrained (mut mean) | Spearman correlation | 0.16 correlation Held-out test-set Spearman from the creator paper. | 50776 |
| Ridge regression (FLIP) | Spearman correlation | 0.03 correlation Held-out test-set Spearman from the creator paper. | 50776 |
| Convolutional network (FLIP) | Spearman correlation | 0.74 correlation Held-out test-set Spearman from the creator paper. | 50776 |
| Levenshtein distance baseline (FLIP) | Spearman correlation | 0.57 correlation Held-out test-set Spearman from the creator paper. | 50776 |
Evaluation run
From FLIP: Benchmark tasks in fitness landscape inference for proteins
Evaluated models / systems: BLOSUM62 baseline (FLIP), Convolutional network (FLIP), ESM-untrained (mean), ESM-untrained (mut mean), ESM-untrained (per AA), ESM-1b (mean; FLIP head), ESM-1b (mut mean; FLIP head), ESM-1b (per AA; FLIP head), ESM-1v (mean; FLIP head), ESM-1v (mut mean; FLIP head), ESM-1v (per AA; FLIP head), Levenshtein distance baseline (FLIP), Ridge regression (FLIP)
| Metric | Kind / baseline | Unit | Aggregation | Threshold / tolerance |
|---|---|---|---|---|
| Spearman correlation | absolute | correlation | computed across all examples in this held-out test split | Not reported |
| Mean squared error | absolute | squared fitness units | mean across all examples in this held-out test split | Not reported |
| Model | Metric | Value | n |
|---|---|---|---|
| ESM-1b (per AA; FLIP head) | Spearman correlation | 0.59 correlation Held-out test-set Spearman from the creator paper. | 3644 |
| ESM-1b (mean; FLIP head) | Spearman correlation | 0.13 correlation Held-out test-set Spearman from the creator paper. | 3644 |
| ESM-1b (mut mean; FLIP head) | Spearman correlation | 0.45 correlation Held-out test-set Spearman from the creator paper. | 3644 |
| ESM-1v (per AA; FLIP head) | Spearman correlation | 0.51 correlation Held-out test-set Spearman from the creator paper. | 3644 |
| ESM-1v (mean; FLIP head) | Spearman correlation | 0.1 correlation Held-out test-set Spearman from the creator paper. | 3644 |
| ESM-1v (mut mean; FLIP head) | Spearman correlation | 0.49 correlation Held-out test-set Spearman from the creator paper. | 3644 |
| ESM-untrained (per AA) | Spearman correlation | 0.23 correlation Held-out test-set Spearman from the creator paper. | 3644 |
| ESM-untrained (mean) | Spearman correlation | 0.1 correlation Held-out test-set Spearman from the creator paper. | 3644 |
| ESM-untrained (mut mean) | Spearman correlation | 0.13 correlation Held-out test-set Spearman from the creator paper. | 3644 |
| Ridge regression (FLIP) | Spearman correlation | 0.34 correlation Held-out test-set Spearman from the creator paper. | 3644 |
| Convolutional network (FLIP) | Spearman correlation | 0.51 correlation Held-out test-set Spearman from the creator paper. | 3644 |
| Levenshtein distance baseline (FLIP) | Spearman correlation | -0.1 correlation Held-out test-set Spearman from the creator paper. | 3644 |
| BLOSUM62 baseline (FLIP) | Spearman correlation | -0.13 correlation Held-out test-set Spearman from the creator paper. | 3644 |
Evaluation run
From FLIP: Benchmark tasks in fitness landscape inference for proteins
Evaluated models / systems: BLOSUM62 baseline (FLIP), Convolutional network (FLIP), ESM-untrained (mean), ESM-untrained (mut mean), ESM-untrained (per AA), ESM-1b (mean; FLIP head), ESM-1b (mut mean; FLIP head), ESM-1b (per AA; FLIP head), ESM-1v (mean; FLIP head), ESM-1v (mut mean; FLIP head), ESM-1v (per AA; FLIP head), Levenshtein distance baseline (FLIP), Ridge regression (FLIP)
| Metric | Kind / baseline | Unit | Aggregation | Threshold / tolerance |
|---|---|---|---|---|
| Spearman correlation | absolute | correlation | computed across all examples in this held-out test split | Not reported |
| Mean squared error | absolute | squared fitness units | mean across all examples in this held-out test split | Not reported |
| Model | Metric | Value | n |
|---|---|---|---|
| ESM-1b (per AA; FLIP head) | Spearman correlation | 0.28 correlation Held-out test-set Spearman from the creator paper. | 8704 |
| ESM-1b (mean; FLIP head) | Spearman correlation | 0.32 correlation Held-out test-set Spearman from the creator paper. | 8704 |
| ESM-1b (mut mean; FLIP head) | Spearman correlation | -0.08 correlation Held-out test-set Spearman from the creator paper. | 8704 |
| ESM-1v (per AA; FLIP head) | Spearman correlation | 0.28 correlation Held-out test-set Spearman from the creator paper. | 8704 |
| ESM-1v (mean; FLIP head) | Spearman correlation | 0.32 correlation Held-out test-set Spearman from the creator paper. | 8704 |
| ESM-1v (mut mean; FLIP head) | Spearman correlation | 0.19 correlation Held-out test-set Spearman from the creator paper. | 8704 |
| ESM-untrained (per AA) | Spearman correlation | 0.06 correlation Held-out test-set Spearman from the creator paper. | 8704 |
| ESM-untrained (mean) | Spearman correlation | 0.05 correlation Held-out test-set Spearman from the creator paper. | 8704 |
| ESM-untrained (mut mean) | Spearman correlation | 0.21 correlation Held-out test-set Spearman from the creator paper. | 8704 |
| Ridge regression (FLIP) | Spearman correlation | 0.28 correlation Held-out test-set Spearman from the creator paper. | 8704 |
| Convolutional network (FLIP) | Spearman correlation | 0.17 correlation Held-out test-set Spearman from the creator paper. | 8704 |
| Levenshtein distance baseline (FLIP) | Spearman correlation | 0.17 correlation Held-out test-set Spearman from the creator paper. | 8704 |
| BLOSUM62 baseline (FLIP) | Spearman correlation | 0.15 correlation Held-out test-set Spearman from the creator paper. | 8704 |
Evaluation run
From FLIP: Benchmark tasks in fitness landscape inference for proteins
Evaluated models / systems: Convolutional network (FLIP), ESM-untrained (per AA), ESM-1b (per AA; FLIP head), ESM-1v (per AA; FLIP head), Ridge regression (FLIP)
| Metric | Kind / baseline | Unit | Aggregation | Threshold / tolerance |
|---|---|---|---|---|
| Spearman correlation | absolute | correlation | computed across all examples in this held-out test split | Not reported |
| Mean squared error | absolute | squared fitness units | mean across all examples in this held-out test split | Not reported |
| Model | Metric | Value | n |
|---|---|---|---|
| ESM-1b (per AA; FLIP head) | Spearman correlation | 0.92 correlation Held-out test-set Spearman from the creator paper. | 1772 |
| ESM-1v (per AA; FLIP head) | Spearman correlation | 0.92 correlation Held-out test-set Spearman from the creator paper. | 1772 |
| ESM-untrained (per AA) | Spearman correlation | 0.79 correlation Held-out test-set Spearman from the creator paper. | 1772 |
| Ridge regression (FLIP) | Spearman correlation | 0.82 correlation Held-out test-set Spearman from the creator paper. | 1772 |
| Convolutional network (FLIP) | Spearman correlation | 0.91 correlation Held-out test-set Spearman from the creator paper. | 1772 |
Evaluation run
From FLIP: Benchmark tasks in fitness landscape inference for proteins
Evaluated models / systems: BLOSUM62 baseline (FLIP), Convolutional network (FLIP), ESM-untrained (mean), ESM-untrained (mut mean), ESM-untrained (per AA), ESM-1b (mean; FLIP head), ESM-1b (mut mean; FLIP head), ESM-1b (per AA; FLIP head), ESM-1v (mean; FLIP head), ESM-1v (mut mean; FLIP head), ESM-1v (per AA; FLIP head), Levenshtein distance baseline (FLIP), Ridge regression (FLIP)
| Metric | Kind / baseline | Unit | Aggregation | Threshold / tolerance |
|---|---|---|---|---|
| Spearman correlation | absolute | correlation | computed across all examples in this held-out test split | Not reported |
| Mean squared error | absolute | squared fitness units | mean across all examples in this held-out test split | Not reported |
| Model | Metric | Value | n |
|---|---|---|---|
| ESM-1b (per AA; FLIP head) | Spearman correlation | 0.79 correlation Held-out test-set Spearman from the creator paper. | 5765 |
| ESM-1b (mean; FLIP head) | Spearman correlation | 0.54 correlation Held-out test-set Spearman from the creator paper. | 5765 |
| ESM-1b (mut mean; FLIP head) | Spearman correlation | 0.49 correlation Held-out test-set Spearman from the creator paper. | 5765 |
| ESM-1v (per AA; FLIP head) | Spearman correlation | 0.82 correlation Held-out test-set Spearman from the creator paper. | 5765 |
| ESM-1v (mean; FLIP head) | Spearman correlation | 0.77 correlation Held-out test-set Spearman from the creator paper. | 5765 |
| ESM-1v (mut mean; FLIP head) | Spearman correlation | 0.8 correlation Held-out test-set Spearman from the creator paper. | 5765 |
| ESM-untrained (per AA) | Spearman correlation | 0.48 correlation Held-out test-set Spearman from the creator paper. | 5765 |
| ESM-untrained (mean) | Spearman correlation | 0.46 correlation Held-out test-set Spearman from the creator paper. | 5765 |
| ESM-untrained (mut mean) | Spearman correlation | 0.57 correlation Held-out test-set Spearman from the creator paper. | 5765 |
| Ridge regression (FLIP) | Spearman correlation | 0.76 correlation Held-out test-set Spearman from the creator paper. | 5765 |
| Convolutional network (FLIP) | Spearman correlation | 0.83 correlation Held-out test-set Spearman from the creator paper. | 5765 |
| Levenshtein distance baseline (FLIP) | Spearman correlation | -0.04 correlation Held-out test-set Spearman from the creator paper. | 5765 |
| BLOSUM62 baseline (FLIP) | Spearman correlation | 0.01 correlation Held-out test-set Spearman from the creator paper. | 5765 |
Evaluation run
From FLIP: Benchmark tasks in fitness landscape inference for proteins
Evaluated models / systems: BLOSUM62 baseline (FLIP), Convolutional network (FLIP), ESM-untrained (mean), ESM-untrained (mut mean), ESM-untrained (per AA), ESM-1b (mean; FLIP head), ESM-1b (mut mean; FLIP head), ESM-1b (per AA; FLIP head), ESM-1v (mean; FLIP head), ESM-1v (mut mean; FLIP head), ESM-1v (per AA; FLIP head), Levenshtein distance baseline (FLIP), Ridge regression (FLIP)
| Metric | Kind / baseline | Unit | Aggregation | Threshold / tolerance |
|---|---|---|---|---|
| Spearman correlation | absolute | correlation | computed across all examples in this held-out test split | Not reported |
| Mean squared error | absolute | squared fitness units | mean across all examples in this held-out test split | Not reported |
| Model | Metric | Value | n |
|---|---|---|---|
| ESM-1b (per AA; FLIP head) | Spearman correlation | 0.55 correlation Held-out test-set Spearman from the creator paper. | 8306 |
| ESM-1b (mean; FLIP head) | Spearman correlation | 0.36 correlation Held-out test-set Spearman from the creator paper. | 8306 |
| ESM-1b (mut mean; FLIP head) | Spearman correlation | 0.19 correlation Held-out test-set Spearman from the creator paper. | 8306 |
| ESM-1v (per AA; FLIP head) | Spearman correlation | 0.28 correlation Held-out test-set Spearman from the creator paper. | 8306 |
| ESM-1v (mean; FLIP head) | Spearman correlation | 0.32 correlation Held-out test-set Spearman from the creator paper. | 8306 |
| ESM-1v (mut mean; FLIP head) | Spearman correlation | 0.19 correlation Held-out test-set Spearman from the creator paper. | 8306 |
| ESM-untrained (per AA) | Spearman correlation | 0.06 correlation Held-out test-set Spearman from the creator paper. | 8306 |
| ESM-untrained (mean) | Spearman correlation | 0.05 correlation Held-out test-set Spearman from the creator paper. | 8306 |
| ESM-untrained (mut mean) | Spearman correlation | 0.21 correlation Held-out test-set Spearman from the creator paper. | 8306 |
| Ridge regression (FLIP) | Spearman correlation | 0.59 correlation Held-out test-set Spearman from the creator paper. | 8306 |
| Convolutional network (FLIP) | Spearman correlation | 0.32 correlation Held-out test-set Spearman from the creator paper. | 8306 |
| Levenshtein distance baseline (FLIP) | Spearman correlation | 0.16 correlation Held-out test-set Spearman from the creator paper. | 8306 |
| BLOSUM62 baseline (FLIP) | Spearman correlation | 0.14 correlation Held-out test-set Spearman from the creator paper. | 8306 |
Evaluation run
From FLIP: Benchmark tasks in fitness landscape inference for proteins
Evaluated models / systems: Convolutional network (FLIP), ESM-untrained (mean), ESM-untrained (per AA), ESM-1b (mean; FLIP head), ESM-1b (per AA; FLIP head), ESM-1v (mean; FLIP head), ESM-1v (per AA; FLIP head), Ridge regression (FLIP)
| Metric | Kind / baseline | Unit | Aggregation | Threshold / tolerance |
|---|---|---|---|---|
| Spearman correlation | absolute | correlation | computed across all examples in this held-out test split | Not reported |
| Mean squared error | absolute | squared fitness units | mean across all examples in this held-out test split | Not reported |
| Model | Metric | Value | n |
|---|---|---|---|
| ESM-1b (per AA; FLIP head) | Spearman correlation | 0.71 correlation Held-out test-set Spearman from the creator paper. | 1945 |
| ESM-1b (mean; FLIP head) | Spearman correlation | 0.7 correlation Held-out test-set Spearman from the creator paper. | 1945 |
| ESM-1v (per AA; FLIP head) | Spearman correlation | 0.77 correlation Held-out test-set Spearman from the creator paper. | 1945 |
| ESM-1v (mean; FLIP head) | Spearman correlation | 0.75 correlation Held-out test-set Spearman from the creator paper. | 1945 |
| ESM-untrained (per AA) | Spearman correlation | 0.44 correlation Held-out test-set Spearman from the creator paper. | 1945 |
| ESM-untrained (mean) | Spearman correlation | 0.48 correlation Held-out test-set Spearman from the creator paper. | 1945 |
| Ridge regression (FLIP) | Spearman correlation | 0.15 correlation Held-out test-set Spearman from the creator paper. | 1945 |
| Convolutional network (FLIP) | Spearman correlation | 0.5 correlation Held-out test-set Spearman from the creator paper. | 1945 |
Evaluation run
From FLIP: Benchmark tasks in fitness landscape inference for proteins
Evaluated models / systems: Convolutional network (FLIP), ESM-untrained (mean), ESM-untrained (per AA), ESM-1b (mean; FLIP head), ESM-1b (per AA; FLIP head), ESM-1v (mean; FLIP head), ESM-1v (per AA; FLIP head), Ridge regression (FLIP)
| Metric | Kind / baseline | Unit | Aggregation | Threshold / tolerance |
|---|---|---|---|---|
| Spearman correlation | absolute | correlation | computed across all examples in this held-out test split | Not reported |
| Mean squared error | absolute | squared fitness units | mean across all examples in this held-out test split | Not reported |
| Model | Metric | Value | n |
|---|---|---|---|
| ESM-1b (per AA; FLIP head) | Spearman correlation | 0.76 correlation Held-out test-set Spearman from the creator paper. | 1366 |
| ESM-1b (mean; FLIP head) | Spearman correlation | 0.75 correlation Held-out test-set Spearman from the creator paper. | 1366 |
| ESM-1v (per AA; FLIP head) | Spearman correlation | 0.78 correlation Held-out test-set Spearman from the creator paper. | 1366 |
| ESM-1v (mean; FLIP head) | Spearman correlation | 0.74 correlation Held-out test-set Spearman from the creator paper. | 1366 |
| ESM-untrained (per AA) | Spearman correlation | 0.46 correlation Held-out test-set Spearman from the creator paper. | 1366 |
| ESM-untrained (mean) | Spearman correlation | 0.49 correlation Held-out test-set Spearman from the creator paper. | 1366 |
| Ridge regression (FLIP) | Spearman correlation | 0.24 correlation Held-out test-set Spearman from the creator paper. | 1366 |
| Convolutional network (FLIP) | Spearman correlation | 0.49 correlation Held-out test-set Spearman from the creator paper. | 1366 |
Evaluation run
From FLIP: Benchmark tasks in fitness landscape inference for proteins
Evaluated models / systems: Convolutional network (FLIP), ESM-untrained (mean), ESM-untrained (per AA), ESM-1b (mean; FLIP head), ESM-1b (per AA; FLIP head), ESM-1v (mean; FLIP head), ESM-1v (per AA; FLIP head), Ridge regression (FLIP)
| Metric | Kind / baseline | Unit | Aggregation | Threshold / tolerance |
|---|---|---|---|---|
| Spearman correlation | absolute | correlation | computed across all examples in this held-out test split | Not reported |
| Mean squared error | absolute | squared fitness units | mean across all examples in this held-out test split | Not reported |
| Model | Metric | Value | n |
|---|---|---|---|
| ESM-1b (per AA; FLIP head) | Spearman correlation | 0.68 correlation Held-out test-set Spearman from the creator paper. | 3134 |
| ESM-1b (mean; FLIP head) | Spearman correlation | 0.68 correlation Held-out test-set Spearman from the creator paper. | 3134 |
| ESM-1v (per AA; FLIP head) | Spearman correlation | 0.65 correlation Held-out test-set Spearman from the creator paper. | 3134 |
| ESM-1v (mean; FLIP head) | Spearman correlation | 0.67 correlation Held-out test-set Spearman from the creator paper. | 3134 |
| ESM-untrained (per AA) | Spearman correlation | 0.44 correlation Held-out test-set Spearman from the creator paper. | 3134 |
| ESM-untrained (mean) | Spearman correlation | 0.36 correlation Held-out test-set Spearman from the creator paper. | 3134 |
| Ridge regression (FLIP) | Spearman correlation | 0.17 correlation Held-out test-set Spearman from the creator paper. | 3134 |
| Convolutional network (FLIP) | Spearman correlation | 0.34 correlation Held-out test-set Spearman from the creator paper. | 3134 |
flip-aav-des-mut · flip-aav-des-mut-spearman
| Model | Value | Comparability group |
|---|---|---|
| ESM-1b (mean; FLIP head) | 0.59 | flip-aav-des-mut-spearman |
| ESM-1b (mut mean; FLIP head) | 0.7 | flip-aav-des-mut-spearman |
| ESM-1v (mean; FLIP head) | 0.44 | flip-aav-des-mut-spearman |
| ESM-1v (mut mean; FLIP head) | 0.71 | flip-aav-des-mut-spearman |
| ESM-untrained (mean) | 0.34 | flip-aav-des-mut-spearman |
| ESM-untrained (mut mean) | 0.64 | flip-aav-des-mut-spearman |
| Ridge regression (FLIP) | 0.53 | flip-aav-des-mut-spearman |
| Convolutional network (FLIP) | 0.75 | flip-aav-des-mut-spearman |
| Levenshtein distance baseline (FLIP) | -0.07 | flip-aav-des-mut-spearman |
flip-aav-low-vs-high · flip-aav-low-vs-high-spearman
| Model | Value | Comparability group |
|---|---|---|
| ESM-1b (per AA; FLIP head) | 0.39 | flip-aav-low-vs-high-spearman |
| ESM-1b (mean; FLIP head) | 0.18 | flip-aav-low-vs-high-spearman |
| ESM-1b (mut mean; FLIP head) | 0.33 | flip-aav-low-vs-high-spearman |
| ESM-1v (per AA; FLIP head) | 0.34 | flip-aav-low-vs-high-spearman |
| ESM-1v (mean; FLIP head) | 0.2 | flip-aav-low-vs-high-spearman |
| ESM-1v (mut mean; FLIP head) | 0.31 | flip-aav-low-vs-high-spearman |
| ESM-untrained (per AA) | 0.08 | flip-aav-low-vs-high-spearman |
| ESM-untrained (mean) | 0.22 | flip-aav-low-vs-high-spearman |
| ESM-untrained (mut mean) | 0.24 | flip-aav-low-vs-high-spearman |
| Ridge regression (FLIP) | 0.12 | flip-aav-low-vs-high-spearman |
| Convolutional network (FLIP) | 0.34 | flip-aav-low-vs-high-spearman |
| Levenshtein distance baseline (FLIP) | 0.25 | flip-aav-low-vs-high-spearman |
flip-aav-mut-des · flip-aav-mut-des-spearman
| Model | Value | Comparability group |
|---|---|---|
| ESM-1b (per AA; FLIP head) | 0.76 | flip-aav-mut-des-spearman |
| ESM-1b (mean; FLIP head) | 0.63 | flip-aav-mut-des-spearman |
| ESM-1b (mut mean; FLIP head) | 0.7 | flip-aav-mut-des-spearman |
| ESM-1v (per AA; FLIP head) | 0.79 | flip-aav-mut-des-spearman |
| ESM-1v (mean; FLIP head) | 0.55 | flip-aav-mut-des-spearman |
| ESM-1v (mut mean; FLIP head) | 0.7 | flip-aav-mut-des-spearman |
| ESM-untrained (per AA) | 0.56 | flip-aav-mut-des-spearman |
| ESM-untrained (mean) | 0.27 | flip-aav-mut-des-spearman |
| ESM-untrained (mut mean) | 0.62 | flip-aav-mut-des-spearman |
| Ridge regression (FLIP) | 0.64 | flip-aav-mut-des-spearman |
| Convolutional network (FLIP) | 0.71 | flip-aav-mut-des-spearman |
| Levenshtein distance baseline (FLIP) | 0.6 | flip-aav-mut-des-spearman |
flip-aav-one-vs-rest · flip-aav-one-vs-rest-spearman
| Model | Value | Comparability group |
|---|---|---|
| ESM-1b (per AA; FLIP head) | 0.03 | flip-aav-one-vs-rest-spearman |
| ESM-1b (mean; FLIP head) | 0.04 | flip-aav-one-vs-rest-spearman |
| ESM-1b (mut mean; FLIP head) | 0.31 | flip-aav-one-vs-rest-spearman |
| ESM-1v (per AA; FLIP head) | 0.1 | flip-aav-one-vs-rest-spearman |
| ESM-1v (mean; FLIP head) | 0.18 | flip-aav-one-vs-rest-spearman |
| ESM-1v (mut mean; FLIP head) | 0.44 | flip-aav-one-vs-rest-spearman |
| ESM-untrained (per AA) | 0.18 | flip-aav-one-vs-rest-spearman |
| ESM-untrained (mean) | 0.01 | flip-aav-one-vs-rest-spearman |
| ESM-untrained (mut mean) | 0.26 | flip-aav-one-vs-rest-spearman |
| Ridge regression (FLIP) | 0.22 | flip-aav-one-vs-rest-spearman |
| Convolutional network (FLIP) | 0.48 | flip-aav-one-vs-rest-spearman |
| Levenshtein distance baseline (FLIP) | -0.11 | flip-aav-one-vs-rest-spearman |
flip-aav-sampled · flip-aav-sampled-spearman
| Model | Value | Comparability group |
|---|---|---|
| ESM-1b (per AA; FLIP head) | 0.9 | flip-aav-sampled-spearman |
| ESM-1v (per AA; FLIP head) | 0.92 | flip-aav-sampled-spearman |
| ESM-untrained (per AA) | 0.78 | flip-aav-sampled-spearman |
| Ridge regression (FLIP) | 0.83 | flip-aav-sampled-spearman |
| Convolutional network (FLIP) | 0.92 | flip-aav-sampled-spearman |
flip-aav-seven-vs-rest · flip-aav-seven-vs-rest-spearman
| Model | Value | Comparability group |
|---|---|---|
| ESM-1b (per AA; FLIP head) | 0.65 | flip-aav-seven-vs-rest-spearman |
| ESM-1b (mean; FLIP head) | 0.46 | flip-aav-seven-vs-rest-spearman |
| ESM-1b (mut mean; FLIP head) | 0.61 | flip-aav-seven-vs-rest-spearman |
| ESM-1v (per AA; FLIP head) | 0.7 | flip-aav-seven-vs-rest-spearman |
| ESM-1v (mean; FLIP head) | 0.45 | flip-aav-seven-vs-rest-spearman |
| ESM-1v (mut mean; FLIP head) | 0.64 | flip-aav-seven-vs-rest-spearman |
| ESM-untrained (per AA) | 0.42 | flip-aav-seven-vs-rest-spearman |
| ESM-untrained (mean) | 0.22 | flip-aav-seven-vs-rest-spearman |
| ESM-untrained (mut mean) | 0.56 | flip-aav-seven-vs-rest-spearman |
| Ridge regression (FLIP) | 0.65 | flip-aav-seven-vs-rest-spearman |
| Convolutional network (FLIP) | 0.74 | flip-aav-seven-vs-rest-spearman |
| Levenshtein distance baseline (FLIP) | 0.53 | flip-aav-seven-vs-rest-spearman |
flip-aav-two-vs-rest · flip-aav-two-vs-rest-spearman
| Model | Value | Comparability group |
|---|---|---|
| ESM-1b (per AA; FLIP head) | 0.65 | flip-aav-two-vs-rest-spearman |
| ESM-1b (mean; FLIP head) | 0.26 | flip-aav-two-vs-rest-spearman |
| ESM-1b (mut mean; FLIP head) | 0.65 | flip-aav-two-vs-rest-spearman |
| ESM-1v (per AA; FLIP head) | 0.7 | flip-aav-two-vs-rest-spearman |
| ESM-1v (mean; FLIP head) | 0.16 | flip-aav-two-vs-rest-spearman |
| ESM-1v (mut mean; FLIP head) | 0.64 | flip-aav-two-vs-rest-spearman |
| ESM-untrained (per AA) | 0.22 | flip-aav-two-vs-rest-spearman |
| ESM-untrained (mean) | 0.14 | flip-aav-two-vs-rest-spearman |
| ESM-untrained (mut mean) | 0.16 | flip-aav-two-vs-rest-spearman |
| Ridge regression (FLIP) | 0.03 | flip-aav-two-vs-rest-spearman |
| Convolutional network (FLIP) | 0.74 | flip-aav-two-vs-rest-spearman |
| Levenshtein distance baseline (FLIP) | 0.57 | flip-aav-two-vs-rest-spearman |
flip-gb1-low-vs-high · flip-gb1-low-vs-high-spearman
| Model | Value | Comparability group |
|---|---|---|
| ESM-1b (per AA; FLIP head) | 0.59 | flip-gb1-low-vs-high-spearman |
| ESM-1b (mean; FLIP head) | 0.13 | flip-gb1-low-vs-high-spearman |
| ESM-1b (mut mean; FLIP head) | 0.45 | flip-gb1-low-vs-high-spearman |
| ESM-1v (per AA; FLIP head) | 0.51 | flip-gb1-low-vs-high-spearman |
| ESM-1v (mean; FLIP head) | 0.1 | flip-gb1-low-vs-high-spearman |
| ESM-1v (mut mean; FLIP head) | 0.49 | flip-gb1-low-vs-high-spearman |
| ESM-untrained (per AA) | 0.23 | flip-gb1-low-vs-high-spearman |
| ESM-untrained (mean) | 0.1 | flip-gb1-low-vs-high-spearman |
| ESM-untrained (mut mean) | 0.13 | flip-gb1-low-vs-high-spearman |
| Ridge regression (FLIP) | 0.34 | flip-gb1-low-vs-high-spearman |
| Convolutional network (FLIP) | 0.51 | flip-gb1-low-vs-high-spearman |
| Levenshtein distance baseline (FLIP) | -0.1 | flip-gb1-low-vs-high-spearman |
| BLOSUM62 baseline (FLIP) | -0.13 | flip-gb1-low-vs-high-spearman |
flip-gb1-one-vs-rest · flip-gb1-one-vs-rest-spearman
| Model | Value | Comparability group |
|---|---|---|
| ESM-1b (per AA; FLIP head) | 0.28 | flip-gb1-one-vs-rest-spearman |
| ESM-1b (mean; FLIP head) | 0.32 | flip-gb1-one-vs-rest-spearman |
| ESM-1b (mut mean; FLIP head) | -0.08 | flip-gb1-one-vs-rest-spearman |
| ESM-1v (per AA; FLIP head) | 0.28 | flip-gb1-one-vs-rest-spearman |
| ESM-1v (mean; FLIP head) | 0.32 | flip-gb1-one-vs-rest-spearman |
| ESM-1v (mut mean; FLIP head) | 0.19 | flip-gb1-one-vs-rest-spearman |
| ESM-untrained (per AA) | 0.06 | flip-gb1-one-vs-rest-spearman |
| ESM-untrained (mean) | 0.05 | flip-gb1-one-vs-rest-spearman |
| ESM-untrained (mut mean) | 0.21 | flip-gb1-one-vs-rest-spearman |
| Ridge regression (FLIP) | 0.28 | flip-gb1-one-vs-rest-spearman |
| Convolutional network (FLIP) | 0.17 | flip-gb1-one-vs-rest-spearman |
| Levenshtein distance baseline (FLIP) | 0.17 | flip-gb1-one-vs-rest-spearman |
| BLOSUM62 baseline (FLIP) | 0.15 | flip-gb1-one-vs-rest-spearman |
flip-gb1-sampled · flip-gb1-sampled-spearman
| Model | Value | Comparability group |
|---|---|---|
| ESM-1b (per AA; FLIP head) | 0.92 | flip-gb1-sampled-spearman |
| ESM-1v (per AA; FLIP head) | 0.92 | flip-gb1-sampled-spearman |
| ESM-untrained (per AA) | 0.79 | flip-gb1-sampled-spearman |
| Ridge regression (FLIP) | 0.82 | flip-gb1-sampled-spearman |
| Convolutional network (FLIP) | 0.91 | flip-gb1-sampled-spearman |
flip-gb1-three-vs-rest · flip-gb1-three-vs-rest-spearman
| Model | Value | Comparability group |
|---|---|---|
| ESM-1b (per AA; FLIP head) | 0.79 | flip-gb1-three-vs-rest-spearman |
| ESM-1b (mean; FLIP head) | 0.54 | flip-gb1-three-vs-rest-spearman |
| ESM-1b (mut mean; FLIP head) | 0.49 | flip-gb1-three-vs-rest-spearman |
| ESM-1v (per AA; FLIP head) | 0.82 | flip-gb1-three-vs-rest-spearman |
| ESM-1v (mean; FLIP head) | 0.77 | flip-gb1-three-vs-rest-spearman |
| ESM-1v (mut mean; FLIP head) | 0.8 | flip-gb1-three-vs-rest-spearman |
| ESM-untrained (per AA) | 0.48 | flip-gb1-three-vs-rest-spearman |
| ESM-untrained (mean) | 0.46 | flip-gb1-three-vs-rest-spearman |
| ESM-untrained (mut mean) | 0.57 | flip-gb1-three-vs-rest-spearman |
| Ridge regression (FLIP) | 0.76 | flip-gb1-three-vs-rest-spearman |
| Convolutional network (FLIP) | 0.83 | flip-gb1-three-vs-rest-spearman |
| Levenshtein distance baseline (FLIP) | -0.04 | flip-gb1-three-vs-rest-spearman |
| BLOSUM62 baseline (FLIP) | 0.01 | flip-gb1-three-vs-rest-spearman |
flip-gb1-two-vs-rest · flip-gb1-two-vs-rest-spearman
| Model | Value | Comparability group |
|---|---|---|
| ESM-1b (per AA; FLIP head) | 0.55 | flip-gb1-two-vs-rest-spearman |
| ESM-1b (mean; FLIP head) | 0.36 | flip-gb1-two-vs-rest-spearman |
| ESM-1b (mut mean; FLIP head) | 0.19 | flip-gb1-two-vs-rest-spearman |
| ESM-1v (per AA; FLIP head) | 0.28 | flip-gb1-two-vs-rest-spearman |
| ESM-1v (mean; FLIP head) | 0.32 | flip-gb1-two-vs-rest-spearman |
| ESM-1v (mut mean; FLIP head) | 0.19 | flip-gb1-two-vs-rest-spearman |
| ESM-untrained (per AA) | 0.06 | flip-gb1-two-vs-rest-spearman |
| ESM-untrained (mean) | 0.05 | flip-gb1-two-vs-rest-spearman |
| ESM-untrained (mut mean) | 0.21 | flip-gb1-two-vs-rest-spearman |
| Ridge regression (FLIP) | 0.59 | flip-gb1-two-vs-rest-spearman |
| Convolutional network (FLIP) | 0.32 | flip-gb1-two-vs-rest-spearman |
| Levenshtein distance baseline (FLIP) | 0.16 | flip-gb1-two-vs-rest-spearman |
| BLOSUM62 baseline (FLIP) | 0.14 | flip-gb1-two-vs-rest-spearman |
flip-meltome-human · flip-meltome-human-spearman
| Model | Value | Comparability group |
|---|---|---|
| ESM-1b (per AA; FLIP head) | 0.71 | flip-meltome-human-spearman |
| ESM-1b (mean; FLIP head) | 0.7 | flip-meltome-human-spearman |
| ESM-1v (per AA; FLIP head) | 0.77 | flip-meltome-human-spearman |
| ESM-1v (mean; FLIP head) | 0.75 | flip-meltome-human-spearman |
| ESM-untrained (per AA) | 0.44 | flip-meltome-human-spearman |
| ESM-untrained (mean) | 0.48 | flip-meltome-human-spearman |
| Ridge regression (FLIP) | 0.15 | flip-meltome-human-spearman |
| Convolutional network (FLIP) | 0.5 | flip-meltome-human-spearman |
flip-meltome-human-cell · flip-meltome-human-cell-spearman
| Model | Value | Comparability group |
|---|---|---|
| ESM-1b (per AA; FLIP head) | 0.76 | flip-meltome-human-cell-spearman |
| ESM-1b (mean; FLIP head) | 0.75 | flip-meltome-human-cell-spearman |
| ESM-1v (per AA; FLIP head) | 0.78 | flip-meltome-human-cell-spearman |
| ESM-1v (mean; FLIP head) | 0.74 | flip-meltome-human-cell-spearman |
| ESM-untrained (per AA) | 0.46 | flip-meltome-human-cell-spearman |
| ESM-untrained (mean) | 0.49 | flip-meltome-human-cell-spearman |
| Ridge regression (FLIP) | 0.24 | flip-meltome-human-cell-spearman |
| Convolutional network (FLIP) | 0.49 | flip-meltome-human-cell-spearman |
flip-meltome-mixed · flip-meltome-mixed-spearman
| Model | Value | Comparability group |
|---|---|---|
| ESM-1b (per AA; FLIP head) | 0.68 | flip-meltome-mixed-spearman |
| ESM-1b (mean; FLIP head) | 0.68 | flip-meltome-mixed-spearman |
| ESM-1v (per AA; FLIP head) | 0.65 | flip-meltome-mixed-spearman |
| ESM-1v (mean; FLIP head) | 0.67 | flip-meltome-mixed-spearman |
| ESM-untrained (per AA) | 0.44 | flip-meltome-mixed-spearman |
| ESM-untrained (mean) | 0.36 | flip-meltome-mixed-spearman |
| Ridge regression (FLIP) | 0.17 | flip-meltome-mixed-spearman |
| Convolutional network (FLIP) | 0.34 | flip-meltome-mixed-spearman |
Source locators remain visible; expand an item to inspect the exact Registry fields it supports.
/name/aliases/summary/kind/organizations/release_date/domains/capabilities/modalities/task_formats/task_counts/total/task_counts/basis/task_counts/subsets/coverage_notes/access/level/access/tasks/access/artifacts/access/grader/access/license/access/biosafety_notes/versions/0/release_date/versions/0/task_counts/total/versions/0/task_counts/basis/versions/0/task_counts/subsets/scientific_task_classification/entries/0/scientific_task_classification/entries/1/latest_version/task_counts/subsets/resources/implementations/versions/0/as_of/versions/0/task_counts/subsets/versions/0/formal_tracks