paper · benchmark creator

Evaluating Protein Transfer Learning with TAPE

University of California Berkeley · 2019-12-08

Relationship layer

Benchmark usage

This table records what the work did with each benchmark before attempting to normalize a run. Partial claims remain visible without being treated as comparable evaluations.

No BenchmarkUse relation is normalized for this legacy work yet. Existing EvaluationRuns remain available below.

Normalized evaluation runs

TAPE1 run

Open benchmark record →

Evaluation run

tape-creator-full

From Evaluating Protein Transfer Learning with TAPE

tape-creator-task-nativevoriginal-2019
Scopefull · n=5
ShotsNot applicable
TurnsNot applicable
System prompt publicNot applicable
Reasoning / effortNot applicable
BrowserNot applicable
InternetNot applicable
DatabasesNot applicable
Code executionNot applicable
ContainerNot reported
External toolstask-specific supervised heads over frozen or fine-tuned protein representations
Token budgetNot applicable
Time / cost budgetNot reported
TemperatureNot applicable
SeedNot reported
RepeatsNot reported
Graderdeterministic task-specific scorer · human review: no
StatisticsTask-level evaluation only; no cross-task normalized aggregate.
ContaminationSequence-identity filtering and biologically motivated held-out splits.
Metrics, results, and full protocol

Metrics

MetricKind / baselineUnitAggregationThreshold / tolerance
Per-amino-acid accuracyabsoluteproportionacross labeled residuesNot reported
L/5 medium and long-range precisionabsoluteproportionper protein then task summaryNot reported
Fold-level accuracyabsoluteproportionheld-out test examplesNot reported
Spearman's rhoabsolutecorrelationheld-out test examplesNot reported

No numeric result rows are published yet; the verified protocol remains useful.

Evidence

  • section: Sections 4.2-5 and Table 2; Appendix A (Defines all five task splits, architectures, training procedures, native metrics, and creator comparison table.) — supports /scope, /protocol, /metrics