suite · audited-with-caveats · verified 2026-08-20

AbBiBench

A framework using antibody–antigen complexes to evaluate affinity prediction and antibody redesign.

Audited with caveats: 2 field(s) are marked provisional or conflicted. Warnings are shown next to affected values and these claims are excluded from unqualified comparisons.

Benchmark definition

What is counted

Version
initial-release
Total
184500 (Standardized benchmark data compiling 184,500 mutated antibodies.)
Task formats
unclassified
Capabilities
PredictionDesignGenerationOptimization
Modalities
Protein sequence3D structureWet-lab output

Version history

VersionStatusRelease / as-ofTotalFormal tracks
initial-release
abbibench-initial-release-version
current2025-05-23184500 (Standardized benchmark data compiling 184,500 mutated antibodies.)None registered

Scientific Task Atlas

Scientific task classification

partial for initial-release. No Scientific Task claim passed independent high-confidence verification; task mapping remains pending a targeted official-source audit.

The source names only a broad direction; no more specific leaf task can be assigned without inference.

Relationship registry

How works use this benchmark

Partial claims, non-evaluation uses, and third-party summaries stay visible without entering model comparisons.

Partial evaluation claims

evaluation

abbibench-a-benchmark-for-antibody-binding-affinity-ma-abbibench-2-use

Partialunknown

Work: AbBiBench: A Benchmark for Antibody Binding Affinity Maturation and Design · source version abbibench-a-benchmark-for-antibody-binding-affinity-ma-arxiv-v2

Selection
not reported
Metrics
Not reported / not applicable
Linked runs
None

Not reported / unresolved: Exact model versions, providers, seeds, repeats, and confidence intervals are not reported.; The source conflicts on light-chain inclusion and on the 1mlc evaluation count.; Exact model and tool versions, compute time, and confidence intervals are not reported.; DiffAb seeds are described only as reaching up to 15; the exact seed list is absent.; The numeric ELISA detection threshold is not reported.; The final Pareto-candidate count conflicts between 18 and 21.; Exact model versions, seeds, repeats, and exact p-values are not reported.; benchmark version; realized n/scope; metric; numeric result; prompt and tools; grader and repeats

Owner-reviewed conservative publication: the creator evaluation is retained only as a partial relationship; conflicted settings and outcomes are omitted pending manual reconciliation.

Evidence
  • section: Section 3
    Supports: /relation_type
  • section: Section 3
    Supports: /benchmark_id
  • table: Table S2
    Supports: /model_ids
  • table: Table S2
    Supports: /model_ids
  • table: Table S2
    Supports: /model_ids
  • table: Table S2
    Supports: /model_ids
  • table: Table S2
    Supports: /model_ids
  • table: Table S2
    Supports: /model_ids
  • figure: Figure 3
    Supports: /model_ids
  • table: Table S2
    Supports: /model_ids
  • other: Table S2 and registry-context.json model proteinmpnn
    Supports: /model_ids
  • figure: Figure 3
    Supports: /model_ids
  • table: Table S2
    Supports: /model_ids
  • table: Table S2
    Supports: /model_ids
  • table: Table S2
    Supports: /model_ids
  • table: Table S2
    Supports: /model_ids
  • table: Table S2
    Supports: /model_ids
  • table: Table S2
    Supports: /model_ids
  • table: Table S2
    Supports: /model_ids

evaluation

explainable-protein-protein-binding-affinity-predictio-abbibench-9-use

Partialsubset · n=9

Work: Explainable protein-protein binding affinity prediction via fine-tuning protein language models · source version explainable-protein-protein-binding-affinity-predictio-biorxiv-version-posted-2

Selection
filtered · Nine selected DMS assays: 1n8z, 1mhp_LC, 3gbn_h1, 3gbn_h9, 4fqi_h1, aayl49_ml, aayl50_LC, aayl51, aayl52
Models
BALM-PPI
Metrics
RMSE
Linked runs
None

Not reported / unresolved: Benchmark version is not reported.; The source reports assay inventories, but not an exact realized evaluation n for every assay.; BALM-PPI provider and version are not explicitly reported.; benchmark version

AI-assisted double-pass extraction; n=9 counts selected assay tracks, not individual examples. Values are limited to independently supported claims.

Evidence
  • section: 2.4
    Supports: /relation_type
  • section: 2.4
    Supports: /benchmark_id
  • section: 2.4
    Supports: /scope
  • section: 2.4
    Supports: /scope
  • table: Table S5
    Supports: /scope
  • table: Table S5
    Supports: /scope
  • table: Table S6
    Supports: /model_ids
  • table: Table S6
    Supports: /metric_labels

Creation, training, validation, or model-selection uses

benchmark creation

abbibench-a-benchmark-for-antibody-binding-affinity-ma-abbibench-1-use

Non-evaluationunknown

Work: AbBiBench: A Benchmark for Antibody Binding Affinity Maturation and Design · source version abbibench-a-benchmark-for-antibody-binding-affinity-ma-arxiv-v2

Selection
not applicable
Models
Not reported / not applicable
Metrics
Not reported / not applicable
Linked runs
None

Not reported / unresolved: A benchmark version and artifact release date are not reported.; The source conflicts on whether AbBiBench contains 14 or 16 binding-affinity assays.

AI-assisted double-pass extraction; values are limited to independently supported claims.

Evidence
  • section: Introduction
    Supports: /relation_type
  • section: Introduction
    Supports: /benchmark_id

fine tuning

explainable-protein-protein-binding-affinity-predictio-abbibench-10-use

Non-evaluationunknown

Work: Explainable protein-protein binding affinity prediction via fine-tuning protein language models · source version explainable-protein-protein-binding-affinity-predictio-biorxiv-version-posted-2

Selection
not applicable · Few-Shot 10%
Models
BALM-PPI
Metrics
RMSE
Linked runs
None

Not reported / unresolved: Benchmark version and exact realized training/evaluation n per assay are not reported.; BALM-PPI provider and version are not explicitly reported.

AI-assisted double-pass extraction; values are limited to independently supported claims.

Evidence
  • table: Table S6
    Supports: /relation_type
  • table: Table S6
    Supports: /benchmark_id
  • section: 2.4
    Supports: /scope
  • section: 2.4
    Supports: /scope
  • table: Table S5
    Supports: /scope
  • table: Table S6
    Supports: /scope
  • table: Table S6
    Supports: /model_ids
  • table: Table S6
    Supports: /metric_labels

fine tuning

explainable-protein-protein-binding-affinity-predictio-abbibench-11-use

Non-evaluationunknown

Work: Explainable protein-protein binding affinity prediction via fine-tuning protein language models · source version explainable-protein-protein-binding-affinity-predictio-biorxiv-version-posted-2

Selection
not applicable · Few-Shot 20%
Models
BALM-PPI
Metrics
RMSE
Linked runs
None

Not reported / unresolved: Benchmark version and exact realized training/evaluation n per assay are not reported.; BALM-PPI provider and version are not explicitly reported.

AI-assisted double-pass extraction; values are limited to independently supported claims.

Evidence
  • table: Table S6
    Supports: /relation_type
  • table: Table S6
    Supports: /benchmark_id
  • section: 2.4
    Supports: /scope
  • section: 2.4
    Supports: /scope
  • table: Table S5
    Supports: /scope
  • table: Table S6
    Supports: /scope
  • table: Table S6
    Supports: /model_ids
  • table: Table S6
    Supports: /metric_labels

fine tuning

explainable-protein-protein-binding-affinity-predictio-abbibench-12-use

Non-evaluationunknown

Work: Explainable protein-protein binding affinity prediction via fine-tuning protein language models · source version explainable-protein-protein-binding-affinity-predictio-biorxiv-version-posted-2

Selection
not applicable · Few-Shot 30%
Models
BALM-PPI
Metrics
RMSE
Linked runs
None

Not reported / unresolved: Benchmark version and exact realized training/evaluation n per assay are not reported.; BALM-PPI provider and version are not explicitly reported.

AI-assisted double-pass extraction; values are limited to independently supported claims.

Evidence
  • table: Table S6
    Supports: /relation_type
  • table: Table S6
    Supports: /benchmark_id
  • section: 2.4
    Supports: /scope
  • section: 2.4
    Supports: /scope
  • table: Table S5
    Supports: /scope
  • table: Table S6
    Supports: /scope
  • table: Table S6
    Supports: /model_ids
  • table: Table S6
    Supports: /metric_labels

Evaluation registry

Works and run settings

A setting change—scope, prompt, tools, budget, grader, or repeats—creates a separate run. Charts never cross a comparability group.

No normalized evaluation run is published yet. Creator evidence is still attached below.

Evidence and change history

Source locators remain visible; expand an item to inspect the exact Registry fields it supports.

AbBiBench: A Benchmark for Antibody Binding Affinity Maturation and Design · section: Supplement 1 · Supports 2 fields

Open source →

  • /access/artifacts
  • /access/level
AbBiBench: A Benchmark for Antibody Binding Affinity Maturation and Design · repository-path: official-artifact-context.json: AbBibench/Antibody_Binding_Benchmark_Dataset · Supports 1 field

Open source →

  • /access/level
AbBiBench: A Benchmark for Antibody Binding Affinity Maturation and Design · repository-path: official-artifact-context.json: AbBibench/Antibody_Binding_Benchmark_Dataset · Supports 1 field

Open source →

  • /access/license
AbBiBench: A Benchmark for Antibody Binding Affinity Maturation and Design · section: Section 3.3 · Supports 2 fields

Open source →

  • /access/tasks
  • /access/level
AbBiBench: A Benchmark for Antibody Binding Affinity Maturation and Design · section: Introduction · Supports 1 field

Open source →

  • /aliases
AbBiBench: A Benchmark for Antibody Binding Affinity Maturation and Design · section: Section 3 · Supports 1 field

Open source →

  • /capabilities
AbBiBench: A Benchmark for Antibody Binding Affinity Maturation and Design · section: Introduction · Supports 1 field

Open source →

  • /domains
AbBiBench: A Benchmark for Antibody Binding Affinity Maturation and Design · section: Section 3.3 · Supports 1 field

Open source →

  • /kind
AbBiBench: A Benchmark for Antibody Binding Affinity Maturation and Design · figure: Figure 2 · Supports 1 field

Open source →

  • /modalities
AbBiBench: A Benchmark for Antibody Binding Affinity Maturation and Design · section: Section 3 · Supports 1 field

Open source →

  • /name
AbBiBench: A Benchmark for Antibody Binding Affinity Maturation and Design · page: Author affiliations · Supports 1 field

Open source →

  • /organizations
AbBiBench: A Benchmark for Antibody Binding Affinity Maturation and Design · figure: Figure 2 · Supports 1 field

Open source →

  • /summary
AbBiBench: A Benchmark for Antibody Binding Affinity Maturation and Design · other: arXiv API bibliographic metadata (Resolved from the canonical paper identifier during intake.) · Supports 1 field

Open source →

  • /release_date
AbBiBench: A Benchmark for Antibody Binding Affinity Maturation and Design · section: Introduction · Supports 4 fields

Open source →

  • /task_counts/total
  • /task_counts/basis
  • /versions/0/task_counts/total
  • /versions/0/task_counts/basis
AbBiBench: A Benchmark for Antibody Binding Affinity Maturation and Design · repository-path: official-artifact-context.json: AbBibench/Antibody_Binding_Benchmark_Dataset · Supports 2 fields

Open source →

  • /resources
  • /access/level
AbBiBench: A Benchmark for Antibody Binding Affinity Maturation and Design · page: arXiv footer · Supports 1 field

Open source →

  • /resources/0
AbBiBench: A Benchmark for Antibody Binding Affinity Maturation and Design · page: arXiv footer · Supports 2 fields

Open source →

  • /latest_version
  • /versions/0
AbBiBench: A Benchmark for Antibody Binding Affinity Maturation and Design · table: Table 1 · Supports 1 field

Open source →

  • /task_counts/subsets

Unresolved field claims

  • /access/levelProvisional · medium — The official resource and source-located access descriptions are public, but fully-open is an Atlas-controlled classification rather than a label stated verbatim by the creator source.
    Evidence: abbibench-automated-metadata-2-evidence, abbibench-automated-metadata-4-evidence, abbibench-automated-metadata-1-evidence, abbibench-automated-resource-evidence
  • /task_counts/subsetsConflicted · high — The owner approved the independently supported root total while all conflicted inventory subcounts were excluded from publication.
    Evidence: abbibench-automated-count-conflict-evidence

View source-level modification history on GitHub →