AI-assisted double-pass paper intake (paper-intake:explainable-protein-protein-binding-affinity-predictio); production inclusion required owner approval of the final head SHA.
Versioned change log
The registry remembers what changed.
Additions, corrections, deprecations, version changes, and source failures are retained rather than silently overwritten.
Monthly archive
AI-assisted double-pass paper intake (paper-intake:abbibench-a-benchmark-for-antibody-binding-affinity-ma); production inclusion required owner approval of the final head SHA.
AI-assisted double-pass paper intake (paper-intake:ab-bind-antibody-binding-mutational-database-for-compu); production inclusion required owner approval of the final head SHA.
Commit-pinned SOAR official repository results normalized for the 1,191-entry SOAR-RNA formal subset with separate zero-shot and two-call chain-of-thought protocols.
AI-assisted double-pass paper intake (paper-intake:single-cell-omics-arena-evaluation-of-large-language-m); production inclusion required owner approval of the final head SHA.
Commit-pinned BioSecBench-Surveillance official results normalized with an independent local Codex double-pass; the creator preprint and repository result snapshot remain separate source versions.
AI-assisted double-pass paper intake (paper-intake:comprehensive-benchmark-of-differential-transcript-usa); production inclusion required owner approval of the final head SHA.
AI-assisted double-pass paper intake (paper-intake:a-benchmark-comparison-of-crisprn-guide-rna-design-alg); production inclusion required owner approval of the final head SHA.
AI-assisted double-pass paper intake (paper-intake:enhancing-molecular-property-prediction-with-auxiliary); production inclusion required owner approval of the final head SHA.
AI-assisted double-pass paper intake (paper-intake:an-integrative-approach-to-protein-sequence-design-thr); production inclusion required owner approval of the final head SHA.
AI-assisted double-pass paper intake (paper-intake:crafted-experiments-to-evaluate-feature-selection-meth); production inclusion required owner approval of the final head SHA.
AI-assisted double-pass paper intake (paper-intake:biosecbench-surveillance-a-verifiable-benchmark-for-ai); production inclusion required owner approval of the final head SHA.
AI-assisted double-pass paper intake (paper-intake:ppb-affinity-protein-protein-binding-affinity-dataset); production inclusion required owner approval of the final head SHA.
AI-assisted double-pass paper intake (paper-intake:system-card-claude-opus-5); production inclusion required owner approval of the final head SHA.
AI-assisted double-pass paper intake (paper-intake:scbench-evaluating-ai-agents-on-single-cell-rna-seq-an); production inclusion required owner approval of the final head SHA.
Migrated guarded paper intake to an owner-triggered local Codex workflow while retaining weekly Europe PMC, Crossref, and arXiv candidate discovery. Two fresh read-only local sessions now perform extraction and independent verification; a local golden receipt gates production, deterministic code writes records, and an exact-head-SHA owner comment gates each paper PR. GitHub Actions no longer sends paper content to a model or requires an OpenAI API secret or GitHub App.
Released v1.3.1 as a dependency-only maintenance snapshot with grouped compatible Python, Astro, Node type, pnpm, and GitHub Pages action upgrades; retained TypeScript 5.9.3 after the TypeScript 7 update failed the project test suite.
Released v1.3.0 with versioned Works and BenchmarkUse relationships; normalized Anthropic's partial BixBench claim, SpatialBench paper/current snapshots, an explicitly third-party SpatialBench summary, and three private internal life-science delta evaluations; added a draft-PR paper intake system and Paper Explorer without inferring missing settings or plot values.
Released v1.2.0 with the Scientific Task Atlas taxonomy, evidence-backed benchmark mappings, normalized task-coverage exports, task explorer and detail pages, Chinese task-system guidance, and seven additional creator-audited benchmark families.
Added seven creator-audited benchmark families spanning protein representation learning, genomic sequence classification, RNA representation learning, molecular property prediction, 3D molecular learning, molecular generation, and single-cell integration; froze living implementations to commit-pinned snapshots and classified scientific tasks without flattening heterogeneous endpoints or metrics.
Released v1.1.0 with field-level audits for fourteen launch families, structured versions and evidence, normalized official evaluation settings and results, and an explicit machine-validated VirBench legacy exception.
Closed the current audit pass with 14 of 15 launch families field-audited; retained VirBench as an intentionally deferred legacy record and added a machine-validated registry exception so no other unaudited record can enter production silently.
Audited SCIGYM against the final NeurIPS paper, official project page, commit-pinned creator implementations, and immutable benchmark/evaluation Parquet snapshots; separated the 350 released SBML systems into 137 small and 213 large formal tracks, limited creator evaluation claims to the small track, registered all six exact model versions and 42 printed Table 1 values, reproduced three episodes per model-system pair, separated the no-tool zero-shot baseline, corrected the first public release date, and removed unsupported wet-lab, omics, protein/binding, and license claims.
Audited BLADE against the current creator manuscript, official project site, peer-reviewed paper, and commit-pinned package; separated 12 source research-question/dataset pairs from 188 MCQs and 536 ground-truth decision references, registered MCQ and end-to-end generation as formal tracks, corrected code/data licensing, limited biological coverage to four explicit source questions with no protein/binding/omics tasks, and split the creator evaluation into zero-shot MCQ, 40-repeat one-shot generation, and 20-repeat ten-step ReAct protocols with all 14 printed F1 results and confidence intervals.
Audited CompBioBench v1 against the immutable 100-row Zenodo TSV, public Hugging Face data mirror, official runner, leaderboard Space, and creator preprint; registered eight domain, five style, four difficulty, and internet-required partitions, corrected component licenses and access to partially open because answers/grading remain private, treated the single PDB projection task as protein structure without inferring binding coverage, and split all creator results by exact agent, effort, repeat, timeout, full/hardest scope, cost, time, and non-agentic protocol.
Audited BioMysteryBench across versions; separated the current gated v11 release (90 problems, 73/17) from the superseded v8 creator evaluation (99 problems, 76/23), pinned the open preview and full-set commits plus three official result figures, corrected access and grader claims, and registered all ten printed subset scores across five models without inferring numeric confidence bounds.
Audited GeneBench-Pro as 129 synthetic problems partitioned into 10 public, 50 Artificial Analysis, and 69 internal-holdout problems; registered the 10-domain and 21-terminal-subdomain atlas plus 82/47 review strata, pinned the ten-problem public package and its reference grader, retained its CC-BY-4.0 versus MIT license conflict, and normalized all 60 creator-reported model configurations across 13 exact effort/repeat groups with confidence intervals, valid-attempt handling, tools, and binary-grader semantics.
Audited LAB-Bench as 2,457 questions in eight broad categories; reproduced the 1,967 public and 490 private split across 31 versioned task files, preserved the README's conflicting 30-subtask claim, registered every formal DbQA and SeqQA child track, transcribed all 31 creator multiple-choice result rows and three expert-graded open-response studies, and separated Anthropic 10-shot and crop-tool evaluations by protocol and comparability group.
Audited Biology-Instructions as 21 formal tasks rather than 21 questions; registered 6 DNA, 6 RNA, 5 protein, and 4 multi-molecule child tracks with train/validation/test counts, separated open-source, closed-source, and creator prompts into 63 runs, captured 345 creator-paper results, kept protein design out of scope, corrected access and license claims, and preserved the paper/repository conflicts over aggregate test totals, Stage-3 rows, and evaluator registration keys.
Audited ProteinLMBench against arXiv v2, its complete commit-pinned Hugging Face history and current JSON, and the official runner; corrected the release date to 2024-04-29, recorded the 2-to-10-option distribution and official license conflicts, kept topical design/binding counts unreported, and registered all 18 creator-evaluated systems with 36 accuracy and inference-time results.
Audited original FLIP as 15 dataset-by-split tasks across three formal landscape tracks; separated task counts from sequence-example counts, registered all 15 creator evaluation protocols and 155 main-table Spearman results, isolated the two optimistic sampled splits, pinned the official implementation, and resolved the Meltome Human-cell total to 7,158 using the versioned CSV while retaining the paper's 7,156 discrepancy.
Audited CAMEO as a weekly rolling service, corrected the creator-paper DOI and 2012 origin, separated the current complex-only category from discontinued legacy categories, preserved the bounded 7,150-target 2024 study and common subsets, and registered exact evaluated systems even where no scalar result is available.
Audited CASP as a round-based competition; separated the active CASP17 rolling snapshot from completed CASP16, registered monomer, multimer, ligand, and immune-complex tracks, preserved release/target/evaluation-unit count bases, and normalized five creator-assessor protocols without inventing a cross-category total or leaderboard.
Audited ProteinGym v1.0-v1.3, separated assay/protein/variant count units into four formal tracks, recorded the Binding function category change from 14 to 13 DMS substitution assays, retained the official v1.3 indel count conflict (66 in the release archive versus 74 in the README), and normalized the 50-model v1.0 zero-shot DMS substitution evaluation.
Completed the LifeSciBench field-level audit; preserved 750/136/62, retained binding counts as Not reported, replaced an unsupported numbered version with an initial-release snapshot, and corrected unreported shot/repeat/grader settings.
Began the v1.1 field-level audit cycle with compatible versions, resource pins, structured evidence, audit status, provisional/conflicted claims, result confidence, and expanded source monitoring.
Initial v1.0 registry with fifteen benchmark families, normalized official evaluations, schemas, exports, and the public atlas.
Recorded 136 protein-primary-domain tasks, 62 protein design/optimization tasks, and explicit Not reported status for binding counts.
Recorded the 99-question total, 76/23 human split, five-episode protocol, and accuracy/consistency metrics.
Audited BixBench across the original v1.0 paper snapshot and current v1.5 release; corrected the code/data license to Apache-2.0, separated 296 questions/53 capsules from 205 current questions, retained the unresolved 60-notebook claim beside 59 referenced capsule UUIDs and 64 stored archives, registered the 83/61/61 verifier partition and overlapping omics categories, and split creator agentic, ablation, and zero-shot protocols without digitizing unlabeled plots.