rewire.it

Biological benchmarks

Explore evaluation suites and challenges, what they measure, and the models tested against them.

30 benchmark records in release 2026-09-22-f58a0f1d267f.

Broad tasks, specific protocols and evaluators have their own records. Follow each benchmark to its procedures and results.

  • ATOM3D

    Molecular interactions

    ATOM3D provides molecular-structure datasets and utilities for task-specific evaluation.

  • BEACON

    RNA

    BEACON compares RNA representations across structural and functional downstream tasks.

  • BEELINE

    Biological networks

    BEELINE compares inferred gene-regulatory edge rankings with reference networks.

  • BEND

    Genomics

    BEND evaluates DNA representations using task-specific genomic annotations and explicit split membership.

  • CAFA

    Protein function

    CAFA evaluates prospective protein-function predictions against annotations that become available after prediction submission.

  • CAMI

    Microbiome

    CAMI is a community benchmark program for metagenomic computational methods.

  • CAPRI

    Protein structure

    CAPRI assesses blind predictions of protein-complex structures supplied before experimental publication.

  • CASP

    Protein structure

    CASP assesses structure-prediction methods through blind predictions and category-specific evaluation.

  • DART-Eval

    Genomics

    DART-Eval measures human regulatory-DNA representations under zero-shot, probing and fine-tuning regimes.

  • FLIP

    Protein function

    FLIP evaluates protein-sequence representations using multiple deliberately defined train/test splits.

  • FLIP2

    Protein function

    FLIP2 evaluates protein fitness prediction under deliberately shifted training and test distributions.

  • GENEB

    Genomics

    GENEB compares frozen genomic representations across classification tasks and label-budget regimes.

  • Genomic Benchmarks

    Genomics

    Genomic Benchmarks packages genomic sequence-classification datasets with explicit versions and supplied train/test folders.

  • GlycanML

    Glycomics

    GlycanML evaluates glycan learning across multiple classification and interaction tasks.

  • GUE

    Genomics

    GUE evaluates genome understanding across multiple datasets, task types and species.

  • HEST-Benchmark

    Spatial omics

    HEST-Benchmark tests prediction of gene expression from histological image representations.

  • MassSpecGym

    Metabolomics

    MassSpecGym separates spectrum-to-structure generation, candidate retrieval and structure-to-spectrum simulation.

  • mRNABench

    RNA

    mRNABench assesses genomic-model embeddings on transcript-specific expression, stability and regulatory tasks.

  • NABench

    RNA

    NABench compares nucleotide foundation models on measured DNA/RNA sequence effects under multiple adaptation settings.

  • Open Problems

    Single cell

    Open Problems is an extensible platform hosting benchmark tasks and their datasets.

  • PerturBench

    Single cell

    PerturBench evaluates predicted single-cell perturbation responses with explicit aggregation and metric choices.

  • PEtab benchmark collection

    Mechanistic biology

    The PEtab collection supports evaluation of computational methods for fitting mathematical models to observations.

  • PFMBench

    Protein function

    PFMBench is a configurable suite of protein-model downstream evaluations.

  • PLINDER

    Molecular interactions

    PLINDER supplies annotated protein–ligand systems and evaluation resources for docking.

  • ProteinBench

    Protein structure

    ProteinBench assesses multiple protein-model tasks using quality, novelty, diversity and robustness dimensions.

  • ProteinGym

    Protein function

    ProteinGym separates experimental variant-effect and clinical annotation tasks under supervised and zero-shot regimes.

  • scIB

    Single cell

    scIB is an atlas-level single-cell integration benchmark study covering RNA, chromatin-accessibility and simulated datasets. It compares batch removal with preservation of biological variation. The scib Python package and scib-pipeline implement its evaluation workflow.

  • TAPE

    Protein function

    TAPE evaluates protein representations through five supervised downstream tasks.

  • TDC molecular tasks

    Molecular interactions

    TDC organizes molecular prediction tasks into datasets and benchmark groups with explicit splitting and evaluation interfaces.

  • Virtual Cell Challenge 2026

    Single cell

    The 2026 Virtual Cell Challenge evaluates perturbation-response prediction in unseen cellular contexts.

This index reflects a dated catalogue, not an exhaustive census. Source checking does not mean independent reproduction; compare results only under compatible protocols, datasets and metrics.