rewire.it
model · family

DNABERT-2

DNABERT-2 is a DNA encoder pretrained on sequences from multiple species. It supplies representations that can be adapted to genomic tasks.

0 evaluations · 0 metric rows

Shared profile: DNABERT-2. This page retains the exact record and its evaluation context.

At a glance

Explanatory profile: source reviewed · Automated source review, 2026-09-16. This does not change the review status of its results.

Inputs, outputs and configuration
PropertyDescription and evidence
Released modelDNABERT-2-117MMAGICS-LAB/DNABERT_2 official source · README.md: Introduction, Model and Data, Quick Start, Pre-Training and Fine-tune
ObjectiveMasked-language pretrainingMAGICS-LAB/DNABERT_2 official source · README.md: Introduction, Model and Data, Quick Start, Pre-Training and Fine-tune
Configuration in this record117MMAGICS-LAB/DNABERT_2 official source · README.md: Introduction, Model and Data, Quick Start, Pre-Training and Fine-tune
Training dataNot extracted or verified for this record.
Context limitsNot extracted or verified for this record.
AccessNot extracted or verified for this record.
Code licenceNot extracted or verified for this record.
Weights licenceNot extracted or verified for this record.

How it works

Conceptual procedure

Schematic of the documented input, computation and output; not an executable configuration.

Conceptual procedureDNA sequence. Then: Byte-pair tokens. Then: BERT encoder with ALiBi. Then: Token or pooled embeddings. Then: Task-specific headDNA sequenceByte-pair tokensBERT encoder with ALiBiToken or pooled embeddingsTask-specific head
Read the diagram as text
  1. DNA sequence
  2. Byte-pair tokens
  3. BERT encoder with ALiBi
  4. Token or pooled embeddings
  5. Task-specific head
MAGICS-LAB/DNABERT_2 official source · README.md: Introduction, Model and Data, Quick Start, Pre-Training and Fine-tune

DNA is compressed into variable-length byte-pair tokens. A BERT-style encoder uses ALiBi positional biases; masked-language pretraining learns contextual features. Sequence pooling or a separately trained prediction head produces task outputs.

MAGICS-LAB/DNABERT_2 official source · README.md: Introduction, Model and Data, Quick Start, Pre-Training and Fine-tune

Benchmarks and results

Release 2026-09-16-d74d282221a9 · 0 evaluations · 0 metric rows. Different protocols are not a single leaderboard.

No evaluations linked in this release.

Strengths and limitations

Strengths supported by sources

  • One released encoder can support embedding extraction and supervised adaptation across several genomic tasks.MAGICS-LAB/DNABERT_2 official source · README.md: Introduction, Model and Data, Quick Start, Pre-Training and Fine-tune

Limitations and conditions

  • An encoder embedding is not a splice-impact prediction. Pooling, sequence context and the supervised head are part of the evaluated pipeline.MAGICS-LAB/DNABERT_2 official source · README.md: Introduction, Model and Data, Quick Start, Pre-Training and Fine-tune
Profile review details

Primary project documentation or paper inspected for the explanatory claims and cited locations. Reviewed coverage concerns this narrative, not complete metadata, independent reproduction or a performance ranking.

Stable record: catalog-model-dnabert-2

Sources and history

Release 2026-09-16-d74d282221a9 · Record review: discovered

Download this release
Technical metadata and extraction receipts

Stable ID: catalog-model-dnabert-2

areas
dna-genomes
method types
foundation model
entity level
family
version
117M
reported name
DNABERT-2
access
Public checkpoint; remote model code needs review before local use.
method type
foundation model
missing metadata
checkpoint revision: not_yet_extracted; training data: not_yet_extracted; licence: not_yet_extracted
Related records

Suggest a correction