rewire.it
model · family

DNABERT-2

DNABERT-2 is a DNA encoder pretrained on sequences from multiple species. It supplies representations that can be adapted to genomic tasks.

0 evaluations · 0 metric rows

At a glance

Explanatory profile: source reviewed · Automated source review, 2026-09-16. This does not change the review status of its results.

Inputs, outputs and configuration
PropertyDescription and evidence
Released modelDNABERT-2-117MMAGICS-LAB/DNABERT_2 official source · README.md: Introduction, Model and Data, Quick Start, Pre-Training and Fine-tune
ObjectiveMasked-language pretrainingMAGICS-LAB/DNABERT_2 official source · README.md: Introduction, Model and Data, Quick Start, Pre-Training and Fine-tune
Training dataNot extracted or verified for this record.
Context limitsNot extracted or verified for this record.
AccessNot extracted or verified for this record.
Code licenceNot extracted or verified for this record.
Weights licenceNot extracted or verified for this record.

Versions and evaluated configurations

How it works

Conceptual procedure

Schematic of the documented input, computation and output; not an executable configuration.

Conceptual procedureDNA sequence. Then: Byte-pair tokens. Then: BERT encoder with ALiBi. Then: Token or pooled embeddings. Then: Task-specific headDNA sequenceByte-pair tokensBERT encoder with ALiBiToken or pooled embeddingsTask-specific head
Read the diagram as text
  1. DNA sequence
  2. Byte-pair tokens
  3. BERT encoder with ALiBi
  4. Token or pooled embeddings
  5. Task-specific head
MAGICS-LAB/DNABERT_2 official source · README.md: Introduction, Model and Data, Quick Start, Pre-Training and Fine-tune

DNA is compressed into variable-length byte-pair tokens. A BERT-style encoder uses ALiBi positional biases; masked-language pretraining learns contextual features. Sequence pooling or a separately trained prediction head produces task outputs.

MAGICS-LAB/DNABERT_2 official source · README.md: Introduction, Model and Data, Quick Start, Pre-Training and Fine-tune

Benchmarks and results

Release 2026-09-16-d74d282221a9 · 0 evaluations · 0 metric rows. Different protocols are not a single leaderboard.

No evaluations linked in this release.

Strengths and limitations

Strengths supported by sources

  • One released encoder can support embedding extraction and supervised adaptation across several genomic tasks.MAGICS-LAB/DNABERT_2 official source · README.md: Introduction, Model and Data, Quick Start, Pre-Training and Fine-tune

Limitations and conditions

  • An encoder embedding is not a splice-impact prediction. Pooling, sequence context and the supervised head are part of the evaluated pipeline.MAGICS-LAB/DNABERT_2 official source · README.md: Introduction, Model and Data, Quick Start, Pre-Training and Fine-tune
Profile review details

Primary project documentation or paper inspected for the explanatory claims and cited locations. Reviewed coverage concerns this narrative, not complete metadata, independent reproduction or a performance ranking.

Stable record: discovery-model-dnabert-2

Pipelines using this model

These evaluated pipelines include additional processing or trained components. Their results are not assigned to the underlying model.

Applicable tests and references

Applicability is distinct from a completed evaluation.

  • GUE · Proposed association

Sources and history

Release 2026-09-16-d74d282221a9 · Record review: discovered

Download this release
Technical metadata and extraction receipts

Stable ID: discovery-model-dnabert-2

areas
genomics
access
official_source_linked
benchmark applicability
candidate; not evidence of a reported evaluation
candidate benchmark ids
discovery-benchmark-gue
entity level
family
missing metadata
checkpoint: unextracted; code licence: unextracted; parameters: unextracted; training cutoff: unextracted; training data: unextracted; version: unextracted; weights licence: unextracted
reported name
DNABERT-2
version
Not reported
Related records

Suggest a correction