rewire.it
Configuration

DNABERT-2

DART-Eval tests DNABERT-2 on regulatory DNA under separately defined zero-shot, probed and fine-tuned protocols.

SourcesDART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA · Table 2 (models); Section 3.2 zero-shot analysis; Table 3 regulatory-element identification; Appendix B evaluation procedures

1 evaluation · 1 metric row

How it worksEvaluated procedure (conceptual)
Evaluated procedure (conceptual)1. Regulatory DNA and matched controls defined by the DART-Eval task. Then: 2. DNABERT-2. Then: 3. Regulatory-element scores or task-specific predictions, according to the evaluation protocolEvaluated procedure (conceptual)1. Regulatory DNA and matched controls defined by the DART-Eval task. Then: 2. DNABERT-2. Then: 3. Regulatory-element scores or task-specific predictions, according to the evaluation protocolEvaluated procedure (conceptual)1. Regulatory DNA and matched controls defined by the DART-Eval task. Then: 2. DNABERT-2. Then: 3. Regulatory-element scores or task-specific predictions, according to the evaluation protocol

Conceptual input–method–output guide. Check the procedure text and linked evaluation for fitted components, additional inputs and exact settings.

SourcesDART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA · Table 2 (models); Section 3.2 zero-shot analysis; Table 3 regulatory-element identification; Appendix B evaluation procedures

At a glance

Model type

DNA sequence transformer; this record is the paper-specific evaluated configuration.

SourcesMAGICS-LAB/DNABERT_2 README.md · README.md model description

limited source coverage · Automated source review, 2026-09-16. All specifications and missing details

Evaluations and results

Release 2026-09-17-d277315f7d76 · 1 evaluation · 1 metric row. Different protocols are not a single leaderboard.

Results grouped by the exact reported evaluation
Metric and findingCoverage and uncertaintyEvidence
DNABERT-2: regulatory element identification

zero-shot likelihood ranking: higher likelihood for cCRE than matched control

Independent external evaluation · Evaluation metadata: needs review

0.876 accuracy

Unit: fraction · Direction: unknown

Uncertainty: not reported in legacy extract

Scored: Not reported · Eligible: Not reported

source checkedDART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA · Table 3 (PDF page 5), DNABERT-2 row, Zero-Shot Accuracy column

Source checking is not independent reproduction.

How it works

How the evaluated method works

The masked DNA transformer uses byte-pair tokenisation. The linked regulatory-element row uses the paper’s zero-shot procedure; trained probing or fine-tuning rows are distinct evaluations.

SourcesDART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA · Table 2 (models); Section 3.2 zero-shot analysis; Table 3 regulatory-element identification; Appendix B evaluation procedures
Underlying method and version boundaries

DNABERT-2 replaces overlapping k-mer tokens with byte-pair encoding and uses ALiBi positional biases. The official 117M model produces 768-dimensional token representations; downstream classifiers and pooling choices are separate configuration details.

SourcesMAGICS-LAB/DNABERT_2 README.md · README.md; introduction, model description, pretrained-model and usage sections at pinned revision
What was evaluated

The linked evaluation record identifies DNABERT-2: regulatory element identification. Its dataset, split, adaptation and evidence origin remain attached to the reported results.

SourcesDART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA · The named evaluation’s methods and comparison table; exact preserved evaluation IDs: evaluation-b2-dart-eval-regulatory-2024

Strengths and limitations

Profile review details

Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.

Stable record: reported-model-28413ae1766316

Specifications

Inputs, training, access and other details

Explanatory profile: limited source coverage · Automated source review, 2026-09-16. Review applies to the cited claims; unresolved fields are listed below. Numerical results retain their own review status.

Inputs, outputs and configuration
PropertyDescription and evidence
Model typeDNA sequence transformer; this record is the paper-specific evaluated configuration.
SourcesMAGICS-LAB/DNABERT_2 README.md · README.md model description
Architecture / procedureThe masked DNA transformer uses byte-pair tokenisation. The linked regulatory-element row uses the paper’s zero-shot procedure; trained probing or fine-tuning rows are distinct evaluations.
SourcesDART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA · Table 2 (models); Section 3.2 zero-shot analysis; Table 3 regulatory-element identification; Appendix B evaluation procedures
Biological inputsRegulatory DNA and matched controls defined by the DART-Eval task
SourcesDART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA · Table 2 (models); Section 3.2 zero-shot analysis; Table 3 regulatory-element identification; Appendix B evaluation procedures
OutputsRegulatory-element scores or task-specific predictions, according to the evaluation protocol
SourcesDART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA · Table 2 (models); Section 3.2 zero-shot analysis; Table 3 regulatory-element identification; Appendix B evaluation procedures
Parameters117 million parameters
SourcesDART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA · Table 2 (models); Section 3.2 zero-shot analysis; Table 3 regulatory-element identification; Appendix B evaluation procedures
Known versions / configurationDNABERT-2 is the comparison-table label; that label does not specify an immutable weight revision. · Not reported in inspected sources
SourcesDART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA · Model identification in the comparison table and corresponding Methods; immutable checkpoint revision is not supplied by the table label.
Training data / fittingThe benchmark identifies multispecies genomic pretraining; task adaptation is reported separately.
SourcesDART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA · Table 2 (models); Section 3.2 zero-shot analysis; Table 3 regulatory-element identification; Appendix B evaluation procedures
Context limitsA maximum input/context length for this exact evaluated configuration is not established by the inspected sources. · Not reported in inspected sources
Sources (2)DART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA; MAGICS-LAB/DNABERT_2 README.md · Complete primary text and named comparison table; inspected for explicit maximum input length (dataset lengths and family-wide limits are not substituted); README.md at pinned repository revision
AccessOfficial upstream implementation and usage documentation: https://github.com/MAGICS-LAB/DNABERT_2/blob/f25bed9ee20db966dff39e5c1571249d04e36404/README.md. This pinned documentation revision is not automatically the evaluated weight revision.
SourcesMAGICS-LAB/DNABERT_2 README.md · README.md; installation, model download and usage instructions
Code licenceApache 2.0 (upstream repository code at the cited revision; this does not establish every dependency or historical checkpoint licence).
SourcesMAGICS-LAB/DNABERT_2 LICENSE · LICENSE; complete licence text
Weights licenceThe inspected model-access documentation does not explicitly identify terms for this exact evaluated checkpoint or fitted head; repository code terms are shown separately. · Not reported in inspected sources
SourcesMAGICS-LAB/DNABERT_2 README.md · README.md; checkpoint/access documentation and licence scope

Evidence table

Inspect claims, sources and review details

Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.

One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.

21 evidence rows matching the loaded filters

Claims, original sources and review scope · Release 2026-09-17-d277315f7d76
Property and statementOriginal source and locationReview and provenance
Diagram caption

Conceptual input–method–output guide. Check the procedure text and linked evaluation for fitted components, additional inputs and exact settings.

Individual claims
DART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA

Original source ↗

Table 2 (models); Section 3.2 zero-shot analysis; Table 3 regulatory-element identification; Appendix B evaluation procedures

Version: NeurIPS 2024 Datasets and Benchmarks Track proceedings
Retrieved: 2026-09-16T10:38:57.558203+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.

Field: attributes.profile.diagram.caption

Source artifact SHA-256: e5aee5b1f7cc6fd961b1d2a131d02cf243b79e091d5e418fbabee7fde9b39b22

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Diagram steps

["Regulatory DNA and matched controls defined by the DART-Eval task","DNABERT-2","Regulatory-element scores or task-specific predictions, according to the evaluation protocol"]

Individual claims
DART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA

Original source ↗

Table 2 (models); Section 3.2 zero-shot analysis; Table 3 regulatory-element identification; Appendix B evaluation procedures

Version: NeurIPS 2024 Datasets and Benchmarks Track proceedings
Retrieved: 2026-09-16T10:38:57.558203+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.

Field: attributes.profile.diagram.steps

Source artifact SHA-256: e5aee5b1f7cc6fd961b1d2a131d02cf243b79e091d5e418fbabee7fde9b39b22

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Diagram title

Evaluated procedure (conceptual)

Individual claims
DART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA

Original source ↗

Table 2 (models); Section 3.2 zero-shot analysis; Table 3 regulatory-element identification; Appendix B evaluation procedures

Version: NeurIPS 2024 Datasets and Benchmarks Track proceedings
Retrieved: 2026-09-16T10:38:57.558203+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.

Field: attributes.profile.diagram.title

Source artifact SHA-256: e5aee5b1f7cc6fd961b1d2a131d02cf243b79e091d5e418fbabee7fde9b39b22

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Model type

DNA sequence transformer; this record is the paper-specific evaluated configuration.

Individual claims
MAGICS-LAB/DNABERT_2 README.md

Original source ↗

README.md model description

Version: f25bed9ee20db966dff39e5c1571249d04e36404
Retrieved: 2026-09-16T20:00:02.624261+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.

Field: attributes.profile.facts.0.value

Source artifact SHA-256: 734a8cec5f667d74d421bf3b273ad7e256216109636da45aa7ceba21cd34de16

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Architecture / procedure

The masked DNA transformer uses byte-pair tokenisation. The linked regulatory-element row uses the paper’s zero-shot procedure; trained probing or fine-tuning rows are distinct evaluations.

Individual claims
DART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA

Original source ↗

Table 2 (models); Section 3.2 zero-shot analysis; Table 3 regulatory-element identification; Appendix B evaluation procedures

Version: NeurIPS 2024 Datasets and Benchmarks Track proceedings
Retrieved: 2026-09-16T10:38:57.558203+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.

Field: attributes.profile.facts.1.value

Source artifact SHA-256: e5aee5b1f7cc6fd961b1d2a131d02cf243b79e091d5e418fbabee7fde9b39b22

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Weights licence

The inspected model-access documentation does not explicitly identify terms for this exact evaluated checkpoint or fitted head; repository code terms are shown separately.

Individual claims
MAGICS-LAB/DNABERT_2 README.md

Original source ↗

README.md; checkpoint/access documentation and licence scope

Version: f25bed9ee20db966dff39e5c1571249d04e36404
Retrieved: 2026-09-16T20:00:02.624261+00:00

unreported

automated source review · 2026-09-16

Audit details

Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.

Field: attributes.profile.facts.10.value

Source artifact SHA-256: 734a8cec5f667d74d421bf3b273ad7e256216109636da45aa7ceba21cd34de16

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Biological inputs

Regulatory DNA and matched controls defined by the DART-Eval task

Individual claims
DART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA

Original source ↗

Table 2 (models); Section 3.2 zero-shot analysis; Table 3 regulatory-element identification; Appendix B evaluation procedures

Version: NeurIPS 2024 Datasets and Benchmarks Track proceedings
Retrieved: 2026-09-16T10:38:57.558203+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.

Field: attributes.profile.facts.2.value

Source artifact SHA-256: e5aee5b1f7cc6fd961b1d2a131d02cf243b79e091d5e418fbabee7fde9b39b22

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Outputs

Regulatory-element scores or task-specific predictions, according to the evaluation protocol

Individual claims
DART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA

Original source ↗

Table 2 (models); Section 3.2 zero-shot analysis; Table 3 regulatory-element identification; Appendix B evaluation procedures

Version: NeurIPS 2024 Datasets and Benchmarks Track proceedings
Retrieved: 2026-09-16T10:38:57.558203+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.

Field: attributes.profile.facts.3.value

Source artifact SHA-256: e5aee5b1f7cc6fd961b1d2a131d02cf243b79e091d5e418fbabee7fde9b39b22

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Parameters

117 million parameters

Individual claims
DART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA

Original source ↗

Table 2 (models); Section 3.2 zero-shot analysis; Table 3 regulatory-element identification; Appendix B evaluation procedures

Version: NeurIPS 2024 Datasets and Benchmarks Track proceedings
Retrieved: 2026-09-16T10:38:57.558203+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.

Field: attributes.profile.facts.4.value

Source artifact SHA-256: e5aee5b1f7cc6fd961b1d2a131d02cf243b79e091d5e418fbabee7fde9b39b22

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Known versions / configuration

DNABERT-2 is the comparison-table label; that label does not specify an immutable weight revision.

Individual claims
DART-Eval: A Comprehensive DNA Language Model Evaluation Benchmark on Regulatory DNA

Original source ↗

Model identification in the comparison table and corresponding Methods; immutable checkpoint revision is not supplied by the table label.

Version: NeurIPS 2024 Datasets and Benchmarks Track proceedings
Retrieved: 2026-09-16T10:38:57.558203+00:00

unreported

automated source review · 2026-09-16

Audit details

Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.

Field: attributes.profile.facts.5.value

Source artifact SHA-256: e5aee5b1f7cc6fd961b1d2a131d02cf243b79e091d5e418fbabee7fde9b39b22

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Sources and history

Release 2026-09-17-d277315f7d76 · Record review: needs review

3 source records and release historyDownload this release
Technical metadata and extraction receipts

Stable ID: reported-model-28413ae1766316

areas
dna-genomes
entity level
method
version
not stated in table
reported name
DNABERT-2
historical missing metadata
checkpoint revision: not_reported_in_legacy_extract; training data: not_reported_in_legacy_extract; licence: not_reported_in_legacy_extract
metadata review scope
historical_missing_metadata preserves the original discovery state. Current descriptive evidence and missingness are recorded in profile.facts; numerical-result review is separate.
legacy kinds
model
entity classification
review date: 2026-09-17; rationale: This source-scoped entry preserves the method/configuration actually named in an evaluation. It is neither a global family identity nor proof of an immutable checkpoint; the linked evaluation retains adaptation, fitting and scoring details.; source ids: dart-eval-regulatory-2024; evidence-reported-base-dnabert2-readme-md; source locator: Table 2 (models); Section 3.2 zero-shot analysis; Table 3 regulatory-element identification; Appendix B evaluation procedures | README.md model description; ambiguities: Configuration means the source-labelled evaluated identity. It does not establish missing checkpoint hashes, default settings or equivalence to same-named records in other papers.
Related records

Suggest a correction