rewire.it
Configuration

DNABERT-2

This genomic model is adapted to predict G-quadruplex-associated sequence regions.

SourcesBenchmarking DNA large language models on quadruplexes · Discussion and conclusions (paragraph 6); Results/LLM performance at the genome-wide level (paragraph 4)

1 evaluation · 1 metric row

How it worksEvaluated procedure (conceptual)
Evaluated procedure (conceptual)1. DNA sequences for G-quadruplex classification and genomic scanning. Then: 2. DNABERT-2. Then: 3. G4-associated sequence predictionsEvaluated procedure (conceptual)1. DNA sequences for G-quadruplex classification and genomic scanning. Then: 2. DNABERT-2. Then: 3. G4-associated sequence predictionsEvaluated procedure (conceptual)1. DNA sequences for G-quadruplex classification and genomic scanning. Then: 2. DNABERT-2. Then: 3. G4-associated sequence predictions

Conceptual input–method–output guide. Check the procedure text and linked evaluation for fitted components, additional inputs and exact settings.

SourcesBenchmarking DNA large language models on quadruplexes · Materials and methods/Data preparation (paragraph 2); Results/Interpretation of LLMs (paragraph 2)

At a glance

Model type

DNA sequence transformer; this record is the paper-specific evaluated configuration.

SourcesMAGICS-LAB/DNABERT_2 README.md · README.md model description

limited source coverage · Automated source review, 2026-09-16. All specifications and missing details

Evaluations and results

Release 2026-09-17-d277315f7d76 · 1 evaluation · 1 metric row. Different protocols are not a single leaderboard.

Results grouped by the exact reported evaluation
Metric and findingCoverage and uncertaintyEvidence
DNABERT-2: G-quadruplex classification

Pretrained model evaluated on KEx as reported in Table 5.

Independent external evaluation · Evaluation metadata: needs review

97.0 Accuracy

Unit: % · Direction: unknown

Uncertainty: ± 0.5

Scored: Not reported · Eligible: Not reported

source checkedBenchmarking DNA large language models on quadruplexes · Table 5, DNABERT-2 (117 M) row, Accuracy column

Source checking is not independent reproduction.

How it works

How the evaluated method works

The study fine-tunes genomic sequence models for G4 classification and applies them to whole-genome annotation, retaining each model’s native tokenisation.

SourcesBenchmarking DNA large language models on quadruplexes · Materials and methods/Data preparation (paragraph 2); Results/Interpretation of LLMs (paragraph 2)
Underlying method and version boundaries

DNABERT-2 replaces overlapping k-mer tokens with byte-pair encoding and uses ALiBi positional biases. The official 117M model produces 768-dimensional token representations; downstream classifiers and pooling choices are separate configuration details.

SourcesMAGICS-LAB/DNABERT_2 README.md · README.md; introduction, model description, pretrained-model and usage sections at pinned revision
What was evaluated

The linked evaluation record identifies DNABERT-2: G-quadruplex classification. Its dataset, split, adaptation and evidence origin remain attached to the reported results.

SourcesBenchmarking DNA large language models on quadruplexes · The named evaluation’s methods and comparison table; exact preserved evaluation IDs: evaluation-lit-005

Strengths and limitations

Limitations and conditions

Profile review details

Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.

Stable record: reported-model-ade36035f58f27

Specifications

Inputs, training, access and other details

Explanatory profile: limited source coverage · Automated source review, 2026-09-16. Review applies to the cited claims; unresolved fields are listed below. Numerical results retain their own review status.

Inputs, outputs and configuration
PropertyDescription and evidence
Model typeDNA sequence transformer; this record is the paper-specific evaluated configuration.
SourcesMAGICS-LAB/DNABERT_2 README.md · README.md model description
Architecture / procedureThe study fine-tunes genomic sequence models for G4 classification and applies them to whole-genome annotation, retaining each model’s native tokenisation.
SourcesBenchmarking DNA large language models on quadruplexes · Materials and methods/Data preparation (paragraph 2); Results/Interpretation of LLMs (paragraph 2)
Biological inputsDNA sequences for G-quadruplex classification and genomic scanning
SourcesBenchmarking DNA large language models on quadruplexes · Results/LLM comparison on quadruplex datasets (paragraph 1); Materials and methods/Data preparation (paragraph 1)
OutputsG4-associated sequence predictions
SourcesBenchmarking DNA large language models on quadruplexes · Introduction (paragraph 3); CRediT authorship contribution statement (paragraph 1)
Parameters117 million parameters, as identified for this row
SourcesBenchmarking DNA large language models on quadruplexes · Materials and methods/Low Rank Adaptation (LoRA) (paragraph 1); Discussion and conclusions (paragraph 3)
Known versions / configuration117M
SourcesBenchmarking DNA large language models on quadruplexes · Table tbl0030 (paragraph 1); Table tbl0025 (paragraph 1)
Training data / fittingTask-specific fine-tuning for approximately four to ten epochs, depending on model performance, with gradient accumulation.
SourcesBenchmarking DNA large language models on quadruplexes · Results/LLM comparison on quadruplex datasets (paragraph 5); Materials and methods/Fine-tuning (paragraph 1)
Context limitsA maximum input/context length for this exact evaluated configuration is not established by the inspected sources. · Not reported in inspected sources
Sources (2)Benchmarking DNA large language models on quadruplexes; MAGICS-LAB/DNABERT_2 README.md · Materials and methods/Data preparation; Materials and methods/Tokenization for G4s; Materials and methods/Metrics of evaluation; Materials and methods/Fine-tuning; Materials and methods/Low Rank Adaptation (LoRA); inspected for explicit maximum input length (dataset lengths and family-wide limits are not substituted); README.md at pinned repository revision
AccessOfficial upstream implementation and usage documentation: https://github.com/MAGICS-LAB/DNABERT_2/blob/f25bed9ee20db966dff39e5c1571249d04e36404/README.md. This pinned documentation revision is not automatically the evaluated weight revision.
SourcesMAGICS-LAB/DNABERT_2 README.md · README.md; installation, model download and usage instructions
Code licenceApache 2.0 (upstream repository code at the cited revision; this does not establish every dependency or historical checkpoint licence).
SourcesMAGICS-LAB/DNABERT_2 LICENSE · LICENSE; complete licence text
Weights licenceThe inspected model-access documentation does not explicitly identify terms for this exact evaluated checkpoint or fitted head; repository code terms are shown separately. · Not reported in inspected sources
SourcesMAGICS-LAB/DNABERT_2 README.md · README.md; checkpoint/access documentation and licence scope

Evidence table

Inspect claims, sources and review details

Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.

One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.

21 evidence rows matching the loaded filters

Claims, original sources and review scope · Release 2026-09-17-d277315f7d76
Property and statementOriginal source and locationReview and provenance
Diagram caption

Conceptual input–method–output guide. Check the procedure text and linked evaluation for fitted components, additional inputs and exact settings.

Individual claims
Benchmarking DNA large language models on quadruplexes

Original source ↗

Materials and methods/Data preparation (paragraph 2); Results/Interpretation of LLMs (paragraph 2)

Version: version of record
Retrieved: 2026-09-16T10:33:35.379Z

source checked

automated source review · 2026-09-16

Audit details

Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.

Field: attributes.profile.diagram.caption

Source artifact SHA-256: c3d7c6d068d3c11a9c8255a932197ece3804d94e8b2d4f858bea373a1b6eb32f

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Diagram steps

["DNA sequences for G-quadruplex classification and genomic scanning","DNABERT-2","G4-associated sequence predictions"]

Individual claims
Benchmarking DNA large language models on quadruplexes

Original source ↗

Materials and methods/Data preparation (paragraph 2); Results/Interpretation of LLMs (paragraph 2)

Version: version of record
Retrieved: 2026-09-16T10:33:35.379Z

source checked

automated source review · 2026-09-16

Audit details

Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.

Field: attributes.profile.diagram.steps

Source artifact SHA-256: c3d7c6d068d3c11a9c8255a932197ece3804d94e8b2d4f858bea373a1b6eb32f

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Diagram title

Evaluated procedure (conceptual)

Individual claims
Benchmarking DNA large language models on quadruplexes

Original source ↗

Materials and methods/Data preparation (paragraph 2); Results/Interpretation of LLMs (paragraph 2)

Version: version of record
Retrieved: 2026-09-16T10:33:35.379Z

source checked

automated source review · 2026-09-16

Audit details

Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.

Field: attributes.profile.diagram.title

Source artifact SHA-256: c3d7c6d068d3c11a9c8255a932197ece3804d94e8b2d4f858bea373a1b6eb32f

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Model type

DNA sequence transformer; this record is the paper-specific evaluated configuration.

Individual claims
MAGICS-LAB/DNABERT_2 README.md

Original source ↗

README.md model description

Version: f25bed9ee20db966dff39e5c1571249d04e36404
Retrieved: 2026-09-16T20:00:02.624261+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.

Field: attributes.profile.facts.0.value

Source artifact SHA-256: 734a8cec5f667d74d421bf3b273ad7e256216109636da45aa7ceba21cd34de16

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Architecture / procedure

The study fine-tunes genomic sequence models for G4 classification and applies them to whole-genome annotation, retaining each model’s native tokenisation.

Individual claims
Benchmarking DNA large language models on quadruplexes

Original source ↗

Materials and methods/Data preparation (paragraph 2); Results/Interpretation of LLMs (paragraph 2)

Version: version of record
Retrieved: 2026-09-16T10:33:35.379Z

source checked

automated source review · 2026-09-16

Audit details

Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.

Field: attributes.profile.facts.1.value

Source artifact SHA-256: c3d7c6d068d3c11a9c8255a932197ece3804d94e8b2d4f858bea373a1b6eb32f

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Weights licence

The inspected model-access documentation does not explicitly identify terms for this exact evaluated checkpoint or fitted head; repository code terms are shown separately.

Individual claims
MAGICS-LAB/DNABERT_2 README.md

Original source ↗

README.md; checkpoint/access documentation and licence scope

Version: f25bed9ee20db966dff39e5c1571249d04e36404
Retrieved: 2026-09-16T20:00:02.624261+00:00

unreported

automated source review · 2026-09-16

Audit details

Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.

Field: attributes.profile.facts.10.value

Source artifact SHA-256: 734a8cec5f667d74d421bf3b273ad7e256216109636da45aa7ceba21cd34de16

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Biological inputs

DNA sequences for G-quadruplex classification and genomic scanning

Individual claims
Benchmarking DNA large language models on quadruplexes

Original source ↗

Results/LLM comparison on quadruplex datasets (paragraph 1); Materials and methods/Data preparation (paragraph 1)

Version: version of record
Retrieved: 2026-09-16T10:33:35.379Z

source checked

automated source review · 2026-09-16

Audit details

Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.

Field: attributes.profile.facts.2.value

Source artifact SHA-256: c3d7c6d068d3c11a9c8255a932197ece3804d94e8b2d4f858bea373a1b6eb32f

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Outputs

G4-associated sequence predictions

Individual claims
Benchmarking DNA large language models on quadruplexes

Original source ↗

Introduction (paragraph 3); CRediT authorship contribution statement (paragraph 1)

Version: version of record
Retrieved: 2026-09-16T10:33:35.379Z

source checked

automated source review · 2026-09-16

Audit details

Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.

Field: attributes.profile.facts.3.value

Source artifact SHA-256: c3d7c6d068d3c11a9c8255a932197ece3804d94e8b2d4f858bea373a1b6eb32f

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Parameters

117 million parameters, as identified for this row

Individual claims
Benchmarking DNA large language models on quadruplexes

Original source ↗

Materials and methods/Low Rank Adaptation (LoRA) (paragraph 1); Discussion and conclusions (paragraph 3)

Version: version of record
Retrieved: 2026-09-16T10:33:35.379Z

source checked

automated source review · 2026-09-16

Audit details

Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.

Field: attributes.profile.facts.4.value

Source artifact SHA-256: c3d7c6d068d3c11a9c8255a932197ece3804d94e8b2d4f858bea373a1b6eb32f

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Known versions / configuration

117M

Individual claims
Benchmarking DNA large language models on quadruplexes

Original source ↗

Table tbl0030 (paragraph 1); Table tbl0025 (paragraph 1)

Version: version of record
Retrieved: 2026-09-16T10:33:35.379Z

source checked

automated source review · 2026-09-16

Audit details

Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.

Field: attributes.profile.facts.5.value

Source artifact SHA-256: c3d7c6d068d3c11a9c8255a932197ece3804d94e8b2d4f858bea373a1b6eb32f

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Sources and history

Release 2026-09-17-d277315f7d76 · Record review: needs review

3 source records and release historyDownload this release
Technical metadata and extraction receipts

Stable ID: reported-model-ade36035f58f27

areas
dna-genomes
entity level
method
version
117M
reported name
DNABERT-2
historical missing metadata
checkpoint revision: not_reported_in_legacy_extract; training data: not_reported_in_legacy_extract; licence: not_reported_in_legacy_extract
metadata review scope
historical_missing_metadata preserves the original discovery state. Current descriptive evidence and missingness are recorded in profile.facts; numerical-result review is separate.
legacy kinds
model
entity classification
review date: 2026-09-17; rationale: This source-scoped entry preserves the method/configuration actually named in an evaluation. It is neither a global family identity nor proof of an immutable checkpoint; the linked evaluation retains adaptation, fitting and scoring details.; source ids: quadruplex-llm-benchmark-2025; evidence-reported-base-dnabert2-readme-md; source locator: Materials and methods/Data preparation (paragraph 2); Results/Interpretation of LLMs (paragraph 2) | README.md model description | Discussion and conclusions (paragraph 6); Results/LLM performance at the genome-wide level (paragraph 4); ambiguities: Configuration means the source-labelled evaluated identity. It does not establish missing checkpoint hashes, default settings or equivalence to same-named records in other papers.
Related records

Suggest a correction