Model type
DNA sequence transformer; this record is the paper-specific evaluated configuration.
This genomic model is adapted to predict G-quadruplex-associated sequence regions.
Conceptual input–method–output guide. Check the procedure text and linked evaluation for fitted components, additional inputs and exact settings.
DNA sequence transformer; this record is the paper-specific evaluated configuration.
DNA sequences for G-quadruplex classification and genomic scanning
G4-associated sequence predictions
limited source coverage · Automated source review, 2026-09-16. All specifications and missing details
Release 2026-09-17-d277315f7d76 · 1 evaluation · 1 metric row. Different protocols are not a single leaderboard.
| Metric and finding | Coverage and uncertainty | Evidence |
|---|---|---|
| DNABERT-2: G-quadruplex classification Pretrained model evaluated on KEx as reported in Table 5. Independent external evaluation · Evaluation metadata: needs review | ||
| 97.0 Accuracy Unit: % · Direction: unknown | Uncertainty: ± 0.5 Scored: Not reported · Eligible: Not reported | source checkedBenchmarking DNA large language models on quadruplexes · Table 5, DNABERT-2 (117 M) row, Accuracy column Source checking is not independent reproduction. |
The study fine-tunes genomic sequence models for G4 classification and applies them to whole-genome annotation, retaining each model’s native tokenisation.
DNABERT-2 replaces overlapping k-mer tokens with byte-pair encoding and uses ALiBi positional biases. The official 117M model produces 768-dimensional token representations; downstream classifiers and pooling choices are separate configuration details.
The linked evaluation record identifies DNABERT-2: G-quadruplex classification. Its dataset, split, adaptation and evidence origin remain attached to the reported results.
Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.
Stable record: reported-model-ade36035f58f27Explanatory profile: limited source coverage · Automated source review, 2026-09-16. Review applies to the cited claims; unresolved fields are listed below. Numerical results retain their own review status.
| Property | Description and evidence |
|---|---|
| Model type | DNA sequence transformer; this record is the paper-specific evaluated configuration.SourcesMAGICS-LAB/DNABERT_2 README.md · README.md model description |
| Architecture / procedure | The study fine-tunes genomic sequence models for G4 classification and applies them to whole-genome annotation, retaining each model’s native tokenisation.SourcesBenchmarking DNA large language models on quadruplexes · Materials and methods/Data preparation (paragraph 2); Results/Interpretation of LLMs (paragraph 2) |
| Biological inputs | DNA sequences for G-quadruplex classification and genomic scanningSourcesBenchmarking DNA large language models on quadruplexes · Results/LLM comparison on quadruplex datasets (paragraph 1); Materials and methods/Data preparation (paragraph 1) |
| Outputs | G4-associated sequence predictionsSourcesBenchmarking DNA large language models on quadruplexes · Introduction (paragraph 3); CRediT authorship contribution statement (paragraph 1) |
| Parameters | 117 million parameters, as identified for this rowSourcesBenchmarking DNA large language models on quadruplexes · Materials and methods/Low Rank Adaptation (LoRA) (paragraph 1); Discussion and conclusions (paragraph 3) |
| Known versions / configuration | 117MSourcesBenchmarking DNA large language models on quadruplexes · Table tbl0030 (paragraph 1); Table tbl0025 (paragraph 1) |
| Training data / fitting | Task-specific fine-tuning for approximately four to ten epochs, depending on model performance, with gradient accumulation.SourcesBenchmarking DNA large language models on quadruplexes · Results/LLM comparison on quadruplex datasets (paragraph 5); Materials and methods/Fine-tuning (paragraph 1) |
| Context limits | A maximum input/context length for this exact evaluated configuration is not established by the inspected sources. · Not reported in inspected sourcesSources (2)Benchmarking DNA large language models on quadruplexes; MAGICS-LAB/DNABERT_2 README.md · Materials and methods/Data preparation; Materials and methods/Tokenization for G4s; Materials and methods/Metrics of evaluation; Materials and methods/Fine-tuning; Materials and methods/Low Rank Adaptation (LoRA); inspected for explicit maximum input length (dataset lengths and family-wide limits are not substituted); README.md at pinned repository revision |
| Access | Official upstream implementation and usage documentation: https://github.com/MAGICS-LAB/DNABERT_2/blob/f25bed9ee20db966dff39e5c1571249d04e36404/README.md. This pinned documentation revision is not automatically the evaluated weight revision.SourcesMAGICS-LAB/DNABERT_2 README.md · README.md; installation, model download and usage instructions |
| Code licence | Apache 2.0 (upstream repository code at the cited revision; this does not establish every dependency or historical checkpoint licence).SourcesMAGICS-LAB/DNABERT_2 LICENSE · LICENSE; complete licence text |
| Weights licence | The inspected model-access documentation does not explicitly identify terms for this exact evaluated checkpoint or fitted head; repository code terms are shown separately. · Not reported in inspected sourcesSourcesMAGICS-LAB/DNABERT_2 README.md · README.md; checkpoint/access documentation and licence scope |
Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.
One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.
21 evidence rows matching the loaded filters
| Property and statement | Original source and location | Review and provenance |
|---|---|---|
| Diagram caption Conceptual input–method–output guide. Check the procedure text and linked evaluation for fitted components, additional inputs and exact settings. Individual claims | Benchmarking DNA large language models on quadruplexes Materials and methods/Data preparation (paragraph 2); Results/Interpretation of LLMs (paragraph 2) Version: version of record | source checked automated source review · 2026-09-16 Audit detailsPrimary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Diagram steps ["DNA sequences for G-quadruplex classification and genomic scanning","DNABERT-2","G4-associated sequence predictions"] Individual claims | Benchmarking DNA large language models on quadruplexes Materials and methods/Data preparation (paragraph 2); Results/Interpretation of LLMs (paragraph 2) Version: version of record | source checked automated source review · 2026-09-16 Audit detailsPrimary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Diagram title Evaluated procedure (conceptual) Individual claims | Benchmarking DNA large language models on quadruplexes Materials and methods/Data preparation (paragraph 2); Results/Interpretation of LLMs (paragraph 2) Version: version of record | source checked automated source review · 2026-09-16 Audit detailsPrimary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Model type DNA sequence transformer; this record is the paper-specific evaluated configuration. Individual claims | MAGICS-LAB/DNABERT_2 README.md README.md model description Version: f25bed9ee20db966dff39e5c1571249d04e36404 | source checked automated source review · 2026-09-16 Audit detailsPrimary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Architecture / procedure The study fine-tunes genomic sequence models for G4 classification and applies them to whole-genome annotation, retaining each model’s native tokenisation. Individual claims | Benchmarking DNA large language models on quadruplexes Materials and methods/Data preparation (paragraph 2); Results/Interpretation of LLMs (paragraph 2) Version: version of record | source checked automated source review · 2026-09-16 Audit detailsPrimary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Weights licence The inspected model-access documentation does not explicitly identify terms for this exact evaluated checkpoint or fitted head; repository code terms are shown separately. Individual claims | MAGICS-LAB/DNABERT_2 README.md README.md; checkpoint/access documentation and licence scope Version: f25bed9ee20db966dff39e5c1571249d04e36404 | unreported automated source review · 2026-09-16 Audit detailsPrimary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Biological inputs DNA sequences for G-quadruplex classification and genomic scanning Individual claims | Benchmarking DNA large language models on quadruplexes Results/LLM comparison on quadruplex datasets (paragraph 1); Materials and methods/Data preparation (paragraph 1) Version: version of record | source checked automated source review · 2026-09-16 Audit detailsPrimary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Outputs G4-associated sequence predictions Individual claims | Benchmarking DNA large language models on quadruplexes Introduction (paragraph 3); CRediT authorship contribution statement (paragraph 1) Version: version of record | source checked automated source review · 2026-09-16 Audit detailsPrimary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Parameters 117 million parameters, as identified for this row Individual claims | Benchmarking DNA large language models on quadruplexes Materials and methods/Low Rank Adaptation (LoRA) (paragraph 1); Discussion and conclusions (paragraph 3) Version: version of record | source checked automated source review · 2026-09-16 Audit detailsPrimary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Known versions / configuration 117M Individual claims | Benchmarking DNA large language models on quadruplexes Table tbl0030 (paragraph 1); Table tbl0025 (paragraph 1) Version: version of record | source checked automated source review · 2026-09-16 Audit detailsPrimary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
Release 2026-09-17-d277315f7d76 · Record review: needs review
Stable ID: reported-model-ade36035f58f27