Model type
DNA sequence transformer; this record is the paper-specific evaluated configuration.
DNABERT-2 is fine-tuned to distinguish natural from GenomeOcean-generated DNA sequences.
Conceptual input–method–output guide. Check the procedure text and linked evaluation for fitted components, additional inputs and exact settings.
DNA sequence transformer; this record is the paper-specific evaluated configuration.
2-kbp natural or generated DNA sequences
Natural-versus-artificial sequence predictions
limited source coverage · Automated source review, 2026-09-16. All specifications and missing details
Release 2026-09-17-d277315f7d76 · 1 evaluation · 3 metric rows. Different protocols are not a single leaderboard.
| Metric and finding | Coverage and uncertainty | Evidence |
|---|---|---|
| DNABERT-2: Natural vs artificial microbial genome sequence Configuration: DNABERT-2Protocol: Natural versus GenomeOcean-generated DNA classification (Natural vs artificial microbial genome sequence)Dataset: GenomeOcean natural/artificial sequence test DNABERT 2 and NTv 2 standard fine-tuning; GenomeOcean LoRA. Negatives generated by GenomeOcean itself. CAMI 2 train 18,000/validation 2,000; GTDB test 20,000; balanced natural/artificial 2 kb sequences. Independent external evaluation · Evaluation metadata: needs review | ||
| 85.12 F1 Unit: % · Direction: higher | Uncertainty: not reported in legacy extract Scored: Not reported · Eligible: Not reported | source checkedGenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies; GenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies · Table 2, DNABERT-2 row, F1 column Source checking is not independent reproduction. |
| 85.02% Recall Unit: percent · Direction: higher | Uncertainty: unreported Scored: Not reported · Eligible: Not reported | source checkedGenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies · Table 2:, row DNABERT-2, column Recall; XML row2 column3 Source checking is not independent reproduction. |
| 85.23% Precision Unit: percent · Direction: higher | Uncertainty: unreported Scored: Not reported · Eligible: Not reported | source checkedGenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies · Table 2:, row DNABERT-2, column Precision; XML row2 column2 Source checking is not independent reproduction. |
The pretrained DNA transformer is adapted with a binary classifier using the study’s natural/artificial sequence labels.
DNABERT-2 replaces overlapping k-mer tokens with byte-pair encoding and uses ALiBi positional biases. The official 117M model produces 768-dimensional token representations; downstream classifiers and pooling choices are separate configuration details.
The linked evaluation record identifies DNABERT-2: Natural vs artificial microbial genome sequence. Its dataset, split, adaptation and evidence origin remain attached to the reported results.
Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.
Stable record: reported-model-7f1165b35f10e2Explanatory profile: limited source coverage · Automated source review, 2026-09-16. Review applies to the cited claims; unresolved fields are listed below. Numerical results retain their own review status.
| Property | Description and evidence |
|---|---|
| Model type | DNA sequence transformer; this record is the paper-specific evaluated configuration.SourcesMAGICS-LAB/DNABERT_2 README.md · README.md model description |
| Architecture / procedure | The pretrained DNA transformer is adapted with a binary classifier using the study’s natural/artificial sequence labels.SourcesGenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies · Results/Model Safety (paragraph 1); Results/Model Safety (paragraph 2) |
| Biological inputs | 2-kbp natural or generated DNA sequencesSourcesGenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies · Results/GenomeOcean Learns Protein-coding Principles from DNA Alone (paragraph 4); Methods/Training/Evaluation Datasets/Generated Sequence Discrimination (paragraph 1) |
| Outputs | Natural-versus-artificial sequence predictionsSourcesGenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies · Methods/Training/Evaluation Datasets/Generated Sequence Discrimination (paragraph 1); Results/Model Safety (paragraph 1) |
| Parameters | An aggregate parameter total for this exact evaluated configuration is not established by the inspected sources. · Not reported in inspected sourcesSources (2)GenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies; MAGICS-LAB/DNABERT_2 README.md · Results/Pre-training; Methods/Model and Training/Tokenization; Methods/Model and Training/Architecture and Pre-training; Methods/Model and Training/Fine-tune GenomeOcean as Biosynthetic Gene Clusters Foundation Model (bgcFM); Methods/Training/Evaluation Datasets/Assembled Metagenome Datasets; Methods/Training/Evaluation Datasets/Evaluation of Dataset Complexity; Methods/Training/Evaluation Datasets/Biosynthetic Gene Cluster (BGC) Collection; Methods/Training/Evaluation Datasets/ZymoBiomics Microbial Community Standard Datasets; inspected for aggregate parameter count (component sizes are not added without an exact configuration); README.md at pinned repository revision |
| Known versions / configuration | DNABERT-2 is the comparison-table label; that label does not specify an immutable weight revision. · Not reported in inspected sourcesSourcesGenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies · Model identification in the comparison table and corresponding Methods; immutable checkpoint revision is not supplied by the table label. |
| Training data / fitting | Balanced 18,000/2,000/20,000 train/validation/test sequences.SourcesGenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies · Methods/Training/Evaluation Datasets/Generated Sequence Discrimination (paragraph 1); Data & Code/Metagenome Raw Reads Datasets Used in Assembly (paragraph 1) |
| Context limits | 2,000-bp benchmark windows.SourcesGenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies · Abstract (paragraph 1); Table T4 (paragraph 1) |
| Access | Official upstream implementation and usage documentation: https://github.com/MAGICS-LAB/DNABERT_2/blob/f25bed9ee20db966dff39e5c1571249d04e36404/README.md. This pinned documentation revision is not automatically the evaluated weight revision.SourcesMAGICS-LAB/DNABERT_2 README.md · README.md; installation, model download and usage instructions |
| Code licence | Apache 2.0 (upstream repository code at the cited revision; this does not establish every dependency or historical checkpoint licence).SourcesMAGICS-LAB/DNABERT_2 LICENSE · LICENSE; complete licence text |
| Weights licence | The inspected model-access documentation does not explicitly identify terms for this exact evaluated checkpoint or fitted head; repository code terms are shown separately. · Not reported in inspected sourcesSourcesMAGICS-LAB/DNABERT_2 README.md · README.md; checkpoint/access documentation and licence scope |
Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.
One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.
21 evidence rows matching the loaded filters
| Property and statement | Original source and location | Review and provenance |
|---|---|---|
| Diagram caption Conceptual input–method–output guide. Check the procedure text and linked evaluation for fitted components, additional inputs and exact settings. Individual claims | GenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies Results/Model Safety (paragraph 1); Results/Model Safety (paragraph 2) Version: preprint archived 2025-02-05 | source checked automated source review · 2026-09-16 Audit detailsPrimary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Diagram steps ["2-kbp natural or generated DNA sequences","DNABERT-2","Natural-versus-artificial sequence predictions"] Individual claims | GenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies Results/Model Safety (paragraph 1); Results/Model Safety (paragraph 2) Version: preprint archived 2025-02-05 | source checked automated source review · 2026-09-16 Audit detailsPrimary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Diagram title Evaluated procedure (conceptual) Individual claims | GenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies Results/Model Safety (paragraph 1); Results/Model Safety (paragraph 2) Version: preprint archived 2025-02-05 | source checked automated source review · 2026-09-16 Audit detailsPrimary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Model type DNA sequence transformer; this record is the paper-specific evaluated configuration. Individual claims | MAGICS-LAB/DNABERT_2 README.md README.md model description Version: f25bed9ee20db966dff39e5c1571249d04e36404 | source checked automated source review · 2026-09-16 Audit detailsPrimary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Architecture / procedure The pretrained DNA transformer is adapted with a binary classifier using the study’s natural/artificial sequence labels. Individual claims | GenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies Results/Model Safety (paragraph 1); Results/Model Safety (paragraph 2) Version: preprint archived 2025-02-05 | source checked automated source review · 2026-09-16 Audit detailsPrimary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Weights licence The inspected model-access documentation does not explicitly identify terms for this exact evaluated checkpoint or fitted head; repository code terms are shown separately. Individual claims | MAGICS-LAB/DNABERT_2 README.md README.md; checkpoint/access documentation and licence scope Version: f25bed9ee20db966dff39e5c1571249d04e36404 | unreported automated source review · 2026-09-16 Audit detailsPrimary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Biological inputs 2-kbp natural or generated DNA sequences Individual claims | GenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies Results/GenomeOcean Learns Protein-coding Principles from DNA Alone (paragraph 4); Methods/Training/Evaluation Datasets/Generated Sequence Discrimination (paragraph 1) Version: preprint archived 2025-02-05 | source checked automated source review · 2026-09-16 Audit detailsPrimary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Outputs Natural-versus-artificial sequence predictions Individual claims | GenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies Methods/Training/Evaluation Datasets/Generated Sequence Discrimination (paragraph 1); Results/Model Safety (paragraph 1) Version: preprint archived 2025-02-05 | source checked automated source review · 2026-09-16 Audit detailsPrimary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Parameters An aggregate parameter total for this exact evaluated configuration is not established by the inspected sources. Individual claims | MAGICS-LAB/DNABERT_2 README.md Results/Pre-training; Methods/Model and Training/Tokenization; Methods/Model and Training/Architecture and Pre-training; Methods/Model and Training/Fine-tune GenomeOcean as Biosynthetic Gene Clusters Foundation Model (bgcFM); Methods/Training/Evaluation Datasets/Assembled Metagenome Datasets; Methods/Training/Evaluation Datasets/Evaluation of Dataset Complexity; Methods/Training/Evaluation Datasets/Biosynthetic Gene Cluster (BGC) Collection; Methods/Training/Evaluation Datasets/ZymoBiomics Microbial Community Standard Datasets; inspected for aggregate parameter count (component sizes are not added without an exact configuration); README.md at pinned repository revision Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: f25bed9ee20db966dff39e5c1571249d04e36404 | unreported automated source review · 2026-09-16 Audit detailsPrimary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Parameters An aggregate parameter total for this exact evaluated configuration is not established by the inspected sources. Individual claims | GenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies Results/Pre-training; Methods/Model and Training/Tokenization; Methods/Model and Training/Architecture and Pre-training; Methods/Model and Training/Fine-tune GenomeOcean as Biosynthetic Gene Clusters Foundation Model (bgcFM); Methods/Training/Evaluation Datasets/Assembled Metagenome Datasets; Methods/Training/Evaluation Datasets/Evaluation of Dataset Complexity; Methods/Training/Evaluation Datasets/Biosynthetic Gene Cluster (BGC) Collection; Methods/Training/Evaluation Datasets/ZymoBiomics Microbial Community Standard Datasets; inspected for aggregate parameter count (component sizes are not added without an exact configuration); README.md at pinned repository revision Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: preprint archived 2025-02-05 | unreported automated source review · 2026-09-16 Audit detailsPrimary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
Release 2026-09-17-d277315f7d76 · Record review: needs review
Stable ID: reported-model-7f1165b35f10e2