rewire.it
Configuration

DNABERT-2

DNABERT-2 is fine-tuned to distinguish natural from GenomeOcean-generated DNA sequences.

SourcesGenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies · Results/Model Safety (paragraph 1); Results/GenomeOcean Learns Protein-coding Principles from DNA Alone (paragraph 4)

1 evaluation · 3 metric rows

How it worksEvaluated procedure (conceptual)
Evaluated procedure (conceptual)1. 2-kbp natural or generated DNA sequences. Then: 2. DNABERT-2. Then: 3. Natural-versus-artificial sequence predictionsEvaluated procedure (conceptual)1. 2-kbp natural or generated DNA sequences. Then: 2. DNABERT-2. Then: 3. Natural-versus-artificial sequence predictionsEvaluated procedure (conceptual)1. 2-kbp natural or generated DNA sequences. Then: 2. DNABERT-2. Then: 3. Natural-versus-artificial sequence predictions

Conceptual input–method–output guide. Check the procedure text and linked evaluation for fitted components, additional inputs and exact settings.

SourcesGenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies · Results/Model Safety (paragraph 1); Results/Model Safety (paragraph 2)

At a glance

Model type

DNA sequence transformer; this record is the paper-specific evaluated configuration.

SourcesMAGICS-LAB/DNABERT_2 README.md · README.md model description

limited source coverage · Automated source review, 2026-09-16. All specifications and missing details

Evaluations and results

Release 2026-09-17-d277315f7d76 · 1 evaluation · 3 metric rows. Different protocols are not a single leaderboard.

Results grouped by the exact reported evaluation
Metric and findingCoverage and uncertaintyEvidence
DNABERT-2: Natural vs artificial microbial genome sequence

DNABERT 2 and NTv 2 standard fine-tuning; GenomeOcean LoRA. Negatives generated by GenomeOcean itself. CAMI 2 train 18,000/validation 2,000; GTDB test 20,000; balanced natural/artificial 2 kb sequences.

Independent external evaluation · Evaluation metadata: needs review

85.12 F1

Unit: % · Direction: higher

Uncertainty: not reported in legacy extract

Scored: Not reported · Eligible: Not reported

source checkedGenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies; GenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies · Table 2, DNABERT-2 row, F1 column

Source checking is not independent reproduction.

85.02% Recall

Unit: percent · Direction: higher

Uncertainty: unreported

Scored: Not reported · Eligible: Not reported

source checkedGenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies · Table 2:, row DNABERT-2, column Recall; XML row2 column3

Source checking is not independent reproduction.

85.23% Precision

Unit: percent · Direction: higher

Uncertainty: unreported

Scored: Not reported · Eligible: Not reported

source checkedGenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies · Table 2:, row DNABERT-2, column Precision; XML row2 column2

Source checking is not independent reproduction.

How it works

How the evaluated method works

The pretrained DNA transformer is adapted with a binary classifier using the study’s natural/artificial sequence labels.

SourcesGenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies · Results/Model Safety (paragraph 1); Results/Model Safety (paragraph 2)
Underlying method and version boundaries

DNABERT-2 replaces overlapping k-mer tokens with byte-pair encoding and uses ALiBi positional biases. The official 117M model produces 768-dimensional token representations; downstream classifiers and pooling choices are separate configuration details.

SourcesMAGICS-LAB/DNABERT_2 README.md · README.md; introduction, model description, pretrained-model and usage sections at pinned revision
What was evaluated

The linked evaluation record identifies DNABERT-2: Natural vs artificial microbial genome sequence. Its dataset, split, adaptation and evidence origin remain attached to the reported results.

SourcesGenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies · The named evaluation’s methods and comparison table; exact preserved evaluation IDs: evaluation-lit-b3-028

Strengths and limitations

Profile review details

Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.

Stable record: reported-model-7f1165b35f10e2

Specifications

Inputs, training, access and other details

Explanatory profile: limited source coverage · Automated source review, 2026-09-16. Review applies to the cited claims; unresolved fields are listed below. Numerical results retain their own review status.

Inputs, outputs and configuration
PropertyDescription and evidence
Model typeDNA sequence transformer; this record is the paper-specific evaluated configuration.
SourcesMAGICS-LAB/DNABERT_2 README.md · README.md model description
Architecture / procedureThe pretrained DNA transformer is adapted with a binary classifier using the study’s natural/artificial sequence labels.
SourcesGenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies · Results/Model Safety (paragraph 1); Results/Model Safety (paragraph 2)
Biological inputs2-kbp natural or generated DNA sequences
SourcesGenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies · Results/GenomeOcean Learns Protein-coding Principles from DNA Alone (paragraph 4); Methods/Training/Evaluation Datasets/Generated Sequence Discrimination (paragraph 1)
OutputsNatural-versus-artificial sequence predictions
SourcesGenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies · Methods/Training/Evaluation Datasets/Generated Sequence Discrimination (paragraph 1); Results/Model Safety (paragraph 1)
ParametersAn aggregate parameter total for this exact evaluated configuration is not established by the inspected sources. · Not reported in inspected sources
Sources (2)GenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies; MAGICS-LAB/DNABERT_2 README.md · Results/Pre-training; Methods/Model and Training/Tokenization; Methods/Model and Training/Architecture and Pre-training; Methods/Model and Training/Fine-tune GenomeOcean as Biosynthetic Gene Clusters Foundation Model (bgcFM); Methods/Training/Evaluation Datasets/Assembled Metagenome Datasets; Methods/Training/Evaluation Datasets/Evaluation of Dataset Complexity; Methods/Training/Evaluation Datasets/Biosynthetic Gene Cluster (BGC) Collection; Methods/Training/Evaluation Datasets/ZymoBiomics Microbial Community Standard Datasets; inspected for aggregate parameter count (component sizes are not added without an exact configuration); README.md at pinned repository revision
Known versions / configurationDNABERT-2 is the comparison-table label; that label does not specify an immutable weight revision. · Not reported in inspected sources
SourcesGenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies · Model identification in the comparison table and corresponding Methods; immutable checkpoint revision is not supplied by the table label.
Training data / fittingBalanced 18,000/2,000/20,000 train/validation/test sequences.
SourcesGenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies · Methods/Training/Evaluation Datasets/Generated Sequence Discrimination (paragraph 1); Data & Code/Metagenome Raw Reads Datasets Used in Assembly (paragraph 1)
Context limits2,000-bp benchmark windows.
SourcesGenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies · Abstract (paragraph 1); Table T4 (paragraph 1)
AccessOfficial upstream implementation and usage documentation: https://github.com/MAGICS-LAB/DNABERT_2/blob/f25bed9ee20db966dff39e5c1571249d04e36404/README.md. This pinned documentation revision is not automatically the evaluated weight revision.
SourcesMAGICS-LAB/DNABERT_2 README.md · README.md; installation, model download and usage instructions
Code licenceApache 2.0 (upstream repository code at the cited revision; this does not establish every dependency or historical checkpoint licence).
SourcesMAGICS-LAB/DNABERT_2 LICENSE · LICENSE; complete licence text
Weights licenceThe inspected model-access documentation does not explicitly identify terms for this exact evaluated checkpoint or fitted head; repository code terms are shown separately. · Not reported in inspected sources
SourcesMAGICS-LAB/DNABERT_2 README.md · README.md; checkpoint/access documentation and licence scope

Evidence table

Inspect claims, sources and review details

Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.

One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.

21 evidence rows matching the loaded filters

Claims, original sources and review scope · Release 2026-09-17-d277315f7d76
Property and statementOriginal source and locationReview and provenance
Diagram caption

Conceptual input–method–output guide. Check the procedure text and linked evaluation for fitted components, additional inputs and exact settings.

Individual claims
GenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies

Original source ↗

Results/Model Safety (paragraph 1); Results/Model Safety (paragraph 2)

Version: preprint archived 2025-02-05
Retrieved: 2026-09-16T10:33:55.224Z

source checked

automated source review · 2026-09-16

Audit details

Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.

Field: attributes.profile.diagram.caption

Source artifact SHA-256: 3cc0df52522fccda23e3958f069c916b87ee50bb5c9a992fa37e25256546e145

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Diagram steps

["2-kbp natural or generated DNA sequences","DNABERT-2","Natural-versus-artificial sequence predictions"]

Individual claims
GenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies

Original source ↗

Results/Model Safety (paragraph 1); Results/Model Safety (paragraph 2)

Version: preprint archived 2025-02-05
Retrieved: 2026-09-16T10:33:55.224Z

source checked

automated source review · 2026-09-16

Audit details

Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.

Field: attributes.profile.diagram.steps

Source artifact SHA-256: 3cc0df52522fccda23e3958f069c916b87ee50bb5c9a992fa37e25256546e145

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Diagram title

Evaluated procedure (conceptual)

Individual claims
GenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies

Original source ↗

Results/Model Safety (paragraph 1); Results/Model Safety (paragraph 2)

Version: preprint archived 2025-02-05
Retrieved: 2026-09-16T10:33:55.224Z

source checked

automated source review · 2026-09-16

Audit details

Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.

Field: attributes.profile.diagram.title

Source artifact SHA-256: 3cc0df52522fccda23e3958f069c916b87ee50bb5c9a992fa37e25256546e145

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Model type

DNA sequence transformer; this record is the paper-specific evaluated configuration.

Individual claims
MAGICS-LAB/DNABERT_2 README.md

Original source ↗

README.md model description

Version: f25bed9ee20db966dff39e5c1571249d04e36404
Retrieved: 2026-09-16T20:00:02.624261+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.

Field: attributes.profile.facts.0.value

Source artifact SHA-256: 734a8cec5f667d74d421bf3b273ad7e256216109636da45aa7ceba21cd34de16

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Architecture / procedure

The pretrained DNA transformer is adapted with a binary classifier using the study’s natural/artificial sequence labels.

Individual claims
GenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies

Original source ↗

Results/Model Safety (paragraph 1); Results/Model Safety (paragraph 2)

Version: preprint archived 2025-02-05
Retrieved: 2026-09-16T10:33:55.224Z

source checked

automated source review · 2026-09-16

Audit details

Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.

Field: attributes.profile.facts.1.value

Source artifact SHA-256: 3cc0df52522fccda23e3958f069c916b87ee50bb5c9a992fa37e25256546e145

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Weights licence

The inspected model-access documentation does not explicitly identify terms for this exact evaluated checkpoint or fitted head; repository code terms are shown separately.

Individual claims
MAGICS-LAB/DNABERT_2 README.md

Original source ↗

README.md; checkpoint/access documentation and licence scope

Version: f25bed9ee20db966dff39e5c1571249d04e36404
Retrieved: 2026-09-16T20:00:02.624261+00:00

unreported

automated source review · 2026-09-16

Audit details

Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.

Field: attributes.profile.facts.10.value

Source artifact SHA-256: 734a8cec5f667d74d421bf3b273ad7e256216109636da45aa7ceba21cd34de16

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Biological inputs

2-kbp natural or generated DNA sequences

Individual claims
GenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies

Original source ↗

Results/GenomeOcean Learns Protein-coding Principles from DNA Alone (paragraph 4); Methods/Training/Evaluation Datasets/Generated Sequence Discrimination (paragraph 1)

Version: preprint archived 2025-02-05
Retrieved: 2026-09-16T10:33:55.224Z

source checked

automated source review · 2026-09-16

Audit details

Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.

Field: attributes.profile.facts.2.value

Source artifact SHA-256: 3cc0df52522fccda23e3958f069c916b87ee50bb5c9a992fa37e25256546e145

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Outputs

Natural-versus-artificial sequence predictions

Individual claims
GenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies

Original source ↗

Methods/Training/Evaluation Datasets/Generated Sequence Discrimination (paragraph 1); Results/Model Safety (paragraph 1)

Version: preprint archived 2025-02-05
Retrieved: 2026-09-16T10:33:55.224Z

source checked

automated source review · 2026-09-16

Audit details

Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.

Field: attributes.profile.facts.3.value

Source artifact SHA-256: 3cc0df52522fccda23e3958f069c916b87ee50bb5c9a992fa37e25256546e145

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Parameters

An aggregate parameter total for this exact evaluated configuration is not established by the inspected sources.

Individual claims
MAGICS-LAB/DNABERT_2 README.md

Original source ↗

Results/Pre-training; Methods/Model and Training/Tokenization; Methods/Model and Training/Architecture and Pre-training; Methods/Model and Training/Fine-tune GenomeOcean as Biosynthetic Gene Clusters Foundation Model (bgcFM); Methods/Training/Evaluation Datasets/Assembled Metagenome Datasets; Methods/Training/Evaluation Datasets/Evaluation of Dataset Complexity; Methods/Training/Evaluation Datasets/Biosynthetic Gene Cluster (BGC) Collection; Methods/Training/Evaluation Datasets/ZymoBiomics Microbial Community Standard Datasets; inspected for aggregate parameter count (component sizes are not added without an exact configuration); README.md at pinned repository revision

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: f25bed9ee20db966dff39e5c1571249d04e36404
Retrieved: 2026-09-16T20:00:02.624261+00:00

unreported

automated source review · 2026-09-16

Audit details

Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.

Field: attributes.profile.facts.4.value

Source artifact SHA-256: 734a8cec5f667d74d421bf3b273ad7e256216109636da45aa7ceba21cd34de16

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Parameters

An aggregate parameter total for this exact evaluated configuration is not established by the inspected sources.

Individual claims
GenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies

Original source ↗

Results/Pre-training; Methods/Model and Training/Tokenization; Methods/Model and Training/Architecture and Pre-training; Methods/Model and Training/Fine-tune GenomeOcean as Biosynthetic Gene Clusters Foundation Model (bgcFM); Methods/Training/Evaluation Datasets/Assembled Metagenome Datasets; Methods/Training/Evaluation Datasets/Evaluation of Dataset Complexity; Methods/Training/Evaluation Datasets/Biosynthetic Gene Cluster (BGC) Collection; Methods/Training/Evaluation Datasets/ZymoBiomics Microbial Community Standard Datasets; inspected for aggregate parameter count (component sizes are not added without an exact configuration); README.md at pinned repository revision

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: preprint archived 2025-02-05
Retrieved: 2026-09-16T10:33:55.224Z

unreported

automated source review · 2026-09-16

Audit details

Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.

Field: attributes.profile.facts.4.value

Source artifact SHA-256: 3cc0df52522fccda23e3958f069c916b87ee50bb5c9a992fa37e25256546e145

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Sources and history

Release 2026-09-17-d277315f7d76 · Record review: needs review

3 source records and release historyDownload this release
Technical metadata and extraction receipts

Stable ID: reported-model-7f1165b35f10e2

areas
microbes-communities
entity level
method
version
Not reported
reported name
DNABERT-2
historical missing metadata
version: not_reported_in_legacy_extract; checkpoint revision: not_reported_in_legacy_extract; training data: not_reported_in_legacy_extract; licence: not_reported_in_legacy_extract
metadata review scope
historical_missing_metadata preserves the original discovery state. Current descriptive evidence and missingness are recorded in profile.facts; numerical-result review is separate.
legacy kinds
model
entity classification
review date: 2026-09-17; rationale: This source-scoped entry preserves the method/configuration actually named in an evaluation. It is neither a global family identity nor proof of an immutable checkpoint; the linked evaluation retains adaptation, fitting and scoring details.; source ids: genomeocean-2025; evidence-reported-base-dnabert2-readme-md; source locator: Results/Model Safety (paragraph 1); Results/Model Safety (paragraph 2) | README.md model description | Results/Model Safety (paragraph 1); Results/GenomeOcean Learns Protein-coding Principles from DNA Alone (paragraph 4); ambiguities: Configuration means the source-labelled evaluated identity. It does not establish missing checkpoint hashes, default settings or equivalence to same-named records in other papers.
Related records

Suggest a correction