Strengths and considerations
No source-reviewed explanatory claims are recorded here yet.
Natural versus GenomeOcean-generated DNA classification · Table 2:. DNABERT 2 and NTv 2 standard fine-tuning; GenomeOcean LoRA. Negatives generated by GenomeOcean itself. CAMI 2 train 18,000/validation 2,000; GTDB test 20,000; balanced natural/artificial 2 kb sequences.
GenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies · Table 2:: Precision, Natural versus GenomeOcean-generated DNA classificationExplanatory profile: limited source coverage · Automated source review, 2026-09-17. Review applies to the cited claims; unresolved fields are listed below. Numerical results retain their own review status.
DNABERT 2 and NTv 2 standard fine-tuning; GenomeOcean LoRA. Negatives generated by GenomeOcean itself. CAMI 2 train 18,000/validation 2,000; GTDB test 20,000; balanced natural/artificial 2 kb sequences.
GenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies · Table 2:: Precision, Natural versus GenomeOcean-generated DNA classificationBenchmarks bring together tasks and protocols. A task describes the biological question; a protocol defines a particular test.
These source-backed links do not make different protocols or scores interchangeable.
Each evaluation records what was tested and under which conditions.
Explore the results reported under one evaluation protocol. Each figure keeps its source, dataset and metric together; it is not a ranking across studies.
Precision (percent) · Higher values are better for this metric.
DNABERT 2 and NTv 2 standard fine-tuning; GenomeOcean LoRA. Negatives generated by GenomeOcean itself. CAMI 2 train 18,000/validation 2,000; GTDB test 20,000; balanced natural/artificial 2 kb sequences.
Evaluation protocol · GenomeOcean natural/artificial sequence test
Source order is preserved. Plotted marks show point estimates; uncertainty, where reported, is retained in the printed values and table. Differences do not establish statistical significance.
GenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies · Table 2:: Precision, Natural versus GenomeOcean-generated DNA classification| Tested entity | Printed value | Uncertainty | Evidence |
|---|---|---|---|
| DNABERT-2 · Configuration | 85.23 percent | Not reported | Independent external evaluation · source checkedGenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies · Table 2:, row DNABERT-2, column Precision; XML row2 column2 |
| Nucleotide Transformers V2 · Configuration | 83.31 percent | Not reported | Author-reported evaluation · source checkedGenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies · Table 2:, row Nucleotide Transformers V2, column Precision; XML row3 column2 |
| GenomeOcean · Configuration | 99.03 percent | Not reported | Author-reported evaluation · source checkedGenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies · Table 2:, row GenomeOcean, column Precision; XML row4 column2 |
Source transcription and grouping reviewed by automated source review on 2026-09-17. These experiments were not independently reproduced by rewire.
Release 2026-09-17-d277315f7d76 · 3 evaluations · 9 metric rows. Different protocols are not a single leaderboard.
| Metric and finding | Coverage and uncertainty | Evidence |
|---|---|---|
| GenomeOcean: Natural vs artificial microbial genome sequence Configuration: GenomeOceanProtocol: Natural versus GenomeOcean-generated DNA classification (Natural vs artificial microbial genome sequence)Dataset: GenomeOcean natural/artificial sequence test DNABERT 2 and NTv 2 standard fine-tuning; GenomeOcean LoRA. Negatives generated by GenomeOcean itself. CAMI 2 train 18,000/validation 2,000; GTDB test 20,000; balanced natural/artificial 2 kb sequences. Author-reported evaluation · Evaluation metadata: needs review | ||
| 99.03 F1 Unit: % · Direction: higher | Uncertainty: not reported in legacy extract Scored: Not reported · Eligible: Not reported | source checkedGenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies; GenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies · Table 2, GenomeOcean row, F1 column Source checking is not independent reproduction. |
| 99.03% Precision Unit: percent · Direction: higher | Uncertainty: unreported Scored: Not reported · Eligible: Not reported | source checkedGenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies · Table 2:, row GenomeOcean, column Precision; XML row4 column2 Source checking is not independent reproduction. |
| 99.03% Recall Unit: percent · Direction: higher | Uncertainty: unreported Scored: Not reported · Eligible: Not reported | source checkedGenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies · Table 2:, row GenomeOcean, column Recall; XML row4 column3 Source checking is not independent reproduction. |
| DNABERT-2: Natural vs artificial microbial genome sequence Configuration: DNABERT-2Protocol: Natural versus GenomeOcean-generated DNA classification (Natural vs artificial microbial genome sequence)Dataset: GenomeOcean natural/artificial sequence test DNABERT 2 and NTv 2 standard fine-tuning; GenomeOcean LoRA. Negatives generated by GenomeOcean itself. CAMI 2 train 18,000/validation 2,000; GTDB test 20,000; balanced natural/artificial 2 kb sequences. Independent external evaluation · Evaluation metadata: needs review | ||
| 85.12 F1 Unit: % · Direction: higher | Uncertainty: not reported in legacy extract Scored: Not reported · Eligible: Not reported | source checkedGenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies; GenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies · Table 2, DNABERT-2 row, F1 column Source checking is not independent reproduction. |
| 85.02% Recall Unit: percent · Direction: higher | Uncertainty: unreported Scored: Not reported · Eligible: Not reported | source checkedGenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies · Table 2:, row DNABERT-2, column Recall; XML row2 column3 Source checking is not independent reproduction. |
| 85.23% Precision Unit: percent · Direction: higher | Uncertainty: unreported Scored: Not reported · Eligible: Not reported | source checkedGenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies · Table 2:, row DNABERT-2, column Precision; XML row2 column2 Source checking is not independent reproduction. |
| Nucleotide Transformers V2: Natural versus GenomeOcean-generated DNA classification Configuration: Nucleotide Transformers V2Protocol: Natural versus GenomeOcean-generated DNA classification (Natural vs artificial microbial genome sequence)Dataset: GenomeOcean natural/artificial sequence test DNABERT 2 and NTv 2 standard fine-tuning; GenomeOcean LoRA. Negatives generated by GenomeOcean itself. CAMI 2 train 18,000/validation 2,000; GTDB test 20,000; balanced natural/artificial 2 kb sequences. Author-reported evaluation · Evaluation metadata: needs review | ||
| 83.14 F1 Unit: % · Direction: higher | Uncertainty: unreported Scored: Not reported · Eligible: Not reported | source checkedGenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies · Table 2:, row Nucleotide Transformers V2, column F1; XML row3 column4 Source checking is not independent reproduction. |
| 82.97% Recall Unit: percent · Direction: higher | Uncertainty: unreported Scored: Not reported · Eligible: Not reported | source checkedGenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies · Table 2:, row Nucleotide Transformers V2, column Recall; XML row3 column3 Source checking is not independent reproduction. |
| 83.31% Precision Unit: percent · Direction: higher | Uncertainty: unreported Scored: Not reported · Eligible: Not reported | source checkedGenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies · Table 2:, row Nucleotide Transformers V2, column Precision; XML row3 column2 Source checking is not independent reproduction. |
Last literature check: 2026-09-17. Primary-source discovery and table/protocol screening; source checked is not independently reproduced. Raw acquisitions not automatically numerical publication approval.
| Paper or primary resource | Version | Reference |
|---|---|---|
| GenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies | preprint archived 2025-02-05 | Read source DOI: 10.1101/2025.01.30.635558 |
complete comparison tables extracted pending publication review
No source-reviewed explanatory claims are recorded here yet.
No source-reviewed explanatory claims are recorded here yet.
Primary-source transcription and separate automated review. No human sign-off or experimental reproduction.
Stable record: paper-protocol-a8c67b393443253db5Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.
One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.
3 evidence rows matching the loaded filters
| Property and statement | Original source and location | Review and provenance |
|---|---|---|
| Evaluation in this paper DNABERT 2 and NTv 2 standard fine-tuning; GenomeOcean LoRA. Negatives generated by GenomeOcean itself. CAMI 2 train 18,000/validation 2,000; GTDB test 20,000; balanced natural/artificial 2 kb sequences. Individual claims | GenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies Table 2:: Precision, Natural versus GenomeOcean-generated DNA classification Version: preprint archived 2025-02-05 | source checked automated source review · 2026-09-17 Audit detailsPrimary-source transcription and separate automated review. No human sign-off or experimental reproduction. Field: Source artifact SHA-256: Hash scope: Exact retrieved primary paper artifact bytes. |
| Introduction Natural versus GenomeOcean-generated DNA classification · Table 2:. DNABERT 2 and NTv 2 standard fine-tuning; GenomeOcean LoRA. Negatives generated by GenomeOcean itself. CAMI 2 train 18,000/validation 2,000; GTDB test 20,000; balanced natural/artificial 2 kb sequences. Individual claims | GenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies Table 2:: Precision, Natural versus GenomeOcean-generated DNA classification Version: preprint archived 2025-02-05 | source checked automated source review · 2026-09-17 Audit detailsPrimary-source transcription and separate automated review. No human sign-off or experimental reproduction. Field: Source artifact SHA-256: Hash scope: Exact retrieved primary paper artifact bytes. |
| Relationship: evaluates task reported-task-9f9ab0090f6522 Individual claims | GenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies Table 2:: Precision, Natural versus GenomeOcean-generated DNA classification Version: preprint archived 2025-02-05 | source checked automated source review · 2026-09-17 Audit detailsField: Claim: paper-claim-734285a668bd8dea20 Source artifact SHA-256: Hash scope: Exact retrieved primary paper artifact bytes. |
Release 2026-09-17-d277315f7d76 · Record review: needs review
Stable ID: paper-protocol-a8c67b393443253db5