rewire.it
Protocol

Natural versus GenomeOcean-generated DNA classification (Natural vs artificial microbial genome sequence)

Natural versus GenomeOcean-generated DNA classification · Table 2:. DNABERT 2 and NTv 2 standard fine-tuning; GenomeOcean LoRA. Negatives generated by GenomeOcean itself. CAMI 2 train 18,000/validation 2,000; GTDB test 20,000; balanced natural/artificial 2 kb sequences.

SourcesGenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies · Table 2:: Precision, Natural versus GenomeOcean-generated DNA classification

3 evaluations · 9 metric rows

At a glance

Inputs, training, access and other details

Explanatory profile: limited source coverage · Automated source review, 2026-09-17. Review applies to the cited claims; unresolved fields are listed below. Numerical results retain their own review status.

Data, procedure and scoring
PropertyDescription and evidence
DatasetsNot extracted or verified for this record.
OrganismsNot extracted or verified for this record.
AssaysNot extracted or verified for this record.
SplitsNot extracted or verified for this record.
Allowed inputsNot extracted or verified for this record.
AdaptationNot extracted or verified for this record.
MetricsNot extracted or verified for this record.
BaselinesNot extracted or verified for this record.

How it works

Evaluation in this paper

DNABERT 2 and NTv 2 standard fine-tuning; GenomeOcean LoRA. Negatives generated by GenomeOcean itself. CAMI 2 train 18,000/validation 2,000; GTDB test 20,000; balanced natural/artificial 2 kb sequences.

SourcesGenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies · Table 2:: Precision, Natural versus GenomeOcean-generated DNA classification

Evaluation design

Benchmarks bring together tasks and protocols. A task describes the biological question; a protocol defines a particular test.

These source-backed links do not make different protocols or scores interchangeable.

Recorded evaluations

Each evaluation records what was tested and under which conditions.

Published comparisons

Explore the results reported under one evaluation protocol. Each figure keeps its source, dataset and metric together; it is not a ranking across studies.

Natural versus GenomeOcean-generated DNA classification · Table 2:

Precision (percent) · Higher values are better for this metric.

DNABERT 2 and NTv 2 standard fine-tuning; GenomeOcean LoRA. Negatives generated by GenomeOcean itself. CAMI 2 train 18,000/validation 2,000; GTDB test 20,000; balanced natural/artificial 2 kb sequences.

Evaluation protocol · GenomeOcean natural/artificial sequence test

  1. DNABERT-2 · Configuration · Independent external evaluation85.23
  2. GenomeOcean · Configuration · Author-reported evaluation99.03

Source order is preserved. Plotted marks show point estimates; uncertainty, where reported, is retained in the printed values and table. Differences do not establish statistical significance.

GenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies · Table 2:: Precision, Natural versus GenomeOcean-generated DNA classification
Values, uncertainty and evidence
Precision: original source values
Tested entityPrinted valueUncertaintyEvidence
DNABERT-2 · Configuration85.23 percentNot reportedIndependent external evaluation · source checkedGenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies · Table 2:, row DNABERT-2, column Precision; XML row2 column2
Nucleotide Transformers V2 · Configuration83.31 percentNot reportedAuthor-reported evaluation · source checkedGenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies · Table 2:, row Nucleotide Transformers V2, column Precision; XML row3 column2
GenomeOcean · Configuration99.03 percentNot reportedAuthor-reported evaluation · source checkedGenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies · Table 2:, row GenomeOcean, column Precision; XML row4 column2
Scope and limitations
  • Self-generator recognition is not generic biological sequence quality.
  • Counts are cohort sizes, not verified per-method successful-prediction coverage.
  • No interval assigned unless printed in source cell.

Source transcription and grouping reviewed by automated source review on 2026-09-17. These experiments were not independently reproduced by rewire.

Tested entities and results

Release 2026-09-17-d277315f7d76 · 3 evaluations · 9 metric rows. Different protocols are not a single leaderboard.

Results grouped by the exact reported evaluation
Metric and findingCoverage and uncertaintyEvidence
GenomeOcean: Natural vs artificial microbial genome sequence

DNABERT 2 and NTv 2 standard fine-tuning; GenomeOcean LoRA. Negatives generated by GenomeOcean itself. CAMI 2 train 18,000/validation 2,000; GTDB test 20,000; balanced natural/artificial 2 kb sequences.

Author-reported evaluation · Evaluation metadata: needs review

99.03 F1

Unit: % · Direction: higher

Uncertainty: not reported in legacy extract

Scored: Not reported · Eligible: Not reported

source checkedGenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies; GenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies · Table 2, GenomeOcean row, F1 column

Source checking is not independent reproduction.

99.03% Precision

Unit: percent · Direction: higher

Uncertainty: unreported

Scored: Not reported · Eligible: Not reported

source checkedGenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies · Table 2:, row GenomeOcean, column Precision; XML row4 column2

Source checking is not independent reproduction.

99.03% Recall

Unit: percent · Direction: higher

Uncertainty: unreported

Scored: Not reported · Eligible: Not reported

source checkedGenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies · Table 2:, row GenomeOcean, column Recall; XML row4 column3

Source checking is not independent reproduction.

DNABERT-2: Natural vs artificial microbial genome sequence

DNABERT 2 and NTv 2 standard fine-tuning; GenomeOcean LoRA. Negatives generated by GenomeOcean itself. CAMI 2 train 18,000/validation 2,000; GTDB test 20,000; balanced natural/artificial 2 kb sequences.

Independent external evaluation · Evaluation metadata: needs review

85.12 F1

Unit: % · Direction: higher

Uncertainty: not reported in legacy extract

Scored: Not reported · Eligible: Not reported

source checkedGenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies; GenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies · Table 2, DNABERT-2 row, F1 column

Source checking is not independent reproduction.

85.02% Recall

Unit: percent · Direction: higher

Uncertainty: unreported

Scored: Not reported · Eligible: Not reported

source checkedGenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies · Table 2:, row DNABERT-2, column Recall; XML row2 column3

Source checking is not independent reproduction.

85.23% Precision

Unit: percent · Direction: higher

Uncertainty: unreported

Scored: Not reported · Eligible: Not reported

source checkedGenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies · Table 2:, row DNABERT-2, column Precision; XML row2 column2

Source checking is not independent reproduction.

Nucleotide Transformers V2: Natural versus GenomeOcean-generated DNA classification

DNABERT 2 and NTv 2 standard fine-tuning; GenomeOcean LoRA. Negatives generated by GenomeOcean itself. CAMI 2 train 18,000/validation 2,000; GTDB test 20,000; balanced natural/artificial 2 kb sequences.

Author-reported evaluation · Evaluation metadata: needs review

83.14 F1

Unit: % · Direction: higher

Uncertainty: unreported

Scored: Not reported · Eligible: Not reported

source checkedGenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies · Table 2:, row Nucleotide Transformers V2, column F1; XML row3 column4

Source checking is not independent reproduction.

82.97% Recall

Unit: percent · Direction: higher

Uncertainty: unreported

Scored: Not reported · Eligible: Not reported

source checkedGenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies · Table 2:, row Nucleotide Transformers V2, column Recall; XML row3 column3

Source checking is not independent reproduction.

83.31% Precision

Unit: percent · Direction: higher

Uncertainty: unreported

Scored: Not reported · Eligible: Not reported

source checkedGenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies · Table 2:, row Nucleotide Transformers V2, column Precision; XML row3 column2

Source checking is not independent reproduction.

Papers and result coverage

Last literature check: 2026-09-17. Primary-source discovery and table/protocol screening; source checked is not independently reproduced. Raw acquisitions not automatically numerical publication approval.

Paper or primary resourceVersionReference
GenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assembliespreprint archived 2025-02-05Read source
DOI: 10.1101/2025.01.30.635558

What is still missing

  • Independent batch review before import; preserve existing observation identities.
Search and extraction details

complete comparison tables extracted pending publication review

Searches

  • GenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies primary paper benchmark results

Evidence locations

  • Table 2:: Precision, Natural versus GenomeOcean-generated DNA classification

Strengths and limitations

Strengths and considerations

No source-reviewed explanatory claims are recorded here yet.

Limitations and conditions

No source-reviewed explanatory claims are recorded here yet.

Profile review details

Primary-source transcription and separate automated review. No human sign-off or experimental reproduction.

Stable record: paper-protocol-a8c67b393443253db5

Evidence table

Inspect claims, sources and review details

Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.

One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.

3 evidence rows matching the loaded filters

Claims, original sources and review scope · Release 2026-09-17-d277315f7d76
Property and statementOriginal source and locationReview and provenance
Evaluation in this paper

DNABERT 2 and NTv 2 standard fine-tuning; GenomeOcean LoRA. Negatives generated by GenomeOcean itself. CAMI 2 train 18,000/validation 2,000; GTDB test 20,000; balanced natural/artificial 2 kb sequences.

Individual claims
GenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies

Original source ↗

Table 2:: Precision, Natural versus GenomeOcean-generated DNA classification

Version: preprint archived 2025-02-05
Retrieved: 2026-09-17T07:56:18.823068+00:00

source checked

automated source review · 2026-09-17

Audit details

Primary-source transcription and separate automated review. No human sign-off or experimental reproduction.

Field: attributes.profile.sections.0.body

Source artifact SHA-256: 3cc0df52522fccda23e3958f069c916b87ee50bb5c9a992fa37e25256546e145

Hash scope: Exact retrieved primary paper artifact bytes.

Inspected artifact

Introduction

Natural versus GenomeOcean-generated DNA classification · Table 2:. DNABERT 2 and NTv 2 standard fine-tuning; GenomeOcean LoRA. Negatives generated by GenomeOcean itself. CAMI 2 train 18,000/validation 2,000; GTDB test 20,000; balanced natural/artificial 2 kb sequences.

Individual claims
GenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies

Original source ↗

Table 2:: Precision, Natural versus GenomeOcean-generated DNA classification

Version: preprint archived 2025-02-05
Retrieved: 2026-09-17T07:56:18.823068+00:00

source checked

automated source review · 2026-09-17

Audit details

Primary-source transcription and separate automated review. No human sign-off or experimental reproduction.

Field: attributes.profile.summary

Source artifact SHA-256: 3cc0df52522fccda23e3958f069c916b87ee50bb5c9a992fa37e25256546e145

Hash scope: Exact retrieved primary paper artifact bytes.

Inspected artifact

Relationship: evaluates task

reported-task-9f9ab0090f6522

Individual claims
GenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies

Original source ↗

Table 2:: Precision, Natural versus GenomeOcean-generated DNA classification

Version: preprint archived 2025-02-05
Retrieved: 2026-09-17T07:56:18.823068+00:00

source checked

automated source review · 2026-09-17

Audit details

Field: links:evaluates_task:reported-task-9f9ab0090f6522

Claim: paper-claim-734285a668bd8dea20

Source artifact SHA-256: 3cc0df52522fccda23e3958f069c916b87ee50bb5c9a992fa37e25256546e145

Hash scope: Exact retrieved primary paper artifact bytes.

Inspected artifact

Sources and history

Release 2026-09-17-d277315f7d76 · Record review: needs review

1 source records and release historyDownload this release
Technical metadata and extraction receipts

Stable ID: paper-protocol-a8c67b393443253db5

areas
microbes-communities
tasks
Natural vs artificial microbial genome sequence
entity level
protocol
protocol
DNABERT 2 and NTv 2 standard fine-tuning; GenomeOcean LoRA. Negatives generated by GenomeOcean itself. CAMI 2 train 18,000/validation 2,000; GTDB test 20,000; balanced natural/artificial 2 kb sequences.
comparison panels
id: part2-genomeocean-2025-T2-d2237c539e; title: Natural versus GenomeOcean-generated DNA classification · Table 2:; protocol id: paper-protocol-a8c67b393443253db5; dataset id: reported-dataset-0c3ac7efe99c37; metric: Precision; unit: percent; direction: higher; result ids: paper-result-da2786542bcec5cf7f; paper-result-fae558ea3909a73b15; paper-result-af3d9eb456d1404344; source ids: part2-genomeocean-2025; source locator: Table 2:: Precision, Natural versus GenomeOcean-generated DNA classification; context: DNABERT 2 and NTv 2 standard fine-tuning; GenomeOcean LoRA. Negatives generated by GenomeOcean itself. CAMI 2 train 18,000/validation 2,000; GTDB test 20,000; balanced natural/artificial 2 kb sequences.; caveats: Self-generator recognition is not generic biological sequence quality.; Counts are cohort sizes, not verified per-method successful-prediction coverage.; No interval assigned unless printed in source cell.; review: method: automated_source_review; date: 2026-09-17; id: part2-genomeocean-2025-T2-3c62335b1d; title: Natural versus GenomeOcean-generated DNA classification · Table 2:; protocol id: paper-protocol-a8c67b393443253db5; dataset id: reported-dataset-0c3ac7efe99c37; metric: Recall; unit: percent; direction: higher; result ids: paper-result-415548dbb583463a98; paper-result-b3378778c23879d6db; paper-result-f356ae1c7c4122384a; source ids: part2-genomeocean-2025; source locator: Table 2:: Recall, Natural versus GenomeOcean-generated DNA classification; context: DNABERT 2 and NTv 2 standard fine-tuning; GenomeOcean LoRA. Negatives generated by GenomeOcean itself. CAMI 2 train 18,000/validation 2,000; GTDB test 20,000; balanced natural/artificial 2 kb sequences.; caveats: Self-generator recognition is not generic biological sequence quality.; Counts are cohort sizes, not verified per-method successful-prediction coverage.; No interval assigned unless printed in source cell.; review: method: automated_source_review; date: 2026-09-17; id: part2-genomeocean-2025-T2-92317bff43; title: Natural versus GenomeOcean-generated DNA classification · Table 2:; protocol id: paper-protocol-a8c67b393443253db5; dataset id: reported-dataset-0c3ac7efe99c37; metric: F1; unit: %; direction: higher; result ids: lit-b3-028; paper-result-0209741f5d0e66b495; lit-b3-027; source ids: part2-genomeocean-2025; source locator: Table 2:: F1, Natural versus GenomeOcean-generated DNA classification; context: DNABERT 2 and NTv 2 standard fine-tuning; GenomeOcean LoRA. Negatives generated by GenomeOcean itself. CAMI 2 train 18,000/validation 2,000; GTDB test 20,000; balanced natural/artificial 2 kb sequences.; caveats: Self-generator recognition is not generic biological sequence quality.; Counts are cohort sizes, not verified per-method successful-prediction coverage.; No interval assigned unless printed in source cell.; review: method: automated_source_review; date: 2026-09-17
benchmark research
review date: 2026-09-17; status: complete_comparison_tables_extracted_pending_publication_review; primary sources: part2-genomeocean-2025; inspected locators: Table 2:: Precision, Natural versus GenomeOcean-generated DNA classification; searched queries: GenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies primary paper benchmark results; gaps: Independent batch review before import; preserve existing observation identities.; claim scope: Primary-source discovery and table/protocol screening; source checked is not independently reproduced. Raw acquisitions not automatically numerical publication approval.
legacy kinds
benchmark
entity classification
review date: 2026-09-17; rationale: The source-backed record identifies a specified evaluated procedure and its dataset/split/scoring context. Classify it as a protocol while preserving version and comparison restrictions.; source ids: part2-genomeocean-2025; source locator: Table 2:: Precision, Natural versus GenomeOcean-generated DNA classification; ambiguities: None recorded
Related records

Suggest a correction