rewire.it
Task

Natural vs artificial microbial genome sequence

This GenomeOcean task distinguishes natural microbial sequence fragments from model-generated fragments. It measures discrimination under the study’s dataset construction.

SourcesGenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies · Methods 4.2.5 Generated Sequence Discrimination; Table 2

3 evaluations · 9 metric rows

At a glance

Inputs, training, access and other details

Explanatory profile: limited source coverage · Automated source review, 2026-09-16. Review applies to the cited claims; unresolved fields are listed below. Numerical results retain their own review status.

Data, procedure and scoring
PropertyDescription and evidence
Entity typePaper-specific evaluation task; this profile is a descriptive evidence summary.
SourcesGenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies · Methods 4.2.5 Generated Sequence Discrimination; Table 2
DatasetsNatural and generated sequence-fragment collections derived from CAMI2 and GTDB references.
SourcesGenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies · Methods 4.2.5 Generated Sequence Discrimination; Table 2
OrganismsMicrobial reference sequence collections; organism membership follows the named source datasets.
SourcesGenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies · Methods 4.2.5 Generated Sequence Discrimination; Table 2
AssaysComputational class labels for sequence origin, without a functional measurement in this particular endpoint.
SourcesGenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies · Methods 4.2.5 Generated Sequence Discrimination; Table 2
SplitsCAMI2-derived training/validation and GTDB-derived test collections are kept distinct. Exact membership and reference releases remain necessary for reproducibility.
SourcesGenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies · Methods 4.2.5 Generated Sequence Discrimination; Table 2
Allowed inputsDNA sequence representations and the natural/generated class label.
SourcesGenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies · Methods 4.2.5 Generated Sequence Discrimination; Table 2
AdaptationThe retained comparison assesses representations through the paper’s classification task; generation and classification are distinct stages.
SourcesGenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies · Methods 4.2.5 Generated Sequence Discrimination; Table 2
MetricsTable 2 reports classification measures including F1; each preserved result keeps its original metric and unit.
SourcesGenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies · Methods 4.2.5 Generated Sequence Discrimination; Table 2
BaselinesTable 2 includes GenomeOcean and DNABERT-2 representation-based comparators. Their classification results do not by themselves assess sequence function.
SourcesGenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies · Methods 4.2.5 Generated Sequence Discrimination; Table 2

How it works

How it worksConceptual assessment outline
Conceptual assessment outline1. Identify reference-derived classes. Then: 2. Separate fitting and test collections. Then: 3. Assess sequence discrimination. Then: 4. Report the specified metricConceptual assessment outline1. Identify reference-derived classes. Then: 2. Separate fitting and test collections. Then: 3. Assess sequence discrimination. Then: 4. Report the specified metricConceptual assessment outline1. Identify reference-derived classes. Then: 2. Separate fitting and test collections. Then: 3. Assess sequence discrimination. Then: 4. Report the specified metric

Conceptual overview of the published statistical assessment.

SourcesGenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies · Methods 4.2.5 Generated Sequence Discrimination; Table 2
What the evaluation establishes

The study uses CAMI2-derived examples for training and validation and GTDB-derived examples for testing. Natural and generated examples form separate classes. This source separation defines the assessment; a discrimination score is not experimental evidence that a generated sequence has biological function.

SourcesGenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies · Methods 4.2.5 Generated Sequence Discrimination; Table 2

Evaluation design

Benchmarks bring together tasks and protocols. A task describes the biological question; a protocol defines a particular test.

These source-backed links do not make different protocols or scores interchangeable.

Published comparisons

Explore the results reported under one evaluation protocol. Each figure keeps its source, dataset and metric together; it is not a ranking across studies.

Natural versus GenomeOcean-generated DNA classification · Table 2:

Precision (percent) · Higher values are better for this metric.

DNABERT 2 and NTv 2 standard fine-tuning; GenomeOcean LoRA. Negatives generated by GenomeOcean itself. CAMI 2 train 18,000/validation 2,000; GTDB test 20,000; balanced natural/artificial 2 kb sequences.

Evaluation protocol · GenomeOcean natural/artificial sequence test

  1. DNABERT-2 · Configuration · Independent external evaluation85.23
  2. GenomeOcean · Configuration · Author-reported evaluation99.03

Source order is preserved. Plotted marks show point estimates; uncertainty, where reported, is retained in the printed values and table. Differences do not establish statistical significance.

GenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies · Table 2:: Precision, Natural versus GenomeOcean-generated DNA classification
Values, uncertainty and evidence
Precision: original source values
Tested entityPrinted valueUncertaintyEvidence
DNABERT-2 · Configuration85.23 percentNot reportedIndependent external evaluation · source checkedGenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies · Table 2:, row DNABERT-2, column Precision; XML row2 column2
Nucleotide Transformers V2 · Configuration83.31 percentNot reportedAuthor-reported evaluation · source checkedGenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies · Table 2:, row Nucleotide Transformers V2, column Precision; XML row3 column2
GenomeOcean · Configuration99.03 percentNot reportedAuthor-reported evaluation · source checkedGenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies · Table 2:, row GenomeOcean, column Precision; XML row4 column2
Scope and limitations
  • Self-generator recognition is not generic biological sequence quality.
  • Counts are cohort sizes, not verified per-method successful-prediction coverage.
  • No interval assigned unless printed in source cell.

Source transcription and grouping reviewed by automated source review on 2026-09-17. These experiments were not independently reproduced by rewire.

Tested entities and results

Release 2026-09-17-d277315f7d76 · 3 evaluations · 9 metric rows. Different protocols are not a single leaderboard.

Results grouped by the exact reported evaluation
Metric and findingCoverage and uncertaintyEvidence
GenomeOcean: Natural vs artificial microbial genome sequence

DNABERT 2 and NTv 2 standard fine-tuning; GenomeOcean LoRA. Negatives generated by GenomeOcean itself. CAMI 2 train 18,000/validation 2,000; GTDB test 20,000; balanced natural/artificial 2 kb sequences.

Author-reported evaluation · Evaluation metadata: needs review

99.03 F1

Unit: % · Direction: higher

Uncertainty: not reported in legacy extract

Scored: Not reported · Eligible: Not reported

source checkedGenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies; GenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies · Table 2, GenomeOcean row, F1 column

Source checking is not independent reproduction.

99.03% Precision

Unit: percent · Direction: higher

Uncertainty: unreported

Scored: Not reported · Eligible: Not reported

source checkedGenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies · Table 2:, row GenomeOcean, column Precision; XML row4 column2

Source checking is not independent reproduction.

99.03% Recall

Unit: percent · Direction: higher

Uncertainty: unreported

Scored: Not reported · Eligible: Not reported

source checkedGenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies · Table 2:, row GenomeOcean, column Recall; XML row4 column3

Source checking is not independent reproduction.

DNABERT-2: Natural vs artificial microbial genome sequence

DNABERT 2 and NTv 2 standard fine-tuning; GenomeOcean LoRA. Negatives generated by GenomeOcean itself. CAMI 2 train 18,000/validation 2,000; GTDB test 20,000; balanced natural/artificial 2 kb sequences.

Independent external evaluation · Evaluation metadata: needs review

85.12 F1

Unit: % · Direction: higher

Uncertainty: not reported in legacy extract

Scored: Not reported · Eligible: Not reported

source checkedGenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies; GenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies · Table 2, DNABERT-2 row, F1 column

Source checking is not independent reproduction.

85.02% Recall

Unit: percent · Direction: higher

Uncertainty: unreported

Scored: Not reported · Eligible: Not reported

source checkedGenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies · Table 2:, row DNABERT-2, column Recall; XML row2 column3

Source checking is not independent reproduction.

85.23% Precision

Unit: percent · Direction: higher

Uncertainty: unreported

Scored: Not reported · Eligible: Not reported

source checkedGenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies · Table 2:, row DNABERT-2, column Precision; XML row2 column2

Source checking is not independent reproduction.

Nucleotide Transformers V2: Natural versus GenomeOcean-generated DNA classification

DNABERT 2 and NTv 2 standard fine-tuning; GenomeOcean LoRA. Negatives generated by GenomeOcean itself. CAMI 2 train 18,000/validation 2,000; GTDB test 20,000; balanced natural/artificial 2 kb sequences.

Author-reported evaluation · Evaluation metadata: needs review

83.14 F1

Unit: % · Direction: higher

Uncertainty: unreported

Scored: Not reported · Eligible: Not reported

source checkedGenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies · Table 2:, row Nucleotide Transformers V2, column F1; XML row3 column4

Source checking is not independent reproduction.

82.97% Recall

Unit: percent · Direction: higher

Uncertainty: unreported

Scored: Not reported · Eligible: Not reported

source checkedGenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies · Table 2:, row Nucleotide Transformers V2, column Recall; XML row3 column3

Source checking is not independent reproduction.

83.31% Precision

Unit: percent · Direction: higher

Uncertainty: unreported

Scored: Not reported · Eligible: Not reported

source checkedGenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies · Table 2:, row Nucleotide Transformers V2, column Precision; XML row3 column2

Source checking is not independent reproduction.

Papers and result coverage

Last literature check: 2026-09-17. Primary-source discovery and table/protocol screening; source checked is not independently reproduced. Raw acquisitions not automatically numerical publication approval.

Paper or primary resourceVersionReference
GenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assembliespreprint archived 2025-02-05Read source
DOI: 10.1101/2025.01.30.635558

What is still missing

  • Independent batch review before import; preserve existing observation identities.
Search and extraction details

complete comparison tables extracted pending publication review

Searches

  • GenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies primary paper benchmark results

Evidence locations

  • Table2; detecting artificial sequences Methods

Strengths and limitations

Profile review details

Inspected the cited primary abstract and descriptive computational-evaluation sections. Review covers the descriptive claims shown; no executable protocol was reconstructed. Original numerical records retain their prior transcription review.

Stable record: reported-task-9f9ab0090f6522

Evidence table

Inspect claims, sources and review details

Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.

One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.

16 evidence rows matching the loaded filters

Claims, original sources and review scope · Release 2026-09-17-d277315f7d76
Property and statementOriginal source and locationReview and provenance
Diagram caption

Conceptual overview of the published statistical assessment.

Individual claims
GenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies

Original source ↗

Methods 4.2.5 Generated Sequence Discrimination; Table 2

Version: preprint archived 2025-02-05
Retrieved: 2026-09-16T10:33:55.224Z

source checked

automated source review · 2026-09-16

Audit details

Inspected the cited primary abstract and descriptive computational-evaluation sections. Review covers the descriptive claims shown; no executable protocol was reconstructed. Original numerical records retain their prior transcription review.

Field: attributes.profile.diagram.caption

Source artifact SHA-256: 3cc0df52522fccda23e3958f069c916b87ee50bb5c9a992fa37e25256546e145

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Diagram steps

["Identify reference-derived classes","Separate fitting and test collections","Assess sequence discrimination","Report the specified metric"]

Individual claims
GenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies

Original source ↗

Methods 4.2.5 Generated Sequence Discrimination; Table 2

Version: preprint archived 2025-02-05
Retrieved: 2026-09-16T10:33:55.224Z

source checked

automated source review · 2026-09-16

Audit details

Inspected the cited primary abstract and descriptive computational-evaluation sections. Review covers the descriptive claims shown; no executable protocol was reconstructed. Original numerical records retain their prior transcription review.

Field: attributes.profile.diagram.steps

Source artifact SHA-256: 3cc0df52522fccda23e3958f069c916b87ee50bb5c9a992fa37e25256546e145

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Diagram title

Conceptual assessment outline

Individual claims
GenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies

Original source ↗

Methods 4.2.5 Generated Sequence Discrimination; Table 2

Version: preprint archived 2025-02-05
Retrieved: 2026-09-16T10:33:55.224Z

source checked

automated source review · 2026-09-16

Audit details

Inspected the cited primary abstract and descriptive computational-evaluation sections. Review covers the descriptive claims shown; no executable protocol was reconstructed. Original numerical records retain their prior transcription review.

Field: attributes.profile.diagram.title

Source artifact SHA-256: 3cc0df52522fccda23e3958f069c916b87ee50bb5c9a992fa37e25256546e145

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Entity type

Paper-specific evaluation task; this profile is a descriptive evidence summary.

Individual claims
GenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies

Original source ↗

Methods 4.2.5 Generated Sequence Discrimination; Table 2

Version: preprint archived 2025-02-05
Retrieved: 2026-09-16T10:33:55.224Z

source checked

automated source review · 2026-09-16

Audit details

Inspected the cited primary abstract and descriptive computational-evaluation sections. Review covers the descriptive claims shown; no executable protocol was reconstructed. Original numerical records retain their prior transcription review.

Field: attributes.profile.facts.0.value

Source artifact SHA-256: 3cc0df52522fccda23e3958f069c916b87ee50bb5c9a992fa37e25256546e145

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Datasets

Natural and generated sequence-fragment collections derived from CAMI2 and GTDB references.

Individual claims
GenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies

Original source ↗

Methods 4.2.5 Generated Sequence Discrimination; Table 2

Version: preprint archived 2025-02-05
Retrieved: 2026-09-16T10:33:55.224Z

source checked

automated source review · 2026-09-16

Audit details

Inspected the cited primary abstract and descriptive computational-evaluation sections. Review covers the descriptive claims shown; no executable protocol was reconstructed. Original numerical records retain their prior transcription review.

Field: attributes.profile.facts.1.value

Source artifact SHA-256: 3cc0df52522fccda23e3958f069c916b87ee50bb5c9a992fa37e25256546e145

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Organisms

Microbial reference sequence collections; organism membership follows the named source datasets.

Individual claims
GenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies

Original source ↗

Methods 4.2.5 Generated Sequence Discrimination; Table 2

Version: preprint archived 2025-02-05
Retrieved: 2026-09-16T10:33:55.224Z

source checked

automated source review · 2026-09-16

Audit details

Inspected the cited primary abstract and descriptive computational-evaluation sections. Review covers the descriptive claims shown; no executable protocol was reconstructed. Original numerical records retain their prior transcription review.

Field: attributes.profile.facts.2.value

Source artifact SHA-256: 3cc0df52522fccda23e3958f069c916b87ee50bb5c9a992fa37e25256546e145

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Assays

Computational class labels for sequence origin, without a functional measurement in this particular endpoint.

Individual claims
GenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies

Original source ↗

Methods 4.2.5 Generated Sequence Discrimination; Table 2

Version: preprint archived 2025-02-05
Retrieved: 2026-09-16T10:33:55.224Z

source checked

automated source review · 2026-09-16

Audit details

Inspected the cited primary abstract and descriptive computational-evaluation sections. Review covers the descriptive claims shown; no executable protocol was reconstructed. Original numerical records retain their prior transcription review.

Field: attributes.profile.facts.3.value

Source artifact SHA-256: 3cc0df52522fccda23e3958f069c916b87ee50bb5c9a992fa37e25256546e145

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Splits

CAMI2-derived training/validation and GTDB-derived test collections are kept distinct. Exact membership and reference releases remain necessary for reproducibility.

Individual claims
GenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies

Original source ↗

Methods 4.2.5 Generated Sequence Discrimination; Table 2

Version: preprint archived 2025-02-05
Retrieved: 2026-09-16T10:33:55.224Z

source checked

automated source review · 2026-09-16

Audit details

Inspected the cited primary abstract and descriptive computational-evaluation sections. Review covers the descriptive claims shown; no executable protocol was reconstructed. Original numerical records retain their prior transcription review.

Field: attributes.profile.facts.4.value

Source artifact SHA-256: 3cc0df52522fccda23e3958f069c916b87ee50bb5c9a992fa37e25256546e145

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Allowed inputs

DNA sequence representations and the natural/generated class label.

Individual claims
GenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies

Original source ↗

Methods 4.2.5 Generated Sequence Discrimination; Table 2

Version: preprint archived 2025-02-05
Retrieved: 2026-09-16T10:33:55.224Z

source checked

automated source review · 2026-09-16

Audit details

Inspected the cited primary abstract and descriptive computational-evaluation sections. Review covers the descriptive claims shown; no executable protocol was reconstructed. Original numerical records retain their prior transcription review.

Field: attributes.profile.facts.5.value

Source artifact SHA-256: 3cc0df52522fccda23e3958f069c916b87ee50bb5c9a992fa37e25256546e145

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Adaptation

The retained comparison assesses representations through the paper’s classification task; generation and classification are distinct stages.

Individual claims
GenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies

Original source ↗

Methods 4.2.5 Generated Sequence Discrimination; Table 2

Version: preprint archived 2025-02-05
Retrieved: 2026-09-16T10:33:55.224Z

source checked

automated source review · 2026-09-16

Audit details

Inspected the cited primary abstract and descriptive computational-evaluation sections. Review covers the descriptive claims shown; no executable protocol was reconstructed. Original numerical records retain their prior transcription review.

Field: attributes.profile.facts.6.value

Source artifact SHA-256: 3cc0df52522fccda23e3958f069c916b87ee50bb5c9a992fa37e25256546e145

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Sources and history

Release 2026-09-17-d277315f7d76 · Record review: needs review

2 source records and release historyDownload this release
Technical metadata and extraction receipts

Stable ID: reported-task-9f9ab0090f6522

areas
microbes-communities
tasks
Natural vs artificial microbial genome sequence
entity level
task
version
Not reported
task
Natural vs artificial microbial genome sequence
scope note
Paper-specific evaluation task; protocol completeness requires further extraction.
comparison panels
id: part2-genomeocean-2025-T2-d2237c539e; title: Natural versus GenomeOcean-generated DNA classification · Table 2:; protocol id: paper-protocol-a8c67b393443253db5; dataset id: reported-dataset-0c3ac7efe99c37; metric: Precision; unit: percent; direction: higher; result ids: paper-result-da2786542bcec5cf7f; paper-result-fae558ea3909a73b15; paper-result-af3d9eb456d1404344; source ids: part2-genomeocean-2025; source locator: Table 2:: Precision, Natural versus GenomeOcean-generated DNA classification; context: DNABERT 2 and NTv 2 standard fine-tuning; GenomeOcean LoRA. Negatives generated by GenomeOcean itself. CAMI 2 train 18,000/validation 2,000; GTDB test 20,000; balanced natural/artificial 2 kb sequences.; caveats: Self-generator recognition is not generic biological sequence quality.; Counts are cohort sizes, not verified per-method successful-prediction coverage.; No interval assigned unless printed in source cell.; review: method: automated_source_review; date: 2026-09-17; id: part2-genomeocean-2025-T2-3c62335b1d; title: Natural versus GenomeOcean-generated DNA classification · Table 2:; protocol id: paper-protocol-a8c67b393443253db5; dataset id: reported-dataset-0c3ac7efe99c37; metric: Recall; unit: percent; direction: higher; result ids: paper-result-415548dbb583463a98; paper-result-b3378778c23879d6db; paper-result-f356ae1c7c4122384a; source ids: part2-genomeocean-2025; source locator: Table 2:: Recall, Natural versus GenomeOcean-generated DNA classification; context: DNABERT 2 and NTv 2 standard fine-tuning; GenomeOcean LoRA. Negatives generated by GenomeOcean itself. CAMI 2 train 18,000/validation 2,000; GTDB test 20,000; balanced natural/artificial 2 kb sequences.; caveats: Self-generator recognition is not generic biological sequence quality.; Counts are cohort sizes, not verified per-method successful-prediction coverage.; No interval assigned unless printed in source cell.; review: method: automated_source_review; date: 2026-09-17; id: part2-genomeocean-2025-T2-92317bff43; title: Natural versus GenomeOcean-generated DNA classification · Table 2:; protocol id: paper-protocol-a8c67b393443253db5; dataset id: reported-dataset-0c3ac7efe99c37; metric: F1; unit: %; direction: higher; result ids: lit-b3-028; paper-result-0209741f5d0e66b495; lit-b3-027; source ids: part2-genomeocean-2025; source locator: Table 2:: F1, Natural versus GenomeOcean-generated DNA classification; context: DNABERT 2 and NTv 2 standard fine-tuning; GenomeOcean LoRA. Negatives generated by GenomeOcean itself. CAMI 2 train 18,000/validation 2,000; GTDB test 20,000; balanced natural/artificial 2 kb sequences.; caveats: Self-generator recognition is not generic biological sequence quality.; Counts are cohort sizes, not verified per-method successful-prediction coverage.; No interval assigned unless printed in source cell.; review: method: automated_source_review; date: 2026-09-17
benchmark research
review date: 2026-09-17; status: complete_comparison_tables_extracted_pending_publication_review; primary sources: part2-genomeocean-2025; inspected locators: Table2; detecting artificial sequences Methods; searched queries: GenomeOcean: An Efficient Genome Foundation Model Trained on Large-Scale Metagenomic Assemblies primary paper benchmark results; gaps: Independent batch review before import; preserve existing observation identities.; claim scope: Primary-source discovery and table/protocol screening; source checked is not independently reproduced. Raw acquisitions not automatically numerical publication approval.
historical missing metadata
protocol version: not_reported_in_legacy_extract; split: not_reported_in_legacy_extract
metadata review scope
historical_missing_metadata preserves the original discovery state. Current descriptive evidence and missingness are recorded in profile.facts; numerical-result review is separate.
legacy kinds
benchmark
entity classification
review date: 2026-09-17; rationale: This source-scoped record identifies the biological prediction task and holds its paper context. Preserve the existing task identity; exact split, model adaptation and scoring remain in linked evaluations or separate protocol records.; source ids: genomeocean-2025; source locator: Methods 4.2.5 Generated Sequence Discrimination; Table 2; ambiguities: A paper- or suite-specific task may constrain some inputs or metrics; that alone does not make it interchangeable with a complete versioned protocol. No protocol equivalence is inferred.; Some legacy profile Entity type facts use the generic phrase computational evaluation protocol. That boilerplate is not sufficient to establish a single fixed protocol identity or to merge this task with another protocol record.
Related records

Suggest a correction