Strengths and considerations
No source-reviewed explanatory claims are recorded here yet.
CAMI II read classification evaluates a computationally limited subsample, with taxonomic-rank-specific interpretation.
Explanatory profile: limited source coverage · Automated source review, 2026-09-16. Review applies to the cited claims; unresolved fields are listed below. Numerical results retain their own review status.
| Property | Description and evidence |
|---|---|
| Datasets | A 10,000-read subset of CAMI II Sample_0 from the human-microbiome collection.SourcesNormalized compression distance for DNA classification · Methods: Evaluation protocol; Results: CAMI II; cached text lines 67–68, 113–116 |
| Splits | Queries are classified against the paper’s separately assembled metagenomic reference/training data.SourcesNormalized compression distance for DNA classification · Methods: Evaluation protocol; Results: CAMI II; cached text lines 67–68, 113–116 |
| Metrics | Macro-F1 is the arithmetic mean of classwise F1; only classes present in the evaluated subset enter the macro calculation.SourcesNormalized compression distance for DNA classification · Methods: Evaluation protocol; Results: CAMI II; cached text lines 67–68, 113–116 |
| Baselines | NCD-gzip and Kraken2.SourcesNormalized compression distance for DNA classification · Methods: Evaluation protocol; Results: CAMI II; cached text lines 67–68, 113–116 |
| Leakage controls | The CAMI experiment classifies a single Sample_0 read subsample with the separately assembled metagenomic reference collection. The CAMI section does not report removal of CAMI source genomes or close relatives from that reference. Gene-out and taxon-out partitions elsewhere in the paper belong to different evaluations. · Not reported in inspected sourcesSourcesNormalized compression distance for DNA classification · Methods: metagenomic reference data; Results: CAMI dataset and Table 5; cached paragraphs 49–58,112–115 |
| Uncertainty | Table 5 reports point metrics for one 10,000-read CAMI subsample. The CAMI results section supplies no repeated-subsampling variability or confidence intervals; five-run cross-validation in Tables 3–4 concerns the separate human-gene classification experiment. · Not reported in inspected sourcesSourcesNormalized compression distance for DNA classification · Results: Human DNA classification, CAMI dataset; Tables 3–5 |
| Entity type | Paper-specific computational evaluation protocol.SourcesNormalized compression distance for DNA classification · Methods: Evaluation protocol; Results: CAMI II; cached text lines 67–68, 113–116 |
| Organisms | CAMI II human-microbiome community taxa.SourcesNormalized compression distance for DNA classification · Methods: Evaluation protocol; Results: CAMI II; cached text lines 67–68, 113–116 |
| Assays | Read-origin taxonomy at the selected evaluation rank.SourcesNormalized compression distance for DNA classification · Methods: Evaluation protocol; Results: CAMI II; cached text lines 67–68, 113–116 |
| Allowed inputs | DNA reads and a separately assembled reference/training collection.SourcesNormalized compression distance for DNA classification · Methods: Evaluation protocol; Results: CAMI II; cached text lines 67–68, 113–116 |
| Adaptation | Reference-based classification using NCD-gzip or Kraken2.SourcesNormalized compression distance for DNA classification · Methods: Evaluation protocol; Results: CAMI II; cached text lines 67–68, 113–116 |
Conceptual summary of the cited evaluation; exact task configuration and source version remain part of the protocol.
A 10,000-read subset of CAMI II Sample_0 from the human-microbiome collection. Queries are classified against the paper’s separately assembled metagenomic reference/training data. Macro-F1 is the arithmetic mean of classwise F1; only classes present in the evaluated subset enter the macro calculation. NCD-gzip and Kraken2.
Each evaluation records what was tested and under which conditions.
Release 2026-09-17-d277315f7d76 · 1 evaluation · 1 metric row. Different protocols are not a single leaderboard.
| Metric and finding | Coverage and uncertainty | Evidence |
|---|---|---|
| NCD-gzip: CAMI II superkingdom read classification Configuration: NCD-gzipTask: CAMI II superkingdom read classificationDataset: CAMI II Sample_0 10,000-read subsample Superkingdom-level macro-averaged F1; NCD assigns every read. Author-reported evaluation · Evaluation metadata: needs review | ||
| 0.9804 Macro F1 Unit: unitless · Direction: unknown | Uncertainty: not reported in legacy extract Scored: Not reported · Eligible: Not reported | source checkedNormalized compression distance for DNA classification · Table 5, NCD Superkingdom row, F1 column Source checking is not independent reproduction. |
Last literature check: 2026-09-17. Primary-paper discovery and source inspection. Source-checked results are not independently reproduced experiments.
| Paper or primary resource | Version | Reference |
|---|---|---|
| Normalized compression distance for DNA classification | version of record | Read source DOI: 10.7717/peerj.20677 |
primary comparison table screened
No source-reviewed explanatory claims are recorded here yet.
Relevant full-paper computational evaluation sections, tables/captions and cited supplementary task passages were reviewed. Reporting omissions are scoped to the inspected sources. Original numerical results are unchanged.
Stable record: reported-task-a2c37b8c420bc3Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.
One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.
17 evidence rows matching the loaded filters
| Property and statement | Original source and location | Review and provenance |
|---|---|---|
| Diagram caption Conceptual summary of the cited evaluation; exact task configuration and source version remain part of the protocol. Individual claims | Normalized compression distance for DNA classification Methods: Evaluation protocol; Results: CAMI II; cached text lines 67–68, 113–116 Version: version of record | source checked automated source review · 2026-09-16 Audit detailsRelevant full-paper computational evaluation sections, tables/captions and cited supplementary task passages were reviewed. Reporting omissions are scoped to the inspected sources. Original numerical results are unchanged. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Diagram steps ["Input: DNA reads and a separately assembled reference/training collection.","Evaluation: Reference-based classification using NCD-gzip or Kraken2.","Readout: Macro-F1 is the arithmetic mean of classwise F1; only classes present in the evaluated subset enter the macro calculation."] Individual claims | Normalized compression distance for DNA classification Methods: Evaluation protocol; Results: CAMI II; cached text lines 67–68, 113–116 Version: version of record | source checked automated source review · 2026-09-16 Audit detailsRelevant full-paper computational evaluation sections, tables/captions and cited supplementary task passages were reviewed. Reporting omissions are scoped to the inspected sources. Original numerical results are unchanged. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Diagram title Computational evaluation flow Individual claims | Normalized compression distance for DNA classification Methods: Evaluation protocol; Results: CAMI II; cached text lines 67–68, 113–116 Version: version of record | source checked automated source review · 2026-09-16 Audit detailsRelevant full-paper computational evaluation sections, tables/captions and cited supplementary task passages were reviewed. Reporting omissions are scoped to the inspected sources. Original numerical results are unchanged. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Datasets A 10,000-read subset of CAMI II Sample_0 from the human-microbiome collection. Individual claims | Normalized compression distance for DNA classification Methods: Evaluation protocol; Results: CAMI II; cached text lines 67–68, 113–116 Version: version of record | source checked automated source review · 2026-09-16 Audit detailsRelevant full-paper computational evaluation sections, tables/captions and cited supplementary task passages were reviewed. Reporting omissions are scoped to the inspected sources. Original numerical results are unchanged. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Splits Queries are classified against the paper’s separately assembled metagenomic reference/training data. Individual claims | Normalized compression distance for DNA classification Methods: Evaluation protocol; Results: CAMI II; cached text lines 67–68, 113–116 Version: version of record | source checked automated source review · 2026-09-16 Audit detailsRelevant full-paper computational evaluation sections, tables/captions and cited supplementary task passages were reviewed. Reporting omissions are scoped to the inspected sources. Original numerical results are unchanged. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Adaptation Reference-based classification using NCD-gzip or Kraken2. Individual claims | Normalized compression distance for DNA classification Methods: Evaluation protocol; Results: CAMI II; cached text lines 67–68, 113–116 Version: version of record | source checked automated source review · 2026-09-16 Audit detailsRelevant full-paper computational evaluation sections, tables/captions and cited supplementary task passages were reviewed. Reporting omissions are scoped to the inspected sources. Original numerical results are unchanged. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Metrics Macro-F1 is the arithmetic mean of classwise F1; only classes present in the evaluated subset enter the macro calculation. Individual claims | Normalized compression distance for DNA classification Methods: Evaluation protocol; Results: CAMI II; cached text lines 67–68, 113–116 Version: version of record | source checked automated source review · 2026-09-16 Audit detailsRelevant full-paper computational evaluation sections, tables/captions and cited supplementary task passages were reviewed. Reporting omissions are scoped to the inspected sources. Original numerical results are unchanged. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Baselines NCD-gzip and Kraken2. Individual claims | Normalized compression distance for DNA classification Methods: Evaluation protocol; Results: CAMI II; cached text lines 67–68, 113–116 Version: version of record | source checked automated source review · 2026-09-16 Audit detailsRelevant full-paper computational evaluation sections, tables/captions and cited supplementary task passages were reviewed. Reporting omissions are scoped to the inspected sources. Original numerical results are unchanged. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Leakage controls The CAMI experiment classifies a single Sample_0 read subsample with the separately assembled metagenomic reference collection. The CAMI section does not report removal of CAMI source genomes or close relatives from that reference. Gene-out and taxon-out partitions elsewhere in the paper belong to different evaluations. Individual claims | Normalized compression distance for DNA classification Methods: metagenomic reference data; Results: CAMI dataset and Table 5; cached paragraphs 49–58,112–115 Version: version of record | unreported automated source review · 2026-09-16 Audit detailsRelevant full-paper computational evaluation sections, tables/captions and cited supplementary task passages were reviewed. Reporting omissions are scoped to the inspected sources. Original numerical results are unchanged. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Uncertainty Table 5 reports point metrics for one 10,000-read CAMI subsample. The CAMI results section supplies no repeated-subsampling variability or confidence intervals; five-run cross-validation in Tables 3–4 concerns the separate human-gene classification experiment. Individual claims | Normalized compression distance for DNA classification Results: Human DNA classification, CAMI dataset; Tables 3–5 Version: version of record | unreported automated source review · 2026-09-16 Audit detailsRelevant full-paper computational evaluation sections, tables/captions and cited supplementary task passages were reviewed. Reporting omissions are scoped to the inspected sources. Original numerical results are unchanged. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
Release 2026-09-17-d277315f7d76 · Record review: needs review
Stable ID: reported-task-a2c37b8c420bc3