rewire.it
Task

Long-read taxonomic profiling

Long-read taxonomic profiling measures both taxon detection and relative-abundance agreement.

SourcesLightweight taxonomic profiling of long-read metagenomic datasets with Lemur and Magnet · Methods: Synthetic and simulated datasets; Method evaluation; cached text lines 75–77, 97–99; uncertainty/repeat-run/statistical-comparison passages

9 evaluations · 23 metric rows

At a glance

Inputs, training, access and other details

Explanatory profile: source reviewed · Automated source review, 2026-09-16. Review applies to the cited claims; unresolved fields are listed below. Numerical results retain their own review status.

Data, procedure and scoring
PropertyDescription and evidence
DatasetsSynthetic and simulated communities with known reference composition.
SourcesLightweight taxonomic profiling of long-read metagenomic datasets with Lemur and Magnet · Methods: Synthetic and simulated datasets; Method evaluation; cached text lines 75–77, 97–99; uncertainty/repeat-run/statistical-comparison passages
SplitsKnown-composition simulated or mock-community samples are profiled against reference resources. The Dilthey simulation has RefSeq species representatives for 94 of 96 strains; the additional simulation selects RefSeq-represented species.
SourcesLightweight taxonomic profiling of long-read metagenomic datasets with Lemur and Magnet · Methods: Synthetic and simulated datasets
MetricsGenus/species precision, recall and F1 after abundance thresholding; normalized L1 loss or Spearman correlation for abundance.
SourcesLightweight taxonomic profiling of long-read metagenomic datasets with Lemur and Magnet · Methods: Synthetic and simulated datasets; Method evaluation; cached text lines 75–77, 97–99; uncertainty/repeat-run/statistical-comparison passages
BaselinesLemur v1.0.1 is evaluated alone and with Magnet against Centrifuger v1.0.0, Kraken 2 v2.1.3, Melon v0.1.0, MetaMaps commit 633d2e0 and Sourmash v4.8.2. Melon lacks fungal references; the paper also reports bacterial-only comparisons for fungal-containing datasets.
SourcesLightweight taxonomic profiling of long-read metagenomic datasets with Lemur and Magnet · Methods: Method Comparison
Leakage controlsReference availability is part of this identification task: 94 of 96 strains in the first simulation have a corresponding RefSeq species representative. The additional metagenome simulation deliberately selects species with RefSeq representative genomes and available MAGs. These settings do not establish novel-species generalization.
SourcesLightweight taxonomic profiling of long-read metagenomic datasets with Lemur and Magnet · Methods: Simulated data from Dilthey et al. 2019; Simulated metagenome
UncertaintyMean and standard deviation across five replicate runs are reported for the Zymo EVEN and LOG comparisons; read subsampling also uses repeated seeds.
SourcesLightweight taxonomic profiling of long-read metagenomic datasets with Lemur and Magnet · Methods: Synthetic and simulated datasets; Method evaluation; cached text lines 75–77, 97–99; uncertainty/repeat-run/statistical-comparison passages
Entity typePaper-specific computational evaluation protocol.
SourcesLightweight taxonomic profiling of long-read metagenomic datasets with Lemur and Magnet · Methods: Synthetic and simulated datasets; Method evaluation; cached text lines 75–77, 97–99; uncertainty/repeat-run/statistical-comparison passages
OrganismsSynthetic and simulated microbial communities.
SourcesLightweight taxonomic profiling of long-read metagenomic datasets with Lemur and Magnet · Methods: Synthetic and simulated datasets; Method evaluation; cached text lines 75–77, 97–99; uncertainty/repeat-run/statistical-comparison passages
AssaysKnown taxonomic composition of sequence mixtures.
SourcesLightweight taxonomic profiling of long-read metagenomic datasets with Lemur and Magnet · Methods: Synthetic and simulated datasets; Method evaluation; cached text lines 75–77, 97–99; uncertainty/repeat-run/statistical-comparison passages
Allowed inputsLong sequencing reads.
SourcesLightweight taxonomic profiling of long-read metagenomic datasets with Lemur and Magnet · Methods: Synthetic and simulated datasets; Method evaluation; cached text lines 75–77, 97–99; uncertainty/repeat-run/statistical-comparison passages
AdaptationLemur and the comparator profilers use their reference resources; Magnet additionally aligns reads to cluster-representative genomes with minimap2. This is reference-based taxonomic profiling, rather than fitting a classifier on labeled train/test folds.
SourcesLightweight taxonomic profiling of long-read metagenomic datasets with Lemur and Magnet · Methods: Competitive read alignment with Magnet; Method Comparison

How it works

How it worksComputational evaluation flow
Computational evaluation flow1. Input: Long sequencing reads.. Then: 2. Evaluation: Known-composition simulated or mock-community samples are profiled against reference resources. The Dilthey simulation has RefSeq species representatives for 94 of 96 strains; the additional simulation selects RefSeq-represented species.. Then: 3. Readout: Genus/species precision, recall and F1 after abundance thresholding; normalized L1 loss or Spearman correlation for abundance.Computational evaluation flow1. Input: Long sequencing reads.. Then: 2. Evaluation: Known-composition simulated or mock-community samples are profiled against reference resources. The Dilthey simulation has RefSeq species representatives for 94 of 96 strains; the additional simulation selects RefSeq-represented species.. Then: 3. Readout: Genus/species precision, recall and F1 after abundance thresholding; normalized L1 loss or Spearman correlation for abundance.Computational evaluation flow1. Input: Long sequencing reads.. Then: 2. Evaluation: Known-composition simulated or mock-community samples are profiled against reference resources. The Dilthey simulation has RefSeq species representatives for 94 of 96 strains; the additional simulation selects RefSeq-represented species.. Then: 3. Readout: Genus/species precision, recall and F1 after abundance thresholding; normalized L1 loss or Spearman correlation for abundance.

Conceptual summary of the cited evaluation; exact task configuration and source version remain part of the protocol.

SourcesLightweight taxonomic profiling of long-read metagenomic datasets with Lemur and Magnet · Methods: Synthetic and simulated datasets; Method evaluation; cached text lines 75–77, 97–99; uncertainty/repeat-run/statistical-comparison passages; Methods: Synthetic and simulated datasets
Evaluation methodology

Synthetic and simulated communities with known reference composition. Known-composition simulated or mock-community samples are profiled against reference resources. The Dilthey simulation has RefSeq species representatives for 94 of 96 strains; the additional simulation selects RefSeq-represented species. Genus/species precision, recall and F1 after abundance thresholding; normalized L1 loss or Spearman correlation for abundance. Lemur v1.0.1 is evaluated alone and with Magnet against Centrifuger v1.0.0, Kraken 2 v2.1.3, Melon v0.1.0, MetaMaps commit 633d2e0 and Sourmash v4.8.2. Melon lacks fungal references; the paper also reports bacterial-only comparisons for fungal-containing datasets. Reference availability is part of this identification task: 94 of 96 strains in the first simulation have a corresponding RefSeq species representative. The additional metagenome simulation deliberately selects species with RefSeq representative genomes and available MAGs. These settings do not establish novel-species generalization.

SourcesLightweight taxonomic profiling of long-read metagenomic datasets with Lemur and Magnet · Methods: Synthetic and simulated datasets; Method evaluation; cached text lines 75–77, 97–99; uncertainty/repeat-run/statistical-comparison passages; Methods: Synthetic and simulated datasets; Methods: Method Comparison; Methods: Simulated data from Dilthey et al. 2019; Simulated metagenome

Evaluation design

Benchmarks bring together tasks and protocols. A task describes the biological question; a protocol defines a particular test.

These source-backed links do not make different protocols or scores interchangeable.

Recorded evaluations

Each evaluation records what was tested and under which conditions.

Published comparisons

Explore the results reported under one evaluation protocol. Each figure keeps its source, dataset and metric together; it is not a ranking across studies.

Species profiling on Dilthey2019 simulated long reads · Table 1:

Recall (fraction) · Higher values are better for this metric.

Species-level recall/precision/F 1. MetaMaps results copied from original manuscript; other tool versions reported in Methods. 200,114 reads;96 strain design,94 strains with RefSeq species representatives; average read 4,997 bp.

Evaluation protocol · Species profiling on Dilthey2019 simulated long reads

  1. Lemur · Configuration · Author-reported evaluation0.951
  2. Lemur + Magnet · Configuration · Author-reported evaluation0.927
  3. Melon · Configuration · Independent external evaluation0.963
  4. MetaMaps · Configuration · Result quoted from another source1.000
  5. Sourmash · Configuration · Independent external evaluation0.927
  6. Centrifuger · Configuration · Independent external evaluation0.774
  7. Kraken 2 · Configuration · Independent external evaluation0.976

Source order is preserved. Plotted marks show point estimates; uncertainty, where reported, is retained in the printed values and table. Differences do not establish statistical significance.

Lightweight taxonomic profiling of long-read metagenomic datasets with Lemur and Magnet · Table 1:: Recall, Species profiling on Dilthey2019 simulated long reads
Values, uncertainty and evidence
Recall: original source values
Tested entityPrinted valueUncertaintyEvidence
Lemur · Configuration0.951 fractionNot reportedAuthor-reported evaluation · source checkedLightweight taxonomic profiling of long-read metagenomic datasets with Lemur and Magnet · Table 1:, row Lemur, column Recall; XML row2 column2
Lemur + Magnet · Configuration0.927 fractionNot reportedAuthor-reported evaluation · source checkedLightweight taxonomic profiling of long-read metagenomic datasets with Lemur and Magnet · Table 1:, row Lemur + Magnet, column Recall; XML row3 column2
Melon · Configuration0.963 fractionNot reportedIndependent external evaluation · source checkedLightweight taxonomic profiling of long-read metagenomic datasets with Lemur and Magnet · Table 1:, row Melon, column Recall; XML row4 column2
MetaMaps · Configuration1.000 fractionNot reportedResult quoted from another source · source checkedLightweight taxonomic profiling of long-read metagenomic datasets with Lemur and Magnet · Table 1:, row MetaMaps, column Recall; XML row5 column2
Sourmash · Configuration0.927 fractionNot reportedIndependent external evaluation · source checkedLightweight taxonomic profiling of long-read metagenomic datasets with Lemur and Magnet · Table 1:, row Sourmash, column Recall; XML row6 column2
Centrifuger · Configuration0.774 fractionNot reportedIndependent external evaluation · source checkedLightweight taxonomic profiling of long-read metagenomic datasets with Lemur and Magnet · Table 1:, row Centrifuger, column Recall; XML row7 column2
Kraken 2 · Configuration0.976 fractionNot reportedIndependent external evaluation · source checkedLightweight taxonomic profiling of long-read metagenomic datasets with Lemur and Magnet · Table 1:, row Kraken 2, column Recall; XML row8 column2
Scope and limitations
  • MetaMaps is quoted prior evidence, not an independent repeated run.
  • Reference databases and classifier-versus-profiler task differ.
  • No interval assigned unless printed in source cell.

Source transcription and grouping reviewed by automated source review on 2026-09-17. These experiments were not independently reproduced by rewire.

Tested entities and results

Release 2026-09-17-d277315f7d76 · 9 evaluations · 23 metric rows. Different protocols are not a single leaderboard.

Results grouped by the exact reported evaluation
Metric and findingCoverage and uncertaintyEvidence
Lemur: Long-read taxonomic profiling

Mean across five replicate runs on Zymo LOG 10%.

Author-reported evaluation · Evaluation metadata: needs review

0.376 F1

Unit: unitless · Direction: unknown

Uncertainty: not reported in legacy extract

Scored: Not reported · Eligible: Not reported

source checkedLightweight taxonomic profiling of long-read metagenomic datasets with Lemur and Magnet · Table 3, LOG 10% / Lemur row, F1 column

Source checking is not independent reproduction.

Kraken 2: Long-read taxonomic profiling

Mean across five replicate runs on Zymo LOG 10%.

Independent external evaluation · Evaluation metadata: needs review

0.375 F1

Unit: unitless · Direction: unknown

Uncertainty: not reported in legacy extract

Scored: Not reported · Eligible: Not reported

source checkedLightweight taxonomic profiling of long-read metagenomic datasets with Lemur and Magnet · Table 3, LOG 10% / Kraken 2 row, F1 column

Source checking is not independent reproduction.

Sourmash: Species profiling on Dilthey2019 simulated long reads

Species-level recall/precision/F 1. MetaMaps results copied from original manuscript; other tool versions reported in Methods. 200,114 reads;96 strain design,94 strains with RefSeq species representatives; average read 4,997 bp.

Independent external evaluation · Evaluation metadata: needs review

0.932 F1 score

Unit: fraction · Direction: higher

Uncertainty: unreported

Scored: Not reported · Eligible: Not reported

source checkedLightweight taxonomic profiling of long-read metagenomic datasets with Lemur and Magnet · Table 1:, row Sourmash, column F1 score; XML row6 column4

Source checking is not independent reproduction.

0.927 Recall

Unit: fraction · Direction: higher

Uncertainty: unreported

Scored: Not reported · Eligible: Not reported

source checkedLightweight taxonomic profiling of long-read metagenomic datasets with Lemur and Magnet · Table 1:, row Sourmash, column Recall; XML row6 column2

Source checking is not independent reproduction.

0.938 Precision

Unit: fraction · Direction: higher

Uncertainty: unreported

Scored: Not reported · Eligible: Not reported

source checkedLightweight taxonomic profiling of long-read metagenomic datasets with Lemur and Magnet · Table 1:, row Sourmash, column Precision; XML row6 column3

Source checking is not independent reproduction.

Lemur + Magnet: Species profiling on Dilthey2019 simulated long reads

Species-level recall/precision/F 1. MetaMaps results copied from original manuscript; other tool versions reported in Methods. 200,114 reads;96 strain design,94 strains with RefSeq species representatives; average read 4,997 bp.

Author-reported evaluation · Evaluation metadata: needs review

0.950 Precision

Unit: fraction · Direction: higher

Uncertainty: unreported

Scored: Not reported · Eligible: Not reported

source checkedLightweight taxonomic profiling of long-read metagenomic datasets with Lemur and Magnet · Table 1:, row Lemur + Magnet, column Precision; XML row3 column3

Source checking is not independent reproduction.

0.938 F1 score

Unit: fraction · Direction: higher

Uncertainty: unreported

Scored: Not reported · Eligible: Not reported

source checkedLightweight taxonomic profiling of long-read metagenomic datasets with Lemur and Magnet · Table 1:, row Lemur + Magnet, column F1 score; XML row3 column4

Source checking is not independent reproduction.

0.927 Recall

Unit: fraction · Direction: higher

Uncertainty: unreported

Scored: Not reported · Eligible: Not reported

source checkedLightweight taxonomic profiling of long-read metagenomic datasets with Lemur and Magnet · Table 1:, row Lemur + Magnet, column Recall; XML row3 column2

Source checking is not independent reproduction.

Kraken 2: Species profiling on Dilthey2019 simulated long reads

Species-level recall/precision/F 1. MetaMaps results copied from original manuscript; other tool versions reported in Methods. 200,114 reads;96 strain design,94 strains with RefSeq species representatives; average read 4,997 bp.

Independent external evaluation · Evaluation metadata: needs review

0.055 Precision

Unit: fraction · Direction: higher

Uncertainty: unreported

Scored: Not reported · Eligible: Not reported

source checkedLightweight taxonomic profiling of long-read metagenomic datasets with Lemur and Magnet · Table 1:, row Kraken 2, column Precision; XML row8 column3

Source checking is not independent reproduction.

0.976 Recall

Unit: fraction · Direction: higher

Uncertainty: unreported

Scored: Not reported · Eligible: Not reported

source checkedLightweight taxonomic profiling of long-read metagenomic datasets with Lemur and Magnet · Table 1:, row Kraken 2, column Recall; XML row8 column2

Source checking is not independent reproduction.

0.104 F1 score

Unit: fraction · Direction: higher

Uncertainty: unreported

Scored: Not reported · Eligible: Not reported

source checkedLightweight taxonomic profiling of long-read metagenomic datasets with Lemur and Magnet · Table 1:, row Kraken 2, column F1 score; XML row8 column4

Source checking is not independent reproduction.

Lemur: Species profiling on Dilthey2019 simulated long reads

Species-level recall/precision/F 1. MetaMaps results copied from original manuscript; other tool versions reported in Methods. 200,114 reads;96 strain design,94 strains with RefSeq species representatives; average read 4,997 bp.

Author-reported evaluation · Evaluation metadata: needs review

0.703 Precision

Unit: fraction · Direction: higher

Uncertainty: unreported

Scored: Not reported · Eligible: Not reported

source checkedLightweight taxonomic profiling of long-read metagenomic datasets with Lemur and Magnet · Table 1:, row Lemur, column Precision; XML row2 column3

Source checking is not independent reproduction.

0.808 F1 score

Unit: fraction · Direction: higher

Uncertainty: unreported

Scored: Not reported · Eligible: Not reported

source checkedLightweight taxonomic profiling of long-read metagenomic datasets with Lemur and Magnet · Table 1:, row Lemur, column F1 score; XML row2 column4

Source checking is not independent reproduction.

0.951 Recall

Unit: fraction · Direction: higher

Uncertainty: unreported

Scored: Not reported · Eligible: Not reported

source checkedLightweight taxonomic profiling of long-read metagenomic datasets with Lemur and Magnet · Table 1:, row Lemur, column Recall; XML row2 column2

Source checking is not independent reproduction.

Melon: Species profiling on Dilthey2019 simulated long reads

Species-level recall/precision/F 1. MetaMaps results copied from original manuscript; other tool versions reported in Methods. 200,114 reads;96 strain design,94 strains with RefSeq species representatives; average read 4,997 bp.

Independent external evaluation · Evaluation metadata: needs review

0.946 F1 score

Unit: fraction · Direction: higher

Uncertainty: unreported

Scored: Not reported · Eligible: Not reported

source checkedLightweight taxonomic profiling of long-read metagenomic datasets with Lemur and Magnet · Table 1:, row Melon, column F1 score; XML row4 column4

Source checking is not independent reproduction.

0.929 Precision

Unit: fraction · Direction: higher

Uncertainty: unreported

Scored: Not reported · Eligible: Not reported

source checkedLightweight taxonomic profiling of long-read metagenomic datasets with Lemur and Magnet · Table 1:, row Melon, column Precision; XML row4 column3

Source checking is not independent reproduction.

0.963 Recall

Unit: fraction · Direction: higher

Uncertainty: unreported

Scored: Not reported · Eligible: Not reported

source checkedLightweight taxonomic profiling of long-read metagenomic datasets with Lemur and Magnet · Table 1:, row Melon, column Recall; XML row4 column2

Source checking is not independent reproduction.

Centrifuger: Species profiling on Dilthey2019 simulated long reads

Species-level recall/precision/F 1. MetaMaps results copied from original manuscript; other tool versions reported in Methods. 200,114 reads;96 strain design,94 strains with RefSeq species representatives; average read 4,997 bp.

Independent external evaluation · Evaluation metadata: needs review

0.774 Recall

Unit: fraction · Direction: higher

Uncertainty: unreported

Scored: Not reported · Eligible: Not reported

source checkedLightweight taxonomic profiling of long-read metagenomic datasets with Lemur and Magnet · Table 1:, row Centrifuger, column Recall; XML row7 column2

Source checking is not independent reproduction.

0.050 Precision

Unit: fraction · Direction: higher

Uncertainty: unreported

Scored: Not reported · Eligible: Not reported

source checkedLightweight taxonomic profiling of long-read metagenomic datasets with Lemur and Magnet · Table 1:, row Centrifuger, column Precision; XML row7 column3

Source checking is not independent reproduction.

0.093 F1 score

Unit: fraction · Direction: higher

Uncertainty: unreported

Scored: Not reported · Eligible: Not reported

source checkedLightweight taxonomic profiling of long-read metagenomic datasets with Lemur and Magnet · Table 1:, row Centrifuger, column F1 score; XML row7 column4

Source checking is not independent reproduction.

MetaMaps: Species profiling on Dilthey2019 simulated long reads

Species-level recall/precision/F 1. MetaMaps results copied from original manuscript; other tool versions reported in Methods. 200,114 reads;96 strain design,94 strains with RefSeq species representatives; average read 4,997 bp.

Result quoted from another source · Evaluation metadata: needs review

0.862 Precision

Unit: fraction · Direction: higher

Uncertainty: unreported

Scored: Not reported · Eligible: Not reported

source checkedLightweight taxonomic profiling of long-read metagenomic datasets with Lemur and Magnet · Table 1:, row MetaMaps, column Precision; XML row5 column3

Source checking is not independent reproduction.

1.000 Recall

Unit: fraction · Direction: higher

Uncertainty: unreported

Scored: Not reported · Eligible: Not reported

source checkedLightweight taxonomic profiling of long-read metagenomic datasets with Lemur and Magnet · Table 1:, row MetaMaps, column Recall; XML row5 column2

Source checking is not independent reproduction.

0.926 F1 score

Unit: fraction · Direction: higher

Uncertainty: unreported

Scored: Not reported · Eligible: Not reported

source checkedLightweight taxonomic profiling of long-read metagenomic datasets with Lemur and Magnet · Table 1:, row MetaMaps, column F1 score; XML row5 column4

Source checking is not independent reproduction.

Papers and result coverage

Last literature check: 2026-09-17. Primary-source discovery and table/protocol screening; source checked is not independently reproduced. Raw acquisitions not automatically numerical publication approval.

Paper or primary resourceVersionReference
Lightweight taxonomic profiling of long-read metagenomic datasets with Lemur and MagnetPMC archival version PMC11185576.2Read source
DOI: 10.1101/2024.06.01.596961

What is still missing

  • Independent batch review before import; preserve existing observation identities.
Search and extraction details

complete comparison tables extracted pending publication review

Searches

  • Lightweight taxonomic profiling of long-read metagenomic datasets with Lemur and Magnet primary paper benchmark results

Evidence locations

  • Tables1–3; simulated Dilthey2019 and ZymoEVEN/LOG cohorts

Strengths and limitations

Strengths supported by sources

No source-reviewed explanatory claims are recorded here yet.

Limitations and conditions

  • Spearman is computed only over taxa present in both truth and prediction, so it omits false-positive and false-negative taxa. Presence/absence thresholds vary for the Zymo EVEN setting.
    SourcesLightweight taxonomic profiling of long-read metagenomic datasets with Lemur and Magnet · Methods: Synthetic and simulated datasets; Method evaluation; cached text lines 75–77, 97–99; uncertainty/repeat-run/statistical-comparison passages
Profile review details

Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.

Stable record: reported-task-6330d593980b5b

Evidence table

Inspect claims, sources and review details

Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.

One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.

17 evidence rows matching the loaded filters

Claims, original sources and review scope · Release 2026-09-17-d277315f7d76
Property and statementOriginal source and locationReview and provenance
Diagram caption

Conceptual summary of the cited evaluation; exact task configuration and source version remain part of the protocol.

Individual claims
Lightweight taxonomic profiling of long-read metagenomic datasets with Lemur and Magnet

Original source ↗

Methods: Synthetic and simulated datasets; Method evaluation; cached text lines 75–77, 97–99; uncertainty/repeat-run/statistical-comparison passages; Methods: Synthetic and simulated datasets

Version: PMC archival version PMC11185576.2
Retrieved: 2026-09-16T10:44:03.420493+00:00

source checked

automated source review · 2026-09-16

Audit details

Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.

Field: attributes.profile.diagram.caption

Source artifact SHA-256: 4afb9195da447916eb6f733816e3640741c7ade08ea8920d205c3be7b3cce27a

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Diagram steps

["Input: Long sequencing reads.","Evaluation: Known-composition simulated or mock-community samples are profiled against reference resources. The Dilthey simulation has RefSeq species representatives for 94 of 96 strains; the additional simulation selects RefSeq-represented species.","Readout: Genus/species precision, recall and F1 after abundance thresholding; normalized L1 loss or Spearman correlation for abundance."]

Individual claims
Lightweight taxonomic profiling of long-read metagenomic datasets with Lemur and Magnet

Original source ↗

Methods: Synthetic and simulated datasets; Method evaluation; cached text lines 75–77, 97–99; uncertainty/repeat-run/statistical-comparison passages; Methods: Synthetic and simulated datasets

Version: PMC archival version PMC11185576.2
Retrieved: 2026-09-16T10:44:03.420493+00:00

source checked

automated source review · 2026-09-16

Audit details

Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.

Field: attributes.profile.diagram.steps

Source artifact SHA-256: 4afb9195da447916eb6f733816e3640741c7ade08ea8920d205c3be7b3cce27a

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Diagram title

Computational evaluation flow

Individual claims
Lightweight taxonomic profiling of long-read metagenomic datasets with Lemur and Magnet

Original source ↗

Methods: Synthetic and simulated datasets; Method evaluation; cached text lines 75–77, 97–99; uncertainty/repeat-run/statistical-comparison passages; Methods: Synthetic and simulated datasets

Version: PMC archival version PMC11185576.2
Retrieved: 2026-09-16T10:44:03.420493+00:00

source checked

automated source review · 2026-09-16

Audit details

Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.

Field: attributes.profile.diagram.title

Source artifact SHA-256: 4afb9195da447916eb6f733816e3640741c7ade08ea8920d205c3be7b3cce27a

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Datasets

Synthetic and simulated communities with known reference composition.

Individual claims
Lightweight taxonomic profiling of long-read metagenomic datasets with Lemur and Magnet

Original source ↗

Methods: Synthetic and simulated datasets; Method evaluation; cached text lines 75–77, 97–99; uncertainty/repeat-run/statistical-comparison passages

Version: PMC archival version PMC11185576.2
Retrieved: 2026-09-16T10:44:03.420493+00:00

source checked

automated source review · 2026-09-16

Audit details

Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.

Field: attributes.profile.facts.0.value

Source artifact SHA-256: 4afb9195da447916eb6f733816e3640741c7ade08ea8920d205c3be7b3cce27a

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Splits

Known-composition simulated or mock-community samples are profiled against reference resources. The Dilthey simulation has RefSeq species representatives for 94 of 96 strains; the additional simulation selects RefSeq-represented species.

Individual claims
Lightweight taxonomic profiling of long-read metagenomic datasets with Lemur and Magnet

Original source ↗

Methods: Synthetic and simulated datasets

Version: PMC archival version PMC11185576.2
Retrieved: 2026-09-16T10:44:03.420493+00:00

source checked

automated source review · 2026-09-16

Audit details

Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.

Field: attributes.profile.facts.1.value

Source artifact SHA-256: 4afb9195da447916eb6f733816e3640741c7ade08ea8920d205c3be7b3cce27a

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Adaptation

Lemur and the comparator profilers use their reference resources; Magnet additionally aligns reads to cluster-representative genomes with minimap2. This is reference-based taxonomic profiling, rather than fitting a classifier on labeled train/test folds.

Individual claims
Lightweight taxonomic profiling of long-read metagenomic datasets with Lemur and Magnet

Original source ↗

Methods: Competitive read alignment with Magnet; Method Comparison

Version: PMC archival version PMC11185576.2
Retrieved: 2026-09-16T10:44:03.420493+00:00

source checked

automated source review · 2026-09-16

Audit details

Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.

Field: attributes.profile.facts.10.value

Source artifact SHA-256: 4afb9195da447916eb6f733816e3640741c7ade08ea8920d205c3be7b3cce27a

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Metrics

Genus/species precision, recall and F1 after abundance thresholding; normalized L1 loss or Spearman correlation for abundance.

Individual claims
Lightweight taxonomic profiling of long-read metagenomic datasets with Lemur and Magnet

Original source ↗

Methods: Synthetic and simulated datasets; Method evaluation; cached text lines 75–77, 97–99; uncertainty/repeat-run/statistical-comparison passages

Version: PMC archival version PMC11185576.2
Retrieved: 2026-09-16T10:44:03.420493+00:00

source checked

automated source review · 2026-09-16

Audit details

Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.

Field: attributes.profile.facts.2.value

Source artifact SHA-256: 4afb9195da447916eb6f733816e3640741c7ade08ea8920d205c3be7b3cce27a

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Baselines

Lemur v1.0.1 is evaluated alone and with Magnet against Centrifuger v1.0.0, Kraken 2 v2.1.3, Melon v0.1.0, MetaMaps commit 633d2e0 and Sourmash v4.8.2. Melon lacks fungal references; the paper also reports bacterial-only comparisons for fungal-containing datasets.

Individual claims
Lightweight taxonomic profiling of long-read metagenomic datasets with Lemur and Magnet

Original source ↗

Methods: Method Comparison

Version: PMC archival version PMC11185576.2
Retrieved: 2026-09-16T10:44:03.420493+00:00

source checked

automated source review · 2026-09-16

Audit details

Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.

Field: attributes.profile.facts.3.value

Source artifact SHA-256: 4afb9195da447916eb6f733816e3640741c7ade08ea8920d205c3be7b3cce27a

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Leakage controls

Reference availability is part of this identification task: 94 of 96 strains in the first simulation have a corresponding RefSeq species representative. The additional metagenome simulation deliberately selects species with RefSeq representative genomes and available MAGs. These settings do not establish novel-species generalization.

Individual claims
Lightweight taxonomic profiling of long-read metagenomic datasets with Lemur and Magnet

Original source ↗

Methods: Simulated data from Dilthey et al. 2019; Simulated metagenome

Version: PMC archival version PMC11185576.2
Retrieved: 2026-09-16T10:44:03.420493+00:00

source checked

automated source review · 2026-09-16

Audit details

Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.

Field: attributes.profile.facts.4.value

Source artifact SHA-256: 4afb9195da447916eb6f733816e3640741c7ade08ea8920d205c3be7b3cce27a

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Uncertainty

Mean and standard deviation across five replicate runs are reported for the Zymo EVEN and LOG comparisons; read subsampling also uses repeated seeds.

Individual claims
Lightweight taxonomic profiling of long-read metagenomic datasets with Lemur and Magnet

Original source ↗

Methods: Synthetic and simulated datasets; Method evaluation; cached text lines 75–77, 97–99; uncertainty/repeat-run/statistical-comparison passages

Version: PMC archival version PMC11185576.2
Retrieved: 2026-09-16T10:44:03.420493+00:00

source checked

automated source review · 2026-09-16

Audit details

Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.

Field: attributes.profile.facts.5.value

Source artifact SHA-256: 4afb9195da447916eb6f733816e3640741c7ade08ea8920d205c3be7b3cce27a

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Sources and history

Release 2026-09-17-d277315f7d76 · Record review: needs review

2 source records and release historyDownload this release
Technical metadata and extraction receipts

Stable ID: reported-task-6330d593980b5b

areas
microbes-communities
tasks
Long-read taxonomic profiling
entity level
task
version
Not reported
task
Long-read taxonomic profiling
scope note
Paper-specific evaluation task; protocol completeness requires further extraction.
comparison panels
id: part2-lemur-magnet-2024-T1-a3172c2928; title: Species profiling on Dilthey2019 simulated long reads · Table 1:; protocol id: paper-protocol-1080592b8d6a8ce0bd; dataset id: paper-dataset-e8eb310c5d40e647ac; metric: Recall; unit: fraction; direction: higher; result ids: paper-result-80f8cca2c2d9379a1a; paper-result-620e5294f3d001ac96; paper-result-a91cd2aadeccebf9b9; paper-result-b70bf4b4c41e6dd812; paper-result-6da2d59412e925bccd; paper-result-74d2e70b442ba9026b; paper-result-8fd5837795f2fbba5f; source ids: part2-lemur-magnet-2024; source locator: Table 1:: Recall, Species profiling on Dilthey2019 simulated long reads; context: Species-level recall/precision/F 1. MetaMaps results copied from original manuscript; other tool versions reported in Methods. 200,114 reads;96 strain design,94 strains with RefSeq species representatives; average read 4,997 bp.; caveats: MetaMaps is quoted prior evidence, not an independent repeated run.; Reference databases and classifier-versus-profiler task differ.; No interval assigned unless printed in source cell.; review: method: automated_source_review; date: 2026-09-17; id: part2-lemur-magnet-2024-T1-81ab2d846d; title: Species profiling on Dilthey2019 simulated long reads · Table 1:; protocol id: paper-protocol-1080592b8d6a8ce0bd; dataset id: paper-dataset-e8eb310c5d40e647ac; metric: Precision; unit: fraction; direction: higher; result ids: paper-result-2268ad4341b9316164; paper-result-06a95cb8e4f61822eb; paper-result-9a39bac49d7dcf2b79; paper-result-a6256bb040cd29f64f; paper-result-f53c61bd32688408d8; paper-result-e2ceeb0f245677c087; paper-result-0dc829636e77fc0d92; source ids: part2-lemur-magnet-2024; source locator: Table 1:: Precision, Species profiling on Dilthey2019 simulated long reads; context: Species-level recall/precision/F 1. MetaMaps results copied from original manuscript; other tool versions reported in Methods. 200,114 reads;96 strain design,94 strains with RefSeq species representatives; average read 4,997 bp.; caveats: MetaMaps is quoted prior evidence, not an independent repeated run.; Reference databases and classifier-versus-profiler task differ.; No interval assigned unless printed in source cell.; review: method: automated_source_review; date: 2026-09-17; id: part2-lemur-magnet-2024-T1-e9572b0ed9; title: Species profiling on Dilthey2019 simulated long reads · Table 1:; protocol id: paper-protocol-1080592b8d6a8ce0bd; dataset id: paper-dataset-e8eb310c5d40e647ac; metric: F1 score; unit: fraction; direction: higher; result ids: paper-result-4468fa08040fb0bd80; paper-result-19f79487e8dff1a781; paper-result-25455d4621f0aeb383; paper-result-f9b8d072704477374e; paper-result-03860ee74178c46356; paper-result-fdc2db2e65a30e98e7; paper-result-f8adf74d2b525a6380; source ids: part2-lemur-magnet-2024; source locator: Table 1:: F1 score, Species profiling on Dilthey2019 simulated long reads; context: Species-level recall/precision/F 1. MetaMaps results copied from original manuscript; other tool versions reported in Methods. 200,114 reads;96 strain design,94 strains with RefSeq species representatives; average read 4,997 bp.; caveats: MetaMaps is quoted prior evidence, not an independent repeated run.; Reference databases and classifier-versus-profiler task differ.; No interval assigned unless printed in source cell.; review: method: automated_source_review; date: 2026-09-17
benchmark research
review date: 2026-09-17; status: complete_comparison_tables_extracted_pending_publication_review; primary sources: part2-lemur-magnet-2024; inspected locators: Tables1–3; simulated Dilthey2019 and ZymoEVEN/LOG cohorts; searched queries: Lightweight taxonomic profiling of long-read metagenomic datasets with Lemur and Magnet primary paper benchmark results; gaps: Independent batch review before import; preserve existing observation identities.; claim scope: Primary-source discovery and table/protocol screening; source checked is not independently reproduced. Raw acquisitions not automatically numerical publication approval.
historical missing metadata
protocol version: not_reported_in_legacy_extract; split: not_reported_in_legacy_extract
metadata review scope
historical_missing_metadata preserves the original discovery state. Current descriptive evidence and missingness are recorded in profile.facts; numerical-result review is separate.
legacy kinds
benchmark
entity classification
review date: 2026-09-17; rationale: This source-scoped record identifies the biological prediction task and holds its paper context. Preserve the existing task identity; exact split, model adaptation and scoring remain in linked evaluations or separate protocol records.; source ids: lemur-magnet-2024; source locator: Methods: Synthetic and simulated datasets; Method evaluation; cached text lines 75–77, 97–99; uncertainty/repeat-run/statistical-comparison passages; ambiguities: A paper- or suite-specific task may constrain some inputs or metrics; that alone does not make it interchangeable with a complete versioned protocol. No protocol equivalence is inferred.; Some legacy profile Entity type facts use the generic phrase computational evaluation protocol. That boilerplate is not sufficient to establish a single fixed protocol identity or to merge this task with another protocol record.
Related records

Suggest a correction