rewire.it
Task

CAMI metagenome assembly

The CAMI assembly task concerns reconstructing metagenomic sequence assemblies; the homepage links MetaQUAST evaluation resources.

Sourcescami official source · Official CAMI homepage: initiative description; challenges; dataset correction notices; toolkit citations

0 evaluations · 0 metric rows

At a glance

Inputs, training, access and other details

Explanatory profile: limited source coverage · Automated source review, 2026-09-16. Review applies to the cited claims; unresolved fields are listed below. Numerical results retain their own review status.

Data, procedure and scoring
PropertyDescription and evidence
DatasetsCAMI II marine, strain-madness and plant-associated simulated metagenomes provide short-read, long-read and hybrid conditions. Genome truth and gold-standard assemblies are available after the blinded challenge; preserve dataset and sequencing condition.
Sourcescami2 primary benchmark evidence · Methods: Challenge datasets and Challenge organization
SplitsCAMI II supplied public-genome practice datasets with ground truth before its blinded challenge. Challenge datasets were marine, strain-madness and plant-associated communities; these are challenge conditions, not a standard supervised train/validation/test partition.
Sourcescami2 primary benchmark evidence · Methods: Challenge datasets, Challenge organization, Evaluation metrics; Figures 2–4; Table 1
MetricsMetaQUAST evaluates genome fraction, mismatches per 100 kb, duplication ratio, NGA50 and misassemblies; strain precision and recall additionally measure high-quality strain reconstruction. Undefined per-genome NGA50 is set to zero before the reported genome average.
Sourcescami2 primary benchmark evidence · Methods: Assembly metrics; Figure 1
BaselinesSubmitted programs are compared under the same data condition. Gold-standard assemblies and MEGAHIT assemblies separate binning performance from upstream assembly error; published method identities and versions are listed in Table 1.
Sourcescami2 primary benchmark evidence · Methods: Challenge datasets, Challenge organization, Evaluation metrics; Figures 2–4; Table 1
Leakage controlsChallenge genome data and metadata were kept confidential until the challenge ended. Public reference collections dated 8 January 2019 were supplied for reference-based methods. CAMI II also includes public genomes, so novelty is stratified rather than assumed for every organism.
Sourcescami2 primary benchmark evidence · Methods: Challenge datasets, Challenge organization, Evaluation metrics; Figures 2–4; Table 1
UncertaintyFigure 1 and the assembly-metrics Methods define descriptive per-genome and dataset summaries, but no universal bootstrap or repeated-seed interval for all assembly scores. · Not reported in inspected sources
Sourcescami2 primary benchmark evidence · Methods: Assembly metrics; Figure 1
Entity typeConstituent benchmark task: CAMI metagenome assembly
Sourcescami official source · Official CAMI homepage: initiative description; challenges; dataset correction notices; toolkit citations
OrganismsMicrobial communities; CAMI III includes longitudinal human-gut samples.
Sourcescami official source · Official CAMI homepage: initiative description; challenges; dataset correction notices; toolkit citations
AssaysChallenge-specific metagenomic sequence data and reference composition.
Sourcescami official source · Official CAMI homepage: initiative description; challenges; dataset correction notices; toolkit citations
Allowed inputsReleased sequence data and track-specific reference resources.
Sourcescami official source · Official CAMI homepage: initiative description; challenges; dataset correction notices; toolkit citations
AdaptationMethods process challenge inputs; a challenge edition and track determine resource rules.
Sourcescami official source · Official CAMI homepage: initiative description; challenges; dataset correction notices; toolkit citations

How it works

How it worksEvaluation procedure
Evaluation procedure1. Allowed inputs: Released sequence data and track-specific reference resources.. Then: 2. Splits: CAMI II supplied public-genome practice datasets with ground truth before its blinded challenge. Challenge datasets were marine, strain-madness and plant-associated communities; these are challenge conditions, not a standard supervised train/validation/test partition.. Then: 3. Metrics: MetaQUAST evaluates genome fraction, mismatches per 100 kb, duplication ratio, NGA50 and misassemblies; strain precision and recall additionally measure high-quality strain reconstruction. Undefined per-genome NGA50 is set to zero before the reported genome average.Evaluation procedure1. Allowed inputs: Released sequence data and track-specific reference resources.. Then: 2. Splits: CAMI II supplied public-genome practice datasets with ground truth before its blinded challenge. Challenge datasets were marine, strain-madness and plant-associated communities; these are challenge conditions, not a standard supervised train/validation/test partition.. Then: 3. Metrics: MetaQUAST evaluates genome fraction, mismatches per 100 kb, duplication ratio, NGA50 and misassemblies; strain precision and recall additionally measure high-quality strain reconstruction. Undefined per-genome NGA50 is set to zero before the reported genome average.Evaluation procedure1. Allowed inputs: Released sequence data and track-specific reference resources.. Then: 2. Splits: CAMI II supplied public-genome practice datasets with ground truth before its blinded challenge. Challenge datasets were marine, strain-madness and plant-associated communities; these are challenge conditions, not a standard supervised train/validation/test partition.. Then: 3. Metrics: MetaQUAST evaluates genome fraction, mismatches per 100 kb, duplication ratio, NGA50 and misassemblies; strain precision and recall additionally measure high-quality strain reconstruction. Undefined per-genome NGA50 is set to zero before the reported genome average.

Conceptual procedure. Task variants and protocol versions retain their separate scoring conditions.

Sources (2)cami official source; cami2 primary benchmark evidence · Official CAMI homepage: initiative description; challenges; dataset correction notices; toolkit citations; Methods: Challenge datasets, Challenge organization, Evaluation metrics; Figures 2–4; Table 1; Methods: Assembly metrics; Figure 1
Evaluation methodology

CAMI is a series of blinded metagenomic software challenges. In CAMI II, participants received simulated short and long reads from defined communities and submitted assemblies, genome bins, taxonomic assignments or abundance profiles. Reference truth was used only for scoring, and software versions, input read types and community conditions were kept distinct.

Sourcescami2 primary benchmark evidence · Methods: Challenge datasets, Challenge organization, Evaluation metrics; Figures 2–4; Table 1

Evaluation design

Benchmarks bring together tasks and protocols. A task describes the biological question; a protocol defines a particular test.

Benchmarks

These source-backed links do not make different protocols or scores interchangeable.

Tested entities and results

Release 2026-09-17-d277315f7d76 · 0 evaluations · 0 metric rows. Different protocols are not a single leaderboard.

No evaluations linked in this release.

Papers and result coverage

Last literature check: 2026-09-17. Primary-source discovery and table/protocol screening; source checked is not independently reproduced. Raw acquisitions not automatically numerical publication approval.

Paper or primary resourceVersionReference
Critical Assessment of Metagenome Interpretation: the second round of challengesPMC9007738Read source
DOI: 10.1038/s41592-022-01431-4

What is still missing

  • Main Table 1 gives best-ranked software, not a complete numerical score table. Assembly, genome binning, taxonomic binning and taxonomic profiling are separate tasks; strain diversity and sample cohorts must be separated. Further source-data extraction remains before new score publication.
Search and extraction details

source found structured extraction pending

Searches

  • CAMI II metagenome benchmarking 2022 supplementary results table

Evidence locations

  • Table1 rankings; main figures and linked supplementary material

Strengths and limitations

Strengths and considerations

  • Multiple tracks distinguish assembly, binning and abundance estimation.
    Sourcescami official source · Official CAMI homepage: initiative description; challenges; dataset correction notices; toolkit citations

Limitations and conditions

  • This profile documents CAMI II as a concrete protocol example. Other CAMI rounds may use different genomes, reference databases and metrics; strain diversity and input assembly quality materially change the task.
    Sourcescami2 primary benchmark evidence · Methods: Challenge datasets, Challenge organization, Evaluation metrics; Figures 2–4; Table 1
Profile review details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Stable record: discovery-benchmark-cami-metagenome-assembly

Evidence table

Inspect claims, sources and review details

Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.

One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.

22 evidence rows matching the loaded filters

Claims, original sources and review scope · Release 2026-09-17-d277315f7d76
Property and statementOriginal source and locationReview and provenance
Diagram caption

Conceptual procedure. Task variants and protocol versions retain their separate scoring conditions.

Individual claims
cami official source

Original source ↗

Official CAMI homepage: initiative description; challenges; dataset correction notices; toolkit citations; Methods: Challenge datasets, Challenge organization, Evaluation metrics; Figures 2–4; Table 1; Methods: Assembly metrics; Figure 1

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: Retrieved website snapshot sha256:17825bf33280f40b596a104c547b57fae5ee5c07f8d60b396d0f4780d47ef9a5
Retrieved: 2026-09-16T10:31:54.339676+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.diagram.caption

Source artifact SHA-256: 17825bf33280f40b596a104c547b57fae5ee5c07f8d60b396d0f4780d47ef9a5

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Diagram caption

Conceptual procedure. Task variants and protocol versions retain their separate scoring conditions.

Individual claims
cami2 primary benchmark evidence

Original source ↗

Official CAMI homepage: initiative description; challenges; dataset correction notices; toolkit citations; Methods: Challenge datasets, Challenge organization, Evaluation metrics; Figures 2–4; Table 1; Methods: Assembly metrics; Figure 1

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: PMC9007738
Retrieved: 2026-09-16T21:04:55.691966+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.diagram.caption

Source artifact SHA-256: da932c1cde8b290e1694fc3cf44d98ed527be1ec7cf41542dd5e9aee75c38704

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Diagram steps

["Allowed inputs: Released sequence data and track-specific reference resources.","Splits: CAMI II supplied public-genome practice datasets with ground truth before its blinded challenge. Challenge datasets were marine, strain-madness and plant-associated communities; these are challenge conditions, not a standard supervised train/validation/test partition.","Metrics: MetaQUAST evaluates genome fraction, mismatches per 100 kb, duplication ratio, NGA50 and misassemblies; strain precision and recall additionally measure high-quality strain reconstruction. Undefined per-genome NGA50 is set to zero before the reported genome average."]

Individual claims
cami official source

Original source ↗

Official CAMI homepage: initiative description; challenges; dataset correction notices; toolkit citations; Methods: Challenge datasets, Challenge organization, Evaluation metrics; Figures 2–4; Table 1; Methods: Assembly metrics; Figure 1

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: Retrieved website snapshot sha256:17825bf33280f40b596a104c547b57fae5ee5c07f8d60b396d0f4780d47ef9a5
Retrieved: 2026-09-16T10:31:54.339676+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.diagram.steps

Source artifact SHA-256: 17825bf33280f40b596a104c547b57fae5ee5c07f8d60b396d0f4780d47ef9a5

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Diagram steps

["Allowed inputs: Released sequence data and track-specific reference resources.","Splits: CAMI II supplied public-genome practice datasets with ground truth before its blinded challenge. Challenge datasets were marine, strain-madness and plant-associated communities; these are challenge conditions, not a standard supervised train/validation/test partition.","Metrics: MetaQUAST evaluates genome fraction, mismatches per 100 kb, duplication ratio, NGA50 and misassemblies; strain precision and recall additionally measure high-quality strain reconstruction. Undefined per-genome NGA50 is set to zero before the reported genome average."]

Individual claims
cami2 primary benchmark evidence

Original source ↗

Official CAMI homepage: initiative description; challenges; dataset correction notices; toolkit citations; Methods: Challenge datasets, Challenge organization, Evaluation metrics; Figures 2–4; Table 1; Methods: Assembly metrics; Figure 1

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: PMC9007738
Retrieved: 2026-09-16T21:04:55.691966+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.diagram.steps

Source artifact SHA-256: da932c1cde8b290e1694fc3cf44d98ed527be1ec7cf41542dd5e9aee75c38704

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Diagram title

Evaluation procedure

Individual claims
cami official source

Original source ↗

Official CAMI homepage: initiative description; challenges; dataset correction notices; toolkit citations; Methods: Challenge datasets, Challenge organization, Evaluation metrics; Figures 2–4; Table 1; Methods: Assembly metrics; Figure 1

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: Retrieved website snapshot sha256:17825bf33280f40b596a104c547b57fae5ee5c07f8d60b396d0f4780d47ef9a5
Retrieved: 2026-09-16T10:31:54.339676+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.diagram.title

Source artifact SHA-256: 17825bf33280f40b596a104c547b57fae5ee5c07f8d60b396d0f4780d47ef9a5

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Diagram title

Evaluation procedure

Individual claims
cami2 primary benchmark evidence

Original source ↗

Official CAMI homepage: initiative description; challenges; dataset correction notices; toolkit citations; Methods: Challenge datasets, Challenge organization, Evaluation metrics; Figures 2–4; Table 1; Methods: Assembly metrics; Figure 1

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: PMC9007738
Retrieved: 2026-09-16T21:04:55.691966+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.diagram.title

Source artifact SHA-256: da932c1cde8b290e1694fc3cf44d98ed527be1ec7cf41542dd5e9aee75c38704

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Datasets

CAMI II marine, strain-madness and plant-associated simulated metagenomes provide short-read, long-read and hybrid conditions. Genome truth and gold-standard assemblies are available after the blinded challenge; preserve dataset and sequencing condition.

Individual claims
cami2 primary benchmark evidence

Original source ↗

Methods: Challenge datasets and Challenge organization

Version: PMC9007738
Retrieved: 2026-09-16T21:04:55.691966+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.facts.0.value

Source artifact SHA-256: da932c1cde8b290e1694fc3cf44d98ed527be1ec7cf41542dd5e9aee75c38704

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Splits

CAMI II supplied public-genome practice datasets with ground truth before its blinded challenge. Challenge datasets were marine, strain-madness and plant-associated communities; these are challenge conditions, not a standard supervised train/validation/test partition.

Individual claims
cami2 primary benchmark evidence

Original source ↗

Methods: Challenge datasets, Challenge organization, Evaluation metrics; Figures 2–4; Table 1

Version: PMC9007738
Retrieved: 2026-09-16T21:04:55.691966+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.facts.1.value

Source artifact SHA-256: da932c1cde8b290e1694fc3cf44d98ed527be1ec7cf41542dd5e9aee75c38704

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Adaptation

Methods process challenge inputs; a challenge edition and track determine resource rules.

Individual claims
cami official source

Original source ↗

Official CAMI homepage: initiative description; challenges; dataset correction notices; toolkit citations

Version: Retrieved website snapshot sha256:17825bf33280f40b596a104c547b57fae5ee5c07f8d60b396d0f4780d47ef9a5
Retrieved: 2026-09-16T10:31:54.339676+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.facts.10.value

Source artifact SHA-256: 17825bf33280f40b596a104c547b57fae5ee5c07f8d60b396d0f4780d47ef9a5

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Metrics

MetaQUAST evaluates genome fraction, mismatches per 100 kb, duplication ratio, NGA50 and misassemblies; strain precision and recall additionally measure high-quality strain reconstruction. Undefined per-genome NGA50 is set to zero before the reported genome average.

Individual claims
cami2 primary benchmark evidence

Original source ↗

Methods: Assembly metrics; Figure 1

Version: PMC9007738
Retrieved: 2026-09-16T21:04:55.691966+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.facts.2.value

Source artifact SHA-256: da932c1cde8b290e1694fc3cf44d98ed527be1ec7cf41542dd5e9aee75c38704

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Sources and history

Release 2026-09-17-d277315f7d76 · Record review: discovered

4 source records and release historyDownload this release
Technical metadata and extraction receipts

Stable ID: discovery-benchmark-cami-metagenome-assembly

areas
microbiome
entity level
task
scope note
Specialist molecular or omics evaluation; protocol details require review before numerical comparison.
task
metagenome assembly
version
Not reported
benchmark research
review date: 2026-09-17; status: source_found_structured_extraction_pending; primary sources: evidence-expansion-p2-evidence-discovery-final-cami2-da932c1cde8b; inspected locators: Table1 rankings; main figures and linked supplementary material; searched queries: CAMI II metagenome benchmarking 2022 supplementary results table; gaps: Main Table 1 gives best-ranked software, not a complete numerical score table. Assembly, genome binning, taxonomic binning and taxonomic profiling are separate tasks; strain diversity and sample cohorts must be separated. Further source-data extraction remains before new score publication.; claim scope: Primary-source discovery and table/protocol screening; source checked is not independently reproduced. Raw acquisitions not automatically numerical publication approval.
historical missing metadata
dataset release: unextracted; metric implementation: unextracted; split manifest: unextracted; version: unextracted
metadata review scope
historical_missing_metadata preserves the original discovery state. Current descriptive evidence and missingness are recorded in profile.facts; numerical-result review is separate.
legacy kinds
benchmark
entity classification
review date: 2026-09-17; rationale: This record identifies the biological prediction question or a suite-specific task, rather than a uniquely fixed evaluated procedure. Preserve its task identity and leave split, model adaptation and scoring details on linked protocols/evaluations.; source ids: evidence-benchmark-cami-snapshot; source locator: Official CAMI homepage: initiative description; challenges; dataset correction notices; toolkit citations; ambiguities: A paper- or suite-specific task may constrain some inputs or metrics; that alone does not make it interchangeable with a complete versioned protocol. No protocol equivalence is inferred.
Related records

Suggest a correction