Model type
Tandem mass-spectral transformer
DreaMS learns molecular representations from tandem mass spectra using self-supervised learning.
Conceptual summary of the documented data flow; optional inputs and configured downstream stages must be reported for a reproducible evaluation.
Tandem mass-spectral transformer
MS/MS spectra with the required peak and acquisition information.
Spectrum embeddings or predictions from separately fine-tuned spectral tasks.
Official project documentation and implementation: https://github.com/pluskal-lab/DreaMS
Source reviewed · Automated source review, 2026-09-16. All specifications and missing details
1 evaluation · 8 metric rows. Different protocols are not a single leaderboard.
Applied filters: All linked evaluations
| Tested configuration | Protocol and dataset | Finding | Evidence and details |
|---|---|---|---|
| Configuration: DreaMS fingerprint (ours) | Protocol: MIST CANOPUS retrieval Cos. similarity: MIST CANOPUS fingerprint retrieval: Cos. similarity Dataset subset: MIST CANOPUS benchmark test set used in DreaMS Extended Data Table 1 (MIST CANOPUS retrieval split) | 0.646 fingerprint_cosine_similarity dimensionless · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · source checkedMethods, coverage and sourceMIST CANOPUS PubChem candidate retrieval; candidate positives share true molecule InChIKey first14 characters, negatives share molecular formula; rank cosine similarity of predicted and candidate fingerprints; accuracy@k is fraction of spectra with positive candidate in top k, printed as percent. Full test set named in Extended Data Table 1; exact split version/hash unextracted; 8000 spectra/7000 molecules describes total benchmark, not confirmed test denominator. Aggregation: Not reported dreams: Journal full-text XML · Extended Data Table 1 (XML Tab1), row DreaMS fingerprint (ours), column Cos. similarity |
| Configuration: DreaMS fingerprint (ours) | Protocol: MIST CANOPUS retrieval Top 1: MIST CANOPUS fingerprint retrieval: Top 1 Dataset subset: MIST CANOPUS benchmark test set used in DreaMS Extended Data Table 1 (MIST CANOPUS retrieval split) | 32.731% accuracy_at_1 percent · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · source checkedMethods, coverage and sourceDreaMS fingerprint (ours) on MIST CANOPUS retrieval Top 1: MIST CANOPUS fingerprint retrieval: Top 1 MIST CANOPUS PubChem candidate retrieval; candidate positives share true molecule InChIKey first14 characters, negatives share molecular formula; rank cosine similarity of predicted and candidate fingerprints; accuracy@k is fraction of spectra with positive candidate in top k, printed as percent. Full test set named in Extended Data Table 1; exact split version/hash unextracted; 8000 spectra/7000 molecules describes total benchmark, not confirmed test denominator. Aggregation: Not reported dreams: Journal full-text XML · Extended Data Table 1 (XML Tab1), row DreaMS fingerprint (ours), column Top 1 |
| Configuration: DreaMS fingerprint (ours) | Protocol: MIST CANOPUS retrieval Top 10: MIST CANOPUS fingerprint retrieval: Top 10 Dataset subset: MIST CANOPUS benchmark test set used in DreaMS Extended Data Table 1 (MIST CANOPUS retrieval split) | 67.719% accuracy_at_10 percent · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · source checkedMethods, coverage and sourceMIST CANOPUS PubChem candidate retrieval; candidate positives share true molecule InChIKey first14 characters, negatives share molecular formula; rank cosine similarity of predicted and candidate fingerprints; accuracy@k is fraction of spectra with positive candidate in top k, printed as percent. Full test set named in Extended Data Table 1; exact split version/hash unextracted; 8000 spectra/7000 molecules describes total benchmark, not confirmed test denominator. Aggregation: Not reported dreams: Journal full-text XML · Extended Data Table 1 (XML Tab1), row DreaMS fingerprint (ours), column Top 10 |
| Configuration: DreaMS fingerprint (ours) | Protocol: MIST CANOPUS retrieval Top 100: MIST CANOPUS fingerprint retrieval: Top 100 Dataset subset: MIST CANOPUS benchmark test set used in DreaMS Extended Data Table 1 (MIST CANOPUS retrieval split) | 87.121% accuracy_at_100 percent · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · source checkedMethods, coverage and sourceMIST CANOPUS PubChem candidate retrieval; candidate positives share true molecule InChIKey first14 characters, negatives share molecular formula; rank cosine similarity of predicted and candidate fingerprints; accuracy@k is fraction of spectra with positive candidate in top k, printed as percent. Full test set named in Extended Data Table 1; exact split version/hash unextracted; 8000 spectra/7000 molecules describes total benchmark, not confirmed test denominator. Aggregation: Not reported dreams: Journal full-text XML · Extended Data Table 1 (XML Tab1), row DreaMS fingerprint (ours), column Top 100 |
| Configuration: DreaMS fingerprint (ours) | Protocol: MIST CANOPUS retrieval Top 20: MIST CANOPUS fingerprint retrieval: Top 20 Dataset subset: MIST CANOPUS benchmark test set used in DreaMS Extended Data Table 1 (MIST CANOPUS retrieval split) | 75.390% accuracy_at_20 percent · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · source checkedMethods, coverage and sourceMIST CANOPUS PubChem candidate retrieval; candidate positives share true molecule InChIKey first14 characters, negatives share molecular formula; rank cosine similarity of predicted and candidate fingerprints; accuracy@k is fraction of spectra with positive candidate in top k, printed as percent. Full test set named in Extended Data Table 1; exact split version/hash unextracted; 8000 spectra/7000 molecules describes total benchmark, not confirmed test denominator. Aggregation: Not reported dreams: Journal full-text XML · Extended Data Table 1 (XML Tab1), row DreaMS fingerprint (ours), column Top 20 |
| Configuration: DreaMS fingerprint (ours) | Protocol: MIST CANOPUS retrieval Top 200: MIST CANOPUS fingerprint retrieval: Top 200 Dataset subset: MIST CANOPUS benchmark test set used in DreaMS Extended Data Table 1 (MIST CANOPUS retrieval split) | 90.771% accuracy_at_200 percent · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · source checkedMethods, coverage and sourceMIST CANOPUS PubChem candidate retrieval; candidate positives share true molecule InChIKey first14 characters, negatives share molecular formula; rank cosine similarity of predicted and candidate fingerprints; accuracy@k is fraction of spectra with positive candidate in top k, printed as percent. Full test set named in Extended Data Table 1; exact split version/hash unextracted; 8000 spectra/7000 molecules describes total benchmark, not confirmed test denominator. Aggregation: Not reported dreams: Journal full-text XML · Extended Data Table 1 (XML Tab1), row DreaMS fingerprint (ours), column Top 200 |
| Configuration: DreaMS fingerprint (ours) | Protocol: MIST CANOPUS retrieval Top 5: MIST CANOPUS fingerprint retrieval: Top 5 Dataset subset: MIST CANOPUS benchmark test set used in DreaMS Extended Data Table 1 (MIST CANOPUS retrieval split) | 59.352% accuracy_at_5 percent · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · source checkedMethods, coverage and sourceDreaMS fingerprint (ours) on MIST CANOPUS retrieval Top 5: MIST CANOPUS fingerprint retrieval: Top 5 MIST CANOPUS PubChem candidate retrieval; candidate positives share true molecule InChIKey first14 characters, negatives share molecular formula; rank cosine similarity of predicted and candidate fingerprints; accuracy@k is fraction of spectra with positive candidate in top k, printed as percent. Full test set named in Extended Data Table 1; exact split version/hash unextracted; 8000 spectra/7000 molecules describes total benchmark, not confirmed test denominator. Aggregation: Not reported dreams: Journal full-text XML · Extended Data Table 1 (XML Tab1), row DreaMS fingerprint (ours), column Top 5 |
| Configuration: DreaMS fingerprint (ours) | Protocol: MIST CANOPUS retrieval Top 50: MIST CANOPUS fingerprint retrieval: Top 50 Dataset subset: MIST CANOPUS benchmark test set used in DreaMS Extended Data Table 1 (MIST CANOPUS retrieval split) | 82.404% accuracy_at_50 percent · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · source checkedMethods, coverage and sourceMIST CANOPUS PubChem candidate retrieval; candidate positives share true molecule InChIKey first14 characters, negatives share molecular formula; rank cosine similarity of predicted and candidate fingerprints; accuracy@k is fraction of spectra with positive candidate in top k, printed as percent. Full test set named in Extended Data Table 1; exact split version/hash unextracted; 8000 spectra/7000 molecules describes total benchmark, not confirmed test denominator. Aggregation: Not reported dreams: Journal full-text XML · Extended Data Table 1 (XML Tab1), row DreaMS fingerprint (ours), column Top 50 |
Source checking is not independent reproduction. Release 2026-09-23-2b89723c6dd9.
Related profile: DreaMS. This page retains the exact record and its evaluation context.
Supervised fine-tuned DreaMS fingerprint prediction configuration, distinct from frozen embeddings.
DreaMS learns molecular representations from tandem mass spectra using self-supervised learning. PeakEncoder maps spectral peaks to continuous features; SpectrumEncoder uses transformer blocks; a task-specific PeakDecoder maps the contextual features to predictions. The documented inputs are MS/MS spectra with the required peak and acquisition information. The output consists of spectrum embeddings or predictions from separately fine-tuned spectral tasks.
Pretrained model and task-specific fine-tunes distributed separately via the linked Zenodo release. The paper describes retaining 60 spectral peaks for the transformer; this is peak-count preprocessing, not a nucleotide or protein context.
Inspected pinned official documentation, relevant implementation files and named primary-paper sections. Claims are limited to those artifacts. Remaining field extraction and identity conflicts are explicit; no new performance claims, model runs or human review are implied.
Stable record: discovery-model-dreamsExplanatory profile: source reviewed · Automated source review, 2026-09-16. Review applies to the cited claims; unresolved fields are listed below. Numerical results retain their own review status.
| Property | Description and evidence |
|---|---|
| Model type | Tandem mass-spectral transformerSources (3)pluskal-lab/DreaMS: README.md; https://zenodo.org/api/records/10997887: page.html; dreams: Journal full-text XML · Paper: DreaMS neural network architecture and Hyperparameters, ablation studies, implementation details and benchmarking; README.md: models/data links; Zenodo record 10997887 metadata.license and files; DreaMS paper Introduction and GeMS mining methods |
| Architecture | PeakEncoder maps spectral peaks to continuous features; SpectrumEncoder uses transformer blocks; a task-specific PeakDecoder maps the contextual features to predictions.Sources (3)pluskal-lab/DreaMS: README.md; https://zenodo.org/api/records/10997887: page.html; dreams: Journal full-text XML · Paper: DreaMS neural network architecture and Hyperparameters, ablation studies, implementation details and benchmarking; README.md: models/data links; Zenodo record 10997887 metadata.license and files; DreaMS paper Introduction and GeMS mining methods |
| Inputs | MS/MS spectra with the required peak and acquisition information.Sources (3)pluskal-lab/DreaMS: README.md; https://zenodo.org/api/records/10997887: page.html; dreams: Journal full-text XML · Paper: DreaMS neural network architecture and Hyperparameters, ablation studies, implementation details and benchmarking; README.md: models/data links; Zenodo record 10997887 metadata.license and files; DreaMS paper Introduction and GeMS mining methods |
| Outputs | Spectrum embeddings or predictions from separately fine-tuned spectral tasks.Sources (3)pluskal-lab/DreaMS: README.md; https://zenodo.org/api/records/10997887: page.html; dreams: Journal full-text XML · Paper: DreaMS neural network architecture and Hyperparameters, ablation studies, implementation details and benchmarking; README.md: models/data links; Zenodo record 10997887 metadata.license and files; DreaMS paper Introduction and GeMS mining methods |
| Parameters | 116M for the complete self-supervised network reported in the DreaMS paper; downstream embedding-only configurations may have fewer parameters.Sources (3)pluskal-lab/DreaMS: README.md; https://zenodo.org/api/records/10997887: page.html; dreams: Journal full-text XML · Paper: DreaMS neural network architecture and Hyperparameters, ablation studies, implementation details and benchmarking; README.md: models/data links; Zenodo record 10997887 metadata.license and files; DreaMS paper Introduction and GeMS mining methods |
| Known versions | Pretrained model and task-specific fine-tunes distributed separately via the linked Zenodo release.Sources (3)pluskal-lab/DreaMS: README.md; https://zenodo.org/api/records/10997887: page.html; dreams: Journal full-text XML · Paper: DreaMS neural network architecture and Hyperparameters, ablation studies, implementation details and benchmarking; README.md: models/data links; Zenodo record 10997887 metadata.license and files; DreaMS paper Introduction and GeMS mining methods |
| Training data | GeMS mined from MassIVE/GNPS; the paper identifies GeMS-A10, approximately 24M spectra, as the high-quality pretraining subset.Sources (3)pluskal-lab/DreaMS: README.md; https://zenodo.org/api/records/10997887: page.html; dreams: Journal full-text XML · Paper: DreaMS neural network architecture and Hyperparameters, ablation studies, implementation details and benchmarking; README.md: models/data links; Zenodo record 10997887 metadata.license and files; DreaMS paper Introduction and GeMS mining methods |
| Training cutoff | GeMS mining selected GNPS-tagged MassIVE studies available as of November 2022; subsequent task-specific datasets have separate provenance.Sources (3)pluskal-lab/DreaMS: README.md; https://zenodo.org/api/records/10997887: page.html; dreams: Journal full-text XML · Paper: DreaMS neural network architecture and Hyperparameters, ablation studies, implementation details and benchmarking; README.md: models/data links; Zenodo record 10997887 metadata.license and files; DreaMS paper Introduction and GeMS mining methods |
| Context limits | The paper describes retaining 60 spectral peaks for the transformer; this is peak-count preprocessing, not a nucleotide or protein context.Sources (3)pluskal-lab/DreaMS: README.md; https://zenodo.org/api/records/10997887: page.html; dreams: Journal full-text XML · Paper: DreaMS neural network architecture and Hyperparameters, ablation studies, implementation details and benchmarking; README.md: models/data links; Zenodo record 10997887 metadata.license and files; DreaMS paper Introduction and GeMS mining methods |
| Weights licence | CC-BY-4.0 for the embedding_model.ckpt and ssl_model.ckpt files in author-linked Zenodo record 10997887; separate from the MIT code licence.Sources (3)pluskal-lab/DreaMS: README.md; https://zenodo.org/api/records/10997887: page.html; dreams: Journal full-text XML · Paper: DreaMS neural network architecture and Hyperparameters, ablation studies, implementation details and benchmarking; README.md: models/data links; Zenodo record 10997887 metadata.license and files; DreaMS paper Introduction and GeMS mining methods |
| Access | Official project documentation and implementation: https://github.com/pluskal-lab/DreaMSSources (3)pluskal-lab/DreaMS: README.md; https://zenodo.org/api/records/10997887: page.html; dreams: Journal full-text XML · Paper: DreaMS neural network architecture and Hyperparameters, ablation studies, implementation details and benchmarking; README.md: models/data links; Zenodo record 10997887 metadata.license and files; DreaMS paper Introduction and GeMS mining methods |
| Code licence | MITSourcespluskal-lab/DreaMS: LICENSE · LICENSE: licence text |
Source checking verifies the cited claim or transcription. It does not establish independent reproduction.
Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.
One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.
1 evidence row matching the loaded filters
| Property and statement | Original source and location | Review and provenance |
|---|---|---|
| Relationship: family discovery-model-dreams Individual claims | dreams: Journal full-text XML Extended Data Table 1 (XML Tab1), DreaMS fingerprint (ours) row Version: Retrieved page snapshot; no immutable publisher revision supplied | source checked automated source review · 2026-09-23 Audit detailsSource-backed evaluated identity only; no independent reproduction. Field: Claim: dreams-mist-retrieval-method-dreams-fingerprint-ours-discovery-model-dreams-identity-claim Source artifact SHA-256: Hash scope: SHA-256 of retrieved original artifact bytes Format: original_artifact |
View linked audit checks and correction history
Release 2026-09-23-2b89723c6dd9 · Record review: source checked
Stable ID: dreams-mist-retrieval-method-dreams-fingerprint-ours