Model type
Frozen spectral/molecular encoders with learned alignment projections
MSAlign retrieves candidate molecules from tandem mass spectra by aligning pretrained molecular and spectral representations.
Conceptual summary of the documented data flow; optional inputs and configured downstream stages must be reported for a reproducible evaluation.
Frozen spectral/molecular encoders with learned alignment projections
MS/MS spectrum and a set of candidate molecular structures.
Candidate-molecule retrieval scores in a shared representation space.
Official project documentation and implementation: https://arxiv.org/abs/2605.19752
limited source coverage · Automated source review, 2026-09-16. All specifications and missing details
4 evaluations · 12 metric rows. Different protocols are not a single leaderboard.
Applied filters: All linked evaluations
| Tested configuration | Protocol and dataset | Finding | Evidence and details |
|---|---|---|---|
| Configuration: MSAlign | Protocol: MSAlign molecular retrieval MassSpecGym formula split: no formula, R@1: MassSpecGym formula split retrieval with no formula Dataset subset: MassSpecGym formula split, no formula candidate pool (MSAlign molecular retrieval split) | 53.8% recall_at_1 percent · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · source checkedMethods, coverage and sourceSection 5.1 candidate retrieval: 256 mass-matched PubChem candidates per spectrum without formula; formula-conditioned methods can exploit true formula and MSAlign+Filter reduces the pool to about 100. Preserve the named source split; exact original split hashes remain unextracted. Aggregation: Not reported msalign: Primary paper PDF · Table 3, MassSpecGym formula split, row R@1, column 8 (MSAlign) |
| Configuration: MSAlign | Protocol: MSAlign molecular retrieval MassSpecGym formula split: no formula, R@20: MassSpecGym formula split retrieval with no formula Dataset subset: MassSpecGym formula split, no formula candidate pool (MSAlign molecular retrieval split) | 87.1% recall_at_20 percent · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · source checkedMethods, coverage and sourceSection 5.1 candidate retrieval: 256 mass-matched PubChem candidates per spectrum without formula; formula-conditioned methods can exploit true formula and MSAlign+Filter reduces the pool to about 100. Preserve the named source split; exact original split hashes remain unextracted. Aggregation: Not reported msalign: Primary paper PDF · Table 3, MassSpecGym formula split, row R@20, column 8 (MSAlign) |
| Configuration: MSAlign | Protocol: MSAlign molecular retrieval MassSpecGym formula split: no formula, R@5: MassSpecGym formula split retrieval with no formula Dataset subset: MassSpecGym formula split, no formula candidate pool (MSAlign molecular retrieval split) | 73.1% recall_at_5 percent · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · source checkedMethods, coverage and sourceSection 5.1 candidate retrieval: 256 mass-matched PubChem candidates per spectrum without formula; formula-conditioned methods can exploit true formula and MSAlign+Filter reduces the pool to about 100. Preserve the named source split; exact original split hashes remain unextracted. Aggregation: Not reported msalign: Primary paper PDF · Table 3, MassSpecGym formula split, row R@5, column 8 (MSAlign) |
| Configuration: MSAlign | Protocol: MSAlign molecular retrieval MassSpecGym MCES split: no formula, R@1: MassSpecGym MCES split retrieval with no formula Dataset subset: MassSpecGym MCES split, no formula candidate pool (MSAlign molecular retrieval split) | 16.2% recall_at_1 percent · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · source checkedMethods, coverage and sourceSection 5.1 candidate retrieval: 256 mass-matched PubChem candidates per spectrum without formula; formula-conditioned methods can exploit true formula and MSAlign+Filter reduces the pool to about 100. Preserve the named source split; exact original split hashes remain unextracted. Aggregation: Not reported msalign: Primary paper PDF · Table 3, MassSpecGym MCES split, row R@1, column 8 (MSAlign) |
| Configuration: MSAlign | Protocol: MSAlign molecular retrieval MassSpecGym MCES split: no formula, R@20: MassSpecGym MCES split retrieval with no formula Dataset subset: MassSpecGym MCES split, no formula candidate pool (MSAlign molecular retrieval split) | 59.9% recall_at_20 percent · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · source checkedMethods, coverage and sourceSection 5.1 candidate retrieval: 256 mass-matched PubChem candidates per spectrum without formula; formula-conditioned methods can exploit true formula and MSAlign+Filter reduces the pool to about 100. Preserve the named source split; exact original split hashes remain unextracted. Aggregation: Not reported msalign: Primary paper PDF · Table 3, MassSpecGym MCES split, row R@20, column 8 (MSAlign) |
| Configuration: MSAlign | Protocol: MSAlign molecular retrieval MassSpecGym MCES split: no formula, R@5: MassSpecGym MCES split retrieval with no formula Dataset subset: MassSpecGym MCES split, no formula candidate pool (MSAlign molecular retrieval split) | 35.6% recall_at_5 percent · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · source checkedMethods, coverage and sourceSection 5.1 candidate retrieval: 256 mass-matched PubChem candidates per spectrum without formula; formula-conditioned methods can exploit true formula and MSAlign+Filter reduces the pool to about 100. Preserve the named source split; exact original split hashes remain unextracted. Aggregation: Not reported msalign: Primary paper PDF · Table 3, MassSpecGym MCES split, row R@5, column 8 (MSAlign) |
| Configuration: MSAlign | Protocol: MSAlign molecular retrieval NPLIB1: no formula, R@1: NPLIB1 retrieval with no formula Dataset subset: NPLIB1, no formula candidate pool (MSAlign molecular retrieval split) | 31.8% recall_at_1 percent · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · source checkedMethods, coverage and sourceMSAlign on MSAlign molecular retrieval NPLIB1: no formula, R@1: NPLIB1 retrieval with no formula Section 5.1 candidate retrieval: 256 mass-matched PubChem candidates per spectrum without formula; formula-conditioned methods can exploit true formula and MSAlign+Filter reduces the pool to about 100. Preserve the named source split; exact original split hashes remain unextracted. Aggregation: Not reported msalign: Primary paper PDF · Table 3, NPLIB1, row R@1, column 8 (MSAlign) |
| Configuration: MSAlign | Protocol: MSAlign molecular retrieval NPLIB1: no formula, R@20: NPLIB1 retrieval with no formula Dataset subset: NPLIB1, no formula candidate pool (MSAlign molecular retrieval split) | 82.9% recall_at_20 percent · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · source checkedMethods, coverage and sourceMSAlign on MSAlign molecular retrieval NPLIB1: no formula, R@20: NPLIB1 retrieval with no formula Section 5.1 candidate retrieval: 256 mass-matched PubChem candidates per spectrum without formula; formula-conditioned methods can exploit true formula and MSAlign+Filter reduces the pool to about 100. Preserve the named source split; exact original split hashes remain unextracted. Aggregation: Not reported msalign: Primary paper PDF · Table 3, NPLIB1, row R@20, column 8 (MSAlign) |
| Configuration: MSAlign | Protocol: MSAlign molecular retrieval NPLIB1: no formula, R@5: NPLIB1 retrieval with no formula Dataset subset: NPLIB1, no formula candidate pool (MSAlign molecular retrieval split) | 60.4% recall_at_5 percent · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · source checkedMethods, coverage and sourceMSAlign on MSAlign molecular retrieval NPLIB1: no formula, R@5: NPLIB1 retrieval with no formula Section 5.1 candidate retrieval: 256 mass-matched PubChem candidates per spectrum without formula; formula-conditioned methods can exploit true formula and MSAlign+Filter reduces the pool to about 100. Preserve the named source split; exact original split hashes remain unextracted. Aggregation: Not reported msalign: Primary paper PDF · Table 3, NPLIB1, row R@5, column 8 (MSAlign) |
| Configuration: MSAlign | Protocol: MSAlign molecular retrieval Spectraverse: no formula, R@1: Spectraverse retrieval with no formula Dataset subset: Spectraverse, no formula candidate pool (MSAlign molecular retrieval split) | 32.3% recall_at_1 percent · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · source checkedMethods, coverage and sourceSection 5.1 candidate retrieval: 256 mass-matched PubChem candidates per spectrum without formula; formula-conditioned methods can exploit true formula and MSAlign+Filter reduces the pool to about 100. Preserve the named source split; exact original split hashes remain unextracted. Aggregation: Not reported msalign: Primary paper PDF · Table 3, Spectraverse, row R@1, column 8 (MSAlign) |
| Configuration: MSAlign | Protocol: MSAlign molecular retrieval Spectraverse: no formula, R@20: Spectraverse retrieval with no formula Dataset subset: Spectraverse, no formula candidate pool (MSAlign molecular retrieval split) | 79.6% recall_at_20 percent · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · source checkedMethods, coverage and sourceSection 5.1 candidate retrieval: 256 mass-matched PubChem candidates per spectrum without formula; formula-conditioned methods can exploit true formula and MSAlign+Filter reduces the pool to about 100. Preserve the named source split; exact original split hashes remain unextracted. Aggregation: Not reported msalign: Primary paper PDF · Table 3, Spectraverse, row R@20, column 8 (MSAlign) |
| Configuration: MSAlign | Protocol: MSAlign molecular retrieval Spectraverse: no formula, R@5: Spectraverse retrieval with no formula Dataset subset: Spectraverse, no formula candidate pool (MSAlign molecular retrieval split) | 59.1% recall_at_5 percent · higher Uncertainty: Not reported Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · source checkedMethods, coverage and sourceSection 5.1 candidate retrieval: 256 mass-matched PubChem candidates per spectrum without formula; formula-conditioned methods can exploit true formula and MSAlign+Filter reduces the pool to about 100. Preserve the named source split; exact original split hashes remain unextracted. Aggregation: Not reported msalign: Primary paper PDF · Table 3, Spectraverse, row R@5, column 8 (MSAlign) |
Source checking is not independent reproduction. Release 2026-09-23-2b89723c6dd9.
Related profile: MSAlign. This page retains the exact record and its evaluation context.
Evaluated retrieval configuration from MSAlign Table 3; model-specific inputs and training are retained in the protocol.
MSAlign retrieves candidate molecules from tandem mass spectra by aligning pretrained molecular and spectral representations. Frozen DreaMS and ChemBERTa encoders connected by lightweight MLP projections trained with a candidate-based contrastive objective. The documented inputs are MS/MS spectrum and a set of candidate molecular structures. The output consists of candidate-molecule retrieval scores in a shared representation space.
MSAlign arXiv:2605.19752v1, submitted 19 May 2026. The applicable input limits require configuration-specific checking.
Inspected pinned official documentation, relevant implementation files and named primary-paper sections. Claims are limited to those artifacts. Remaining field extraction and identity conflicts are explicit; no new performance claims, model runs or human review are implied.
Stable record: discovery-model-msalignExplanatory profile: limited source coverage · Automated source review, 2026-09-16. Review applies to the cited claims; unresolved fields are listed below. Numerical results retain their own review status.
| Property | Description and evidence |
|---|---|
| Model type | Frozen spectral/molecular encoders with learned alignment projectionsSources (2)https://arxiv.org/abs/2605.19752: page.html; msalign: Primary paper PDF · MSAlign paper Section 3 Architecture and Training, Section 4 on splitting, and Section 5.1 Experimental Setting |
| Architecture | Frozen DreaMS and ChemBERTa encoders connected by lightweight MLP projections trained with a candidate-based contrastive objective.Sources (2)https://arxiv.org/abs/2605.19752: page.html; msalign: Primary paper PDF · MSAlign paper Section 3 Architecture and Training, Section 4 on splitting, and Section 5.1 Experimental Setting |
| Inputs | MS/MS spectrum and a set of candidate molecular structures.Sources (2)https://arxiv.org/abs/2605.19752: page.html; msalign: Primary paper PDF · MSAlign paper Section 3 Architecture and Training, Section 4 on splitting, and Section 5.1 Experimental Setting |
| Outputs | Candidate-molecule retrieval scores in a shared representation space.Sources (2)https://arxiv.org/abs/2605.19752: page.html; msalign: Primary paper PDF · MSAlign paper Section 3 Architecture and Training, Section 4 on splitting, and Section 5.1 Experimental Setting |
| Parameters | Approximately 4M trainable projection parameters; frozen DreaMS and ChemBERTa backbones are reported as 96M and 92M respectively.Sources (2)https://arxiv.org/abs/2605.19752: page.html; msalign: Primary paper PDF · MSAlign paper Section 3 Architecture and Training, Section 4 on splitting, and Section 5.1 Experimental Setting |
| Known versions | MSAlign arXiv:2605.19752v1, submitted 19 May 2026.Sources (2)https://arxiv.org/abs/2605.19752: page.html; msalign: Primary paper PDF · MSAlign paper Section 3 Architecture and Training, Section 4 on splitting, and Section 5.1 Experimental Setting |
| Training data | Projection layers are fitted separately on the NPLIB1, MassSpecGym or Spectraverse training splits. The paper distinguishes spectrum/molecule pair counts from unique molecules and controls candidate retrieval using mass matching.Sources (2)https://arxiv.org/abs/2605.19752: page.html; msalign: Primary paper PDF · MSAlign paper Section 3 Architecture and Training, Section 4 on splitting, and Section 5.1 Experimental Setting |
| Training cutoff | NPLIB1, MassSpecGym and Spectraverse are separately split training resources. The inspected paper does not define one latest measurement date covering all three. · Not reported in inspected sourcesSources (2)https://arxiv.org/abs/2605.19752: page.html; msalign: Primary paper PDF · MSAlign paper Section 3 Architecture and Training, Section 4 on splitting, and Section 5.1 Experimental Setting |
| Context limits | The pipeline inherits spectrum and molecule preprocessing from its frozen DreaMS and ChemBERTa encoders. The inspected MSAlign architecture section does not state a single joint input limit. · Not reported in inspected sourcesSources (2)https://arxiv.org/abs/2605.19752: page.html; msalign: Primary paper PDF · MSAlign paper Section 3 Architecture and Training, Section 4 on splitting, and Section 5.1 Experimental Setting |
| Weights licence | The inspected preprint does not state distribution terms for the learned projection checkpoints. Frozen encoder licences remain separate from projection-weight rights. · Not reported in inspected sourcesSources (2)https://arxiv.org/abs/2605.19752: page.html; msalign: Primary paper PDF · MSAlign paper Section 3 Architecture and Training, Section 4 on splitting, and Section 5.1 Experimental Setting; page.html: inspected official source |
| Access | Official project documentation and implementation: https://arxiv.org/abs/2605.19752Sources (2)https://arxiv.org/abs/2605.19752: page.html; msalign: Primary paper PDF · MSAlign paper Section 3 Architecture and Training, Section 4 on splitting, and Section 5.1 Experimental Setting |
| Code licence | The inspected preprint describes the algorithm but does not supply a separate code licence; no code-distribution permission is inferred from the paper licence. · Not reported in inspected sourcesSourceshttps://arxiv.org/abs/2605.19752: page.html · page.html: inspected official source |
Source checking verifies the cited claim or transcription. It does not establish independent reproduction.
Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.
One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.
1 evidence row matching the loaded filters
| Property and statement | Original source and location | Review and provenance |
|---|---|---|
| Relationship: family discovery-model-msalign Individual claims | msalign: Primary paper PDF Table 3, column MSAlign; Section 5.1 Version: 2605.19752v1 | source checked automated source review · 2026-09-23 Audit detailsSource-backed evaluated identity only; no independent reproduction. Field: Claim: msalign-2026-table3-method-msalign-discovery-model-msalign-identity-claim Source artifact SHA-256: Hash scope: SHA-256 of retrieved original artifact bytes Format: original_artifact |
View linked audit checks and correction history
Release 2026-09-23-2b89723c6dd9 · Record review: source checked
Stable ID: msalign-2026-table3-method-msalign