SugarBase taxonomy phylum; GlycanML official motif split (GlycanML split)
Dataset subset reported in GlycanGT published supplementary archive, Table S4. Exact split manifest remains unextracted; source-table identity is retained.
Evaluation results
10 evaluations · 10 metric rows. Different protocols are not a single leaderboard.
Filter evaluations
Applied filters: All linked evaluations
| Tested configuration | Protocol and dataset | Finding | Evidence and details |
|---|---|---|---|
| Configuration: GlycanAA | Protocol: GlycanML phylum Accuracy: GlycanGT study: phylum Accuracy Dataset subset: SugarBase taxonomy phylum; GlycanML official motif split (GlycanML split) | 0.86686666666666601 ± 0.0142043420591498 accuracy fraction · higher Uncertainty: type: standard_deviation; value: 0.0142043420591498 Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · source checkedMethods, coverage and sourceGlycanAA on GlycanML phylum Accuracy: GlycanGT study: phylum Accuracy Taxonomy: 13,209 glycans total across eight levels, 4–1,737 classes per level. Official fixed GlycanML motif-based train/validation/test splits (8:1:1). Section 2.5 reports class-balanced classifiers, train ∪ validation hyperparameter selection by randomized search with 3-fold cross-validation, followed by one evaluation on the held-out test set; complete procedure repeated with three random seeds, reporting mean and standard deviation. GlycanGT large model pretrained with 35% masking provides [Graph] embeddings to SVM/LightGBM; the selected classifier for each S4 row is not identified. Section 2.6 states that all four graph baselines were trained and evaluated on the same datasets/splits; it does not establish that each baseline used the GlycanGT downstream classifier search. SVM search: 10 iterations, RBF/linear, C logU(1e-3,1e2), gamma logU(1e-4,1e-1); LightGBM: 15 randomized iterations. Do not combine with separate original GlycanML-paper protocols. Aggregation: Not reported GlycanGT published supplementary archive, Table S4; glycangt: Journal full-text XML · btag147_supplementary_data.zip / Table_S4.xlsx / Sheet1!A40:E40 (mean D40, SD E40) |
| Configuration: GlycanAA | Protocol: GlycanML phylum Macro-F1: GlycanGT study: phylum Macro-F1 Dataset subset: SugarBase taxonomy phylum; GlycanML official motif split (GlycanML split) | 0.440066666666666 ± 0.0041428653530296202 macro_f1 dimensionless · higher Uncertainty: type: standard_deviation; value: 0.0041428653530296202 Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · source checkedMethods, coverage and sourceGlycanAA on GlycanML phylum Macro-F1: GlycanGT study: phylum Macro-F1 Taxonomy: 13,209 glycans total across eight levels, 4–1,737 classes per level. Official fixed GlycanML motif-based train/validation/test splits (8:1:1). Section 2.5 reports class-balanced classifiers, train ∪ validation hyperparameter selection by randomized search with 3-fold cross-validation, followed by one evaluation on the held-out test set; complete procedure repeated with three random seeds, reporting mean and standard deviation. GlycanGT large model pretrained with 35% masking provides [Graph] embeddings to SVM/LightGBM; the selected classifier for each S4 row is not identified. Section 2.6 states that all four graph baselines were trained and evaluated on the same datasets/splits; it does not establish that each baseline used the GlycanGT downstream classifier search. SVM search: 10 iterations, RBF/linear, C logU(1e-3,1e2), gamma logU(1e-4,1e-1); LightGBM: 15 randomized iterations. Do not combine with separate original GlycanML-paper protocols. Aggregation: Not reported GlycanGT published supplementary archive, Table S4; glycangt: Journal full-text XML · btag147_supplementary_data.zip / Table_S4.xlsx / Sheet1!A41:E41 (mean D41, SD E41) |
| Pipeline: GlycanGT | Protocol: GlycanML phylum Accuracy: GlycanGT study: phylum Accuracy Dataset subset: SugarBase taxonomy phylum; GlycanML official motif split (GlycanML split) | 0.84149437794704296 ± 0.003324320417088 accuracy fraction · higher Uncertainty: type: standard_deviation; value: 0.003324320417088 Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · source checkedMethods, coverage and sourceGlycanGT on GlycanML phylum Accuracy: GlycanGT study: phylum Accuracy Taxonomy: 13,209 glycans total across eight levels, 4–1,737 classes per level. Official fixed GlycanML motif-based train/validation/test splits (8:1:1). Section 2.5 reports class-balanced classifiers, train ∪ validation hyperparameter selection by randomized search with 3-fold cross-validation, followed by one evaluation on the held-out test set; complete procedure repeated with three random seeds, reporting mean and standard deviation. GlycanGT large model pretrained with 35% masking provides [Graph] embeddings to SVM/LightGBM; the selected classifier for each S4 row is not identified. Section 2.6 states that all four graph baselines were trained and evaluated on the same datasets/splits; it does not establish that each baseline used the GlycanGT downstream classifier search. SVM search: 10 iterations, RBF/linear, C logU(1e-3,1e2), gamma logU(1e-4,1e-1); LightGBM: 15 randomized iterations. Do not combine with separate original GlycanML-paper protocols. Aggregation: Not reported GlycanGT published supplementary archive, Table S4; glycangt: Journal full-text XML · btag147_supplementary_data.zip / Table_S4.xlsx / Sheet1!A69:E69 (mean D69, SD E69) |
| Pipeline: GlycanGT | Protocol: GlycanML phylum Macro-F1: GlycanGT study: phylum Macro-F1 Dataset subset: SugarBase taxonomy phylum; GlycanML official motif split (GlycanML split) | 0.49333063993822701 ± 0.080739293560509295 macro_f1 dimensionless · higher Uncertainty: type: standard_deviation; value: 0.080739293560509295 Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · source checkedMethods, coverage and sourceGlycanGT on GlycanML phylum Macro-F1: GlycanGT study: phylum Macro-F1 Taxonomy: 13,209 glycans total across eight levels, 4–1,737 classes per level. Official fixed GlycanML motif-based train/validation/test splits (8:1:1). Section 2.5 reports class-balanced classifiers, train ∪ validation hyperparameter selection by randomized search with 3-fold cross-validation, followed by one evaluation on the held-out test set; complete procedure repeated with three random seeds, reporting mean and standard deviation. GlycanGT large model pretrained with 35% masking provides [Graph] embeddings to SVM/LightGBM; the selected classifier for each S4 row is not identified. Section 2.6 states that all four graph baselines were trained and evaluated on the same datasets/splits; it does not establish that each baseline used the GlycanGT downstream classifier search. SVM search: 10 iterations, RBF/linear, C logU(1e-3,1e2), gamma logU(1e-4,1e-1); LightGBM: 15 randomized iterations. Do not combine with separate original GlycanML-paper protocols. Aggregation: Not reported GlycanGT published supplementary archive, Table S4; glycangt: Journal full-text XML · btag147_supplementary_data.zip / Table_S4.xlsx / Sheet1!A70:E70 (mean D70, SD E70) |
| Configuration: Graphormer | Protocol: GlycanML phylum Accuracy: GlycanGT study: phylum Accuracy Dataset subset: SugarBase taxonomy phylum; GlycanML official motif split (GlycanML split) | 0.81936900000000001 ± 0.025911 accuracy fraction · higher Uncertainty: type: standard_deviation; value: 0.025911 Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · source checkedMethods, coverage and sourceGraphormer on GlycanML phylum Accuracy: GlycanGT study: phylum Accuracy Taxonomy: 13,209 glycans total across eight levels, 4–1,737 classes per level. Official fixed GlycanML motif-based train/validation/test splits (8:1:1). Section 2.5 reports class-balanced classifiers, train ∪ validation hyperparameter selection by randomized search with 3-fold cross-validation, followed by one evaluation on the held-out test set; complete procedure repeated with three random seeds, reporting mean and standard deviation. GlycanGT large model pretrained with 35% masking provides [Graph] embeddings to SVM/LightGBM; the selected classifier for each S4 row is not identified. Section 2.6 states that all four graph baselines were trained and evaluated on the same datasets/splits; it does not establish that each baseline used the GlycanGT downstream classifier search. SVM search: 10 iterations, RBF/linear, C logU(1e-3,1e2), gamma logU(1e-4,1e-1); LightGBM: 15 randomized iterations. Do not combine with separate original GlycanML-paper protocols. Aggregation: Not reported GlycanGT published supplementary archive, Table S4; glycangt: Journal full-text XML · btag147_supplementary_data.zip / Table_S4.xlsx / Sheet1!A102:E102 (mean D102, SD E102) |
| Configuration: Graphormer | Protocol: GlycanML phylum Macro-F1: GlycanGT study: phylum Macro-F1 Dataset subset: SugarBase taxonomy phylum; GlycanML official motif split (GlycanML split) | 0.38569900000000001 ± 0.015233 macro_f1 dimensionless · higher Uncertainty: type: standard_deviation; value: 0.015233 Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · source checkedMethods, coverage and sourceGraphormer on GlycanML phylum Macro-F1: GlycanGT study: phylum Macro-F1 Taxonomy: 13,209 glycans total across eight levels, 4–1,737 classes per level. Official fixed GlycanML motif-based train/validation/test splits (8:1:1). Section 2.5 reports class-balanced classifiers, train ∪ validation hyperparameter selection by randomized search with 3-fold cross-validation, followed by one evaluation on the held-out test set; complete procedure repeated with three random seeds, reporting mean and standard deviation. GlycanGT large model pretrained with 35% masking provides [Graph] embeddings to SVM/LightGBM; the selected classifier for each S4 row is not identified. Section 2.6 states that all four graph baselines were trained and evaluated on the same datasets/splits; it does not establish that each baseline used the GlycanGT downstream classifier search. SVM search: 10 iterations, RBF/linear, C logU(1e-3,1e2), gamma logU(1e-4,1e-1); LightGBM: 15 randomized iterations. Do not combine with separate original GlycanML-paper protocols. Aggregation: Not reported GlycanGT published supplementary archive, Table S4; glycangt: Journal full-text XML · btag147_supplementary_data.zip / Table_S4.xlsx / Sheet1!A103:E103 (mean D103, SD E103) |
| Configuration: RGCN | Protocol: GlycanML phylum Accuracy: GlycanGT study: phylum Accuracy Dataset subset: SugarBase taxonomy phylum; GlycanML official motif split (GlycanML split) | 0.51323902792890796 ± 0.001662160208544 accuracy fraction · higher Uncertainty: type: standard_deviation; value: 0.001662160208544 Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · source checkedMethods, coverage and sourceRGCN on GlycanML phylum Accuracy: GlycanGT study: phylum Accuracy Taxonomy: 13,209 glycans total across eight levels, 4–1,737 classes per level. Official fixed GlycanML motif-based train/validation/test splits (8:1:1). Section 2.5 reports class-balanced classifiers, train ∪ validation hyperparameter selection by randomized search with 3-fold cross-validation, followed by one evaluation on the held-out test set; complete procedure repeated with three random seeds, reporting mean and standard deviation. GlycanGT large model pretrained with 35% masking provides [Graph] embeddings to SVM/LightGBM; the selected classifier for each S4 row is not identified. Section 2.6 states that all four graph baselines were trained and evaluated on the same datasets/splits; it does not establish that each baseline used the GlycanGT downstream classifier search. SVM search: 10 iterations, RBF/linear, C logU(1e-3,1e2), gamma logU(1e-4,1e-1); LightGBM: 15 randomized iterations. Do not combine with separate original GlycanML-paper protocols. Aggregation: Not reported GlycanGT published supplementary archive, Table S4; glycangt: Journal full-text XML · btag147_supplementary_data.zip / Table_S4.xlsx / Sheet1!A7:E7 (mean D7, SD E7) |
| Configuration: RGCN | Protocol: GlycanML phylum Macro-F1: GlycanGT study: phylum Macro-F1 Dataset subset: SugarBase taxonomy phylum; GlycanML official motif split (GlycanML split) | 0.090538684791431095 ± 0.0011545811568653001 macro_f1 dimensionless · higher Uncertainty: type: standard_deviation; value: 0.0011545811568653001 Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · source checkedMethods, coverage and sourceRGCN on GlycanML phylum Macro-F1: GlycanGT study: phylum Macro-F1 Taxonomy: 13,209 glycans total across eight levels, 4–1,737 classes per level. Official fixed GlycanML motif-based train/validation/test splits (8:1:1). Section 2.5 reports class-balanced classifiers, train ∪ validation hyperparameter selection by randomized search with 3-fold cross-validation, followed by one evaluation on the held-out test set; complete procedure repeated with three random seeds, reporting mean and standard deviation. GlycanGT large model pretrained with 35% masking provides [Graph] embeddings to SVM/LightGBM; the selected classifier for each S4 row is not identified. Section 2.6 states that all four graph baselines were trained and evaluated on the same datasets/splits; it does not establish that each baseline used the GlycanGT downstream classifier search. SVM search: 10 iterations, RBF/linear, C logU(1e-3,1e2), gamma logU(1e-4,1e-1); LightGBM: 15 randomized iterations. Do not combine with separate original GlycanML-paper protocols. Aggregation: Not reported GlycanGT published supplementary archive, Table S4; glycangt: Journal full-text XML · btag147_supplementary_data.zip / Table_S4.xlsx / Sheet1!A8:E8 (mean D8, SD E8) |
| Configuration: SweetNet | Protocol: GlycanML phylum Accuracy: GlycanGT study: phylum Accuracy Dataset subset: SugarBase taxonomy phylum; GlycanML official motif split (GlycanML split) | 0.82988755894087696 ± 0.019793636053162401 accuracy fraction · higher Uncertainty: type: standard_deviation; value: 0.019793636053162401 Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · source checkedMethods, coverage and sourceSweetNet on GlycanML phylum Accuracy: GlycanGT study: phylum Accuracy Taxonomy: 13,209 glycans total across eight levels, 4–1,737 classes per level. Official fixed GlycanML motif-based train/validation/test splits (8:1:1). Section 2.5 reports class-balanced classifiers, train ∪ validation hyperparameter selection by randomized search with 3-fold cross-validation, followed by one evaluation on the held-out test set; complete procedure repeated with three random seeds, reporting mean and standard deviation. GlycanGT large model pretrained with 35% masking provides [Graph] embeddings to SVM/LightGBM; the selected classifier for each S4 row is not identified. Section 2.6 states that all four graph baselines were trained and evaluated on the same datasets/splits; it does not establish that each baseline used the GlycanGT downstream classifier search. SVM search: 10 iterations, RBF/linear, C logU(1e-3,1e2), gamma logU(1e-4,1e-1); LightGBM: 15 randomized iterations. Do not combine with separate original GlycanML-paper protocols. Aggregation: Not reported GlycanGT published supplementary archive, Table S4; glycangt: Journal full-text XML · btag147_supplementary_data.zip / Table_S4.xlsx / Sheet1!A46:E46 (mean D46, SD E46) |
| Configuration: SweetNet | Protocol: GlycanML phylum Macro-F1: GlycanGT study: phylum Macro-F1 Dataset subset: SugarBase taxonomy phylum; GlycanML official motif split (GlycanML split) | 0.37055186011768798 ± 0.065574251495964103 macro_f1 dimensionless · higher Uncertainty: type: standard_deviation; value: 0.065574251495964103 Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · source checkedMethods, coverage and sourceSweetNet on GlycanML phylum Macro-F1: GlycanGT study: phylum Macro-F1 Taxonomy: 13,209 glycans total across eight levels, 4–1,737 classes per level. Official fixed GlycanML motif-based train/validation/test splits (8:1:1). Section 2.5 reports class-balanced classifiers, train ∪ validation hyperparameter selection by randomized search with 3-fold cross-validation, followed by one evaluation on the held-out test set; complete procedure repeated with three random seeds, reporting mean and standard deviation. GlycanGT large model pretrained with 35% masking provides [Graph] embeddings to SVM/LightGBM; the selected classifier for each S4 row is not identified. Section 2.6 states that all four graph baselines were trained and evaluated on the same datasets/splits; it does not establish that each baseline used the GlycanGT downstream classifier search. SVM search: 10 iterations, RBF/linear, C logU(1e-3,1e2), gamma logU(1e-4,1e-1); LightGBM: 15 randomized iterations. Do not combine with separate original GlycanML-paper protocols. Aggregation: Not reported GlycanGT published supplementary archive, Table S4; glycangt: Journal full-text XML · btag147_supplementary_data.zip / Table_S4.xlsx / Sheet1!A56:E56 (mean D56, SD E56) |
Source checking is not independent reproduction. Release 2026-09-23-2b89723c6dd9.
Subset and evaluation context
This record describes a particular subset or cohort used in an evaluation. Its results do not describe the full dataset.
Evidence
Source checking verifies the cited claim or transcription. It does not establish independent reproduction.
Evidence table
Inspect claims, sources and review details
Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.
One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.
4 evidence rows matching the loaded filters
| Property and statement | Original source and location | Review and provenance |
|---|---|---|
| description Dataset subset reported in GlycanGT published supplementary archive, Table S4. Exact split manifest remains unextracted; source-table identity is retained. Context-only references | glycangt: Journal full-text XML No field-specific location recorded Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: Primary article XML snapshot | not individually reviewed No individual claim review recorded Audit detailsField: Source artifact SHA-256: Hash scope: SHA-256 of retrieved original artifact bytes Format: original_artifact |
| description Dataset subset reported in GlycanGT published supplementary archive, Table S4. Exact split manifest remains unextracted; source-table identity is retained. Context-only references | GlycanGT published supplementary archive, Table S4 No field-specific location recorded Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: Published Bioinformatics btag147 supplementary archive, retrieved 2026-09-23; Table_S4.xlsx SHA-256 d7c35909bac6bcb78ca8fdb32c0463f05e692f15861a8184e65e415e7216f493 | not individually reviewed No individual claim review recorded Audit detailsField: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| name SugarBase taxonomy phylum; GlycanML official motif split (GlycanML split) Context-only references | glycangt: Journal full-text XML No field-specific location recorded Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: Primary article XML snapshot | not individually reviewed No individual claim review recorded Audit detailsField: Source artifact SHA-256: Hash scope: SHA-256 of retrieved original artifact bytes Format: original_artifact |
| name SugarBase taxonomy phylum; GlycanML official motif split (GlycanML split) Context-only references | GlycanGT published supplementary archive, Table S4 No field-specific location recorded Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: Published Bioinformatics btag147 supplementary archive, retrieved 2026-09-23; Table_S4.xlsx SHA-256 d7c35909bac6bcb78ca8fdb32c0463f05e692f15861a8184e65e415e7216f493 | not individually reviewed No individual claim review recorded Audit detailsField: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
Sources and history
View linked audit checks and correction history
Release 2026-09-23-2b89723c6dd9 · Record review: source checked
2 source records and release history
- GlycanGT published supplementary archive, Table S4 · Original source · Published Bioinformatics btag147 supplementary archive, retrieved 2026-09-23; Table_S4.xlsx SHA-256 d7c35909bac6bcb78ca8fdb32c0463f05e692f15861a8184e65e415e7216f493
- glycangt: Journal full-text XML · Original source · Primary article XML snapshot
Technical metadata and extraction receipts
Stable ID: glycangt-2026-table-s4-dataset-sugarbase-taxonomy-phylum-glycanml-official-motif-split
- areas
- glycans
- missing metadata
- version: unreported; url: unextracted
Related records
- dataset: GlycanAA on GlycanML phylum Accuracy: GlycanGT study: phylum Accuracy
- dataset: GlycanAA on GlycanML phylum Macro-F1: GlycanGT study: phylum Macro-F1
- dataset: GlycanGT on GlycanML phylum Accuracy: GlycanGT study: phylum Accuracy
- dataset: GlycanGT on GlycanML phylum Macro-F1: GlycanGT study: phylum Macro-F1
- dataset: Graphormer on GlycanML phylum Accuracy: GlycanGT study: phylum Accuracy
- dataset: Graphormer on GlycanML phylum Macro-F1: GlycanGT study: phylum Macro-F1
- dataset: RGCN on GlycanML phylum Accuracy: GlycanGT study: phylum Accuracy
- dataset: RGCN on GlycanML phylum Macro-F1: GlycanGT study: phylum Macro-F1
- dataset: SweetNet on GlycanML phylum Accuracy: GlycanGT study: phylum Accuracy
- dataset: SweetNet on GlycanML phylum Macro-F1: GlycanGT study: phylum Macro-F1