rewire.it
Dataset subset

SugarBase taxonomy species; GlycanML official motif split (GlycanML split)

Dataset subset reported in GlycanGT published supplementary archive, Table S4. Exact split manifest remains unextracted; source-table identity is retained.

Evaluation results

10 evaluations · 10 metric rows. Different protocols are not a single leaderboard.

Filter evaluations

Applied filters: All linked evaluations

Exact evaluated configurations and original reported results
Tested configurationProtocol and datasetFindingEvidence and details
Configuration: GlycanAAProtocol: GlycanML species Accuracy: GlycanGT study: species Accuracy
Dataset subset: SugarBase taxonomy species; GlycanML official motif split (GlycanML split)
0.40549999999999897 ± 0.0054836119483420804 accuracy
fraction · higher

Uncertainty: type: standard_deviation; value: 0.0054836119483420804

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · source checked
Methods, coverage and source

GlycanAA on GlycanML species Accuracy: GlycanGT study: species Accuracy

Taxonomy: 13,209 glycans total across eight levels, 4–1,737 classes per level. Official fixed GlycanML motif-based train/validation/test splits (8:1:1). Section 2.5 reports class-balanced classifiers, train ∪ validation hyperparameter selection by randomized search with 3-fold cross-validation, followed by one evaluation on the held-out test set; complete procedure repeated with three random seeds, reporting mean and standard deviation. GlycanGT large model pretrained with 35% masking provides [Graph] embeddings to SVM/LightGBM; the selected classifier for each S4 row is not identified. Section 2.6 states that all four graph baselines were trained and evaluated on the same datasets/splits; it does not establish that each baseline used the GlycanGT downstream classifier search. SVM search: 10 iterations, RBF/linear, C logU(1e-3,1e2), gamma logU(1e-4,1e-1); LightGBM: 15 randomized iterations. Do not combine with separate original GlycanML-paper protocols.

Aggregation: Not reported

GlycanGT published supplementary archive, Table S4; glycangt: Journal full-text XML · btag147_supplementary_data.zip / Table_S4.xlsx / Sheet1!A32:E32 (mean D32, SD E32)
Configuration: GlycanAAProtocol: GlycanML species Macro-F1: GlycanGT study: species Macro-F1
Dataset subset: SugarBase taxonomy species; GlycanML official motif split (GlycanML split)
0.15840000000000001 ± 0.012304064369142401 macro_f1
dimensionless · higher

Uncertainty: type: standard_deviation; value: 0.012304064369142401

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · source checked
Methods, coverage and source

GlycanAA on GlycanML species Macro-F1: GlycanGT study: species Macro-F1

Taxonomy: 13,209 glycans total across eight levels, 4–1,737 classes per level. Official fixed GlycanML motif-based train/validation/test splits (8:1:1). Section 2.5 reports class-balanced classifiers, train ∪ validation hyperparameter selection by randomized search with 3-fold cross-validation, followed by one evaluation on the held-out test set; complete procedure repeated with three random seeds, reporting mean and standard deviation. GlycanGT large model pretrained with 35% masking provides [Graph] embeddings to SVM/LightGBM; the selected classifier for each S4 row is not identified. Section 2.6 states that all four graph baselines were trained and evaluated on the same datasets/splits; it does not establish that each baseline used the GlycanGT downstream classifier search. SVM search: 10 iterations, RBF/linear, C logU(1e-3,1e2), gamma logU(1e-4,1e-1); LightGBM: 15 randomized iterations. Do not combine with separate original GlycanML-paper protocols.

Aggregation: Not reported

GlycanGT published supplementary archive, Table S4; glycangt: Journal full-text XML · btag147_supplementary_data.zip / Table_S4.xlsx / Sheet1!A33:E33 (mean D33, SD E33)
Pipeline: GlycanGTProtocol: GlycanML species Accuracy: GlycanGT study: species Accuracy
Dataset subset: SugarBase taxonomy species; GlycanML official motif split (GlycanML split)
0.39173014145810597 ± 0.021681021594310401 accuracy
fraction · higher

Uncertainty: type: standard_deviation; value: 0.021681021594310401

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · source checked
Methods, coverage and source

GlycanGT on GlycanML species Accuracy: GlycanGT study: species Accuracy

Taxonomy: 13,209 glycans total across eight levels, 4–1,737 classes per level. Official fixed GlycanML motif-based train/validation/test splits (8:1:1). Section 2.5 reports class-balanced classifiers, train ∪ validation hyperparameter selection by randomized search with 3-fold cross-validation, followed by one evaluation on the held-out test set; complete procedure repeated with three random seeds, reporting mean and standard deviation. GlycanGT large model pretrained with 35% masking provides [Graph] embeddings to SVM/LightGBM; the selected classifier for each S4 row is not identified. Section 2.6 states that all four graph baselines were trained and evaluated on the same datasets/splits; it does not establish that each baseline used the GlycanGT downstream classifier search. SVM search: 10 iterations, RBF/linear, C logU(1e-3,1e2), gamma logU(1e-4,1e-1); LightGBM: 15 randomized iterations. Do not combine with separate original GlycanML-paper protocols.

Aggregation: Not reported

GlycanGT published supplementary archive, Table S4; glycangt: Journal full-text XML · btag147_supplementary_data.zip / Table_S4.xlsx / Sheet1!A79:E79 (mean D79, SD E79)
Pipeline: GlycanGTProtocol: GlycanML species Macro-F1: GlycanGT study: species Macro-F1
Dataset subset: SugarBase taxonomy species; GlycanML official motif split (GlycanML split)
0.184571928631623 ± 0.00080738146274129995 macro_f1
dimensionless · higher

Uncertainty: type: standard_deviation; value: 0.00080738146274129995

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · source checked
Methods, coverage and source

GlycanGT on GlycanML species Macro-F1: GlycanGT study: species Macro-F1

Taxonomy: 13,209 glycans total across eight levels, 4–1,737 classes per level. Official fixed GlycanML motif-based train/validation/test splits (8:1:1). Section 2.5 reports class-balanced classifiers, train ∪ validation hyperparameter selection by randomized search with 3-fold cross-validation, followed by one evaluation on the held-out test set; complete procedure repeated with three random seeds, reporting mean and standard deviation. GlycanGT large model pretrained with 35% masking provides [Graph] embeddings to SVM/LightGBM; the selected classifier for each S4 row is not identified. Section 2.6 states that all four graph baselines were trained and evaluated on the same datasets/splits; it does not establish that each baseline used the GlycanGT downstream classifier search. SVM search: 10 iterations, RBF/linear, C logU(1e-3,1e2), gamma logU(1e-4,1e-1); LightGBM: 15 randomized iterations. Do not combine with separate original GlycanML-paper protocols.

Aggregation: Not reported

GlycanGT published supplementary archive, Table S4; glycangt: Journal full-text XML · btag147_supplementary_data.zip / Table_S4.xlsx / Sheet1!A80:E80 (mean D80, SD E80)
Configuration: GraphormerProtocol: GlycanML species Accuracy: GlycanGT study: species Accuracy
Dataset subset: SugarBase taxonomy species; GlycanML official motif split (GlycanML split)
0.35691000000000001 ± 0.019796999999999999 accuracy
fraction · higher

Uncertainty: type: standard_deviation; value: 0.019796999999999999

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · source checked
Methods, coverage and source

Graphormer on GlycanML species Accuracy: GlycanGT study: species Accuracy

Taxonomy: 13,209 glycans total across eight levels, 4–1,737 classes per level. Official fixed GlycanML motif-based train/validation/test splits (8:1:1). Section 2.5 reports class-balanced classifiers, train ∪ validation hyperparameter selection by randomized search with 3-fold cross-validation, followed by one evaluation on the held-out test set; complete procedure repeated with three random seeds, reporting mean and standard deviation. GlycanGT large model pretrained with 35% masking provides [Graph] embeddings to SVM/LightGBM; the selected classifier for each S4 row is not identified. Section 2.6 states that all four graph baselines were trained and evaluated on the same datasets/splits; it does not establish that each baseline used the GlycanGT downstream classifier search. SVM search: 10 iterations, RBF/linear, C logU(1e-3,1e2), gamma logU(1e-4,1e-1); LightGBM: 15 randomized iterations. Do not combine with separate original GlycanML-paper protocols.

Aggregation: Not reported

GlycanGT published supplementary archive, Table S4; glycangt: Journal full-text XML · btag147_supplementary_data.zip / Table_S4.xlsx / Sheet1!A104:E104 (mean D104, SD E104)
Configuration: GraphormerProtocol: GlycanML species Macro-F1: GlycanGT study: species Macro-F1
Dataset subset: SugarBase taxonomy species; GlycanML official motif split (GlycanML split)
0.122909 ± 0.0062329999999999998 macro_f1
dimensionless · higher

Uncertainty: type: standard_deviation; value: 0.0062329999999999998

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · source checked
Methods, coverage and source

Graphormer on GlycanML species Macro-F1: GlycanGT study: species Macro-F1

Taxonomy: 13,209 glycans total across eight levels, 4–1,737 classes per level. Official fixed GlycanML motif-based train/validation/test splits (8:1:1). Section 2.5 reports class-balanced classifiers, train ∪ validation hyperparameter selection by randomized search with 3-fold cross-validation, followed by one evaluation on the held-out test set; complete procedure repeated with three random seeds, reporting mean and standard deviation. GlycanGT large model pretrained with 35% masking provides [Graph] embeddings to SVM/LightGBM; the selected classifier for each S4 row is not identified. Section 2.6 states that all four graph baselines were trained and evaluated on the same datasets/splits; it does not establish that each baseline used the GlycanGT downstream classifier search. SVM search: 10 iterations, RBF/linear, C logU(1e-3,1e2), gamma logU(1e-4,1e-1); LightGBM: 15 randomized iterations. Do not combine with separate original GlycanML-paper protocols.

Aggregation: Not reported

GlycanGT published supplementary archive, Table S4; glycangt: Journal full-text XML · btag147_supplementary_data.zip / Table_S4.xlsx / Sheet1!A105:E105 (mean D105, SD E105)
Configuration: RGCNProtocol: GlycanML species Accuracy: GlycanGT study: species Accuracy
Dataset subset: SugarBase taxonomy species; GlycanML official motif split (GlycanML split)
0.0039898440333695998 ± 0.0069106125800716001 accuracy
fraction · higher

Uncertainty: type: standard_deviation; value: 0.0069106125800716001

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · source checked
Methods, coverage and source

RGCN on GlycanML species Accuracy: GlycanGT study: species Accuracy

Taxonomy: 13,209 glycans total across eight levels, 4–1,737 classes per level. Official fixed GlycanML motif-based train/validation/test splits (8:1:1). Section 2.5 reports class-balanced classifiers, train ∪ validation hyperparameter selection by randomized search with 3-fold cross-validation, followed by one evaluation on the held-out test set; complete procedure repeated with three random seeds, reporting mean and standard deviation. GlycanGT large model pretrained with 35% masking provides [Graph] embeddings to SVM/LightGBM; the selected classifier for each S4 row is not identified. Section 2.6 states that all four graph baselines were trained and evaluated on the same datasets/splits; it does not establish that each baseline used the GlycanGT downstream classifier search. SVM search: 10 iterations, RBF/linear, C logU(1e-3,1e2), gamma logU(1e-4,1e-1); LightGBM: 15 randomized iterations. Do not combine with separate original GlycanML-paper protocols.

Aggregation: Not reported

GlycanGT published supplementary archive, Table S4; glycangt: Journal full-text XML · btag147_supplementary_data.zip / Table_S4.xlsx / Sheet1!A17:E17 (mean D17, SD E17)
Configuration: RGCNProtocol: GlycanML species Macro-F1: GlycanGT study: species Macro-F1
Dataset subset: SugarBase taxonomy species; GlycanML official motif split (GlycanML split)
0.00069761542364280005 ± 0.001208305357893 macro_f1
dimensionless · higher

Uncertainty: type: standard_deviation; value: 0.001208305357893

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · source checked
Methods, coverage and source

RGCN on GlycanML species Macro-F1: GlycanGT study: species Macro-F1

Taxonomy: 13,209 glycans total across eight levels, 4–1,737 classes per level. Official fixed GlycanML motif-based train/validation/test splits (8:1:1). Section 2.5 reports class-balanced classifiers, train ∪ validation hyperparameter selection by randomized search with 3-fold cross-validation, followed by one evaluation on the held-out test set; complete procedure repeated with three random seeds, reporting mean and standard deviation. GlycanGT large model pretrained with 35% masking provides [Graph] embeddings to SVM/LightGBM; the selected classifier for each S4 row is not identified. Section 2.6 states that all four graph baselines were trained and evaluated on the same datasets/splits; it does not establish that each baseline used the GlycanGT downstream classifier search. SVM search: 10 iterations, RBF/linear, C logU(1e-3,1e2), gamma logU(1e-4,1e-1); LightGBM: 15 randomized iterations. Do not combine with separate original GlycanML-paper protocols.

Aggregation: Not reported

GlycanGT published supplementary archive, Table S4; glycangt: Journal full-text XML · btag147_supplementary_data.zip / Table_S4.xlsx / Sheet1!A18:E18 (mean D18, SD E18)
Configuration: SweetNetProtocol: GlycanML species Accuracy: GlycanGT study: species Accuracy
Dataset subset: SugarBase taxonomy species; GlycanML official motif split (GlycanML split)
0.35545883206383699 ± 0.025755183860829499 accuracy
fraction · higher

Uncertainty: type: standard_deviation; value: 0.025755183860829499

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · source checked
Methods, coverage and source

SweetNet on GlycanML species Accuracy: GlycanGT study: species Accuracy

Taxonomy: 13,209 glycans total across eight levels, 4–1,737 classes per level. Official fixed GlycanML motif-based train/validation/test splits (8:1:1). Section 2.5 reports class-balanced classifiers, train ∪ validation hyperparameter selection by randomized search with 3-fold cross-validation, followed by one evaluation on the held-out test set; complete procedure repeated with three random seeds, reporting mean and standard deviation. GlycanGT large model pretrained with 35% masking provides [Graph] embeddings to SVM/LightGBM; the selected classifier for each S4 row is not identified. Section 2.6 states that all four graph baselines were trained and evaluated on the same datasets/splits; it does not establish that each baseline used the GlycanGT downstream classifier search. SVM search: 10 iterations, RBF/linear, C logU(1e-3,1e2), gamma logU(1e-4,1e-1); LightGBM: 15 randomized iterations. Do not combine with separate original GlycanML-paper protocols.

Aggregation: Not reported

GlycanGT published supplementary archive, Table S4; glycangt: Journal full-text XML · btag147_supplementary_data.zip / Table_S4.xlsx / Sheet1!A51:E51 (mean D51, SD E51)
Configuration: SweetNetProtocol: GlycanML species Macro-F1: GlycanGT study: species Macro-F1
Dataset subset: SugarBase taxonomy species; GlycanML official motif split (GlycanML split)
0.11310443922558699 ± 0.0110807900341276 macro_f1
dimensionless · higher

Uncertainty: type: standard_deviation; value: 0.0110807900341276

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · source checked
Methods, coverage and source

SweetNet on GlycanML species Macro-F1: GlycanGT study: species Macro-F1

Taxonomy: 13,209 glycans total across eight levels, 4–1,737 classes per level. Official fixed GlycanML motif-based train/validation/test splits (8:1:1). Section 2.5 reports class-balanced classifiers, train ∪ validation hyperparameter selection by randomized search with 3-fold cross-validation, followed by one evaluation on the held-out test set; complete procedure repeated with three random seeds, reporting mean and standard deviation. GlycanGT large model pretrained with 35% masking provides [Graph] embeddings to SVM/LightGBM; the selected classifier for each S4 row is not identified. Section 2.6 states that all four graph baselines were trained and evaluated on the same datasets/splits; it does not establish that each baseline used the GlycanGT downstream classifier search. SVM search: 10 iterations, RBF/linear, C logU(1e-3,1e2), gamma logU(1e-4,1e-1); LightGBM: 15 randomized iterations. Do not combine with separate original GlycanML-paper protocols.

Aggregation: Not reported

GlycanGT published supplementary archive, Table S4; glycangt: Journal full-text XML · btag147_supplementary_data.zip / Table_S4.xlsx / Sheet1!A61:E61 (mean D61, SD E61)

Source checking is not independent reproduction. Release 2026-09-23-2b89723c6dd9.

Subset and evaluation context

This record describes a particular subset or cohort used in an evaluation. Its results do not describe the full dataset.

Evidence

Source checking verifies the cited claim or transcription. It does not establish independent reproduction.

Evidence table

Inspect claims, sources and review details

Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.

One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.

4 evidence rows matching the loaded filters

Claims, original sources and review scope · Release 2026-09-23-2b89723c6dd9
Property and statementOriginal source and locationReview and provenance
description
Dataset subset reported in GlycanGT published supplementary archive, Table S4. Exact split manifest remains unextracted; source-table identity is retained.
Context-only references
glycangt: Journal full-text XML

Original source ↗

No field-specific location recorded

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: Primary article XML snapshot
Retrieved: 2026-09-16T20:20:56.439056+00:00

not individually reviewed

No individual claim review recorded

Audit details

Field: description

Source artifact SHA-256: 53e89a636c868c0329ee7eb6ae92f1028ec891940bc61730a148981b647fbbe5

Hash scope: SHA-256 of retrieved original artifact bytes

Format: original_artifact

Inspected artifact

description
Dataset subset reported in GlycanGT published supplementary archive, Table S4. Exact split manifest remains unextracted; source-table identity is retained.
Context-only references
GlycanGT published supplementary archive, Table S4

Original source ↗

No field-specific location recorded

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: Published Bioinformatics btag147 supplementary archive, retrieved 2026-09-23; Table_S4.xlsx SHA-256 d7c35909bac6bcb78ca8fdb32c0463f05e692f15861a8184e65e415e7216f493
Retrieved: 2026-09-23T11:31:33.520570+00:00

not individually reviewed

No individual claim review recorded

Audit details

Field: description

Source artifact SHA-256: 27748c6c0c0bb07b0105d274fb745fd4dfe702d34b9ee367d1ddacbc71c56ab0

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

name
SugarBase taxonomy species; GlycanML official motif split (GlycanML split)
Context-only references
glycangt: Journal full-text XML

Original source ↗

No field-specific location recorded

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: Primary article XML snapshot
Retrieved: 2026-09-16T20:20:56.439056+00:00

not individually reviewed

No individual claim review recorded

Audit details

Field: name

Source artifact SHA-256: 53e89a636c868c0329ee7eb6ae92f1028ec891940bc61730a148981b647fbbe5

Hash scope: SHA-256 of retrieved original artifact bytes

Format: original_artifact

Inspected artifact

name
SugarBase taxonomy species; GlycanML official motif split (GlycanML split)
Context-only references
GlycanGT published supplementary archive, Table S4

Original source ↗

No field-specific location recorded

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: Published Bioinformatics btag147 supplementary archive, retrieved 2026-09-23; Table_S4.xlsx SHA-256 d7c35909bac6bcb78ca8fdb32c0463f05e692f15861a8184e65e415e7216f493
Retrieved: 2026-09-23T11:31:33.520570+00:00

not individually reviewed

No individual claim review recorded

Audit details

Field: name

Source artifact SHA-256: 27748c6c0c0bb07b0105d274fb745fd4dfe702d34b9ee367d1ddacbc71c56ab0

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Sources and history

View linked audit checks and correction history

Release 2026-09-23-2b89723c6dd9 · Record review: source checked

2 source records and release historyDownload this release
Technical metadata and extraction receipts

Stable ID: glycangt-2026-table-s4-dataset-sugarbase-taxonomy-species-glycanml-official-motif-split

areas
glycans
missing metadata
version: unreported; url: unextracted
Related records

Suggest a correction