rewire.it
Pipeline

GlycanGT

Large GlycanGT encoder, 35% node/edge masking pretraining; [Graph] embeddings supplied to a trained SVM or LightGBM classifier. Table S4 does not identify the chosen classifier per task. This is a downstream prediction pipeline, not the bare encoder.

21 evaluations · 21 metric rows

Overview

Large GlycanGT encoder, 35% node/edge masking pretraining; [Graph] embeddings supplied to a trained SVM or LightGBM classifier. Table S4 does not identify the chosen classifier per task. This is a downstream prediction pipeline, not the bare encoder.

Consult the linked sources for architecture or protocol details. Missing evidence is not evidence of a missing capability.

Evaluations and results

21 evaluations · 21 metric rows. Different protocols are not a single leaderboard.

Filter evaluations

Applied filters: All linked evaluations

Exact evaluated configurations and original reported results
Tested configurationProtocol and datasetFindingEvidence and details
Pipeline: GlycanGTProtocol: GlycanML class Accuracy: GlycanGT study: class Accuracy
Dataset subset: SugarBase taxonomy class; GlycanML official motif split (GlycanML split)
0.69350743561842498 ± 0.0072451883770178003 accuracy
fraction · higher

Uncertainty: type: standard_deviation; value: 0.0072451883770178003

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · source checked
Methods, coverage and source

GlycanGT on GlycanML class Accuracy: GlycanGT study: class Accuracy

Taxonomy: 13,209 glycans total across eight levels, 4–1,737 classes per level. Official fixed GlycanML motif-based train/validation/test splits (8:1:1). Section 2.5 reports class-balanced classifiers, train ∪ validation hyperparameter selection by randomized search with 3-fold cross-validation, followed by one evaluation on the held-out test set; complete procedure repeated with three random seeds, reporting mean and standard deviation. GlycanGT large model pretrained with 35% masking provides [Graph] embeddings to SVM/LightGBM; the selected classifier for each S4 row is not identified. Section 2.6 states that all four graph baselines were trained and evaluated on the same datasets/splits; it does not establish that each baseline used the GlycanGT downstream classifier search. SVM search: 10 iterations, RBF/linear, C logU(1e-3,1e2), gamma logU(1e-4,1e-1); LightGBM: 15 randomized iterations. Do not combine with separate original GlycanML-paper protocols.

Aggregation: Not reported

GlycanGT published supplementary archive, Table S4; glycangt: Journal full-text XML · btag147_supplementary_data.zip / Table_S4.xlsx / Sheet1!A71:E71 (mean D71, SD E71)
Pipeline: GlycanGTProtocol: GlycanML class Macro-F1: GlycanGT study: class Macro-F1
Dataset subset: SugarBase taxonomy class; GlycanML official motif split (GlycanML split)
0.39715349769695701 ± 0.0084870172846663004 macro_f1
dimensionless · higher

Uncertainty: type: standard_deviation; value: 0.0084870172846663004

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · source checked
Methods, coverage and source

GlycanGT on GlycanML class Macro-F1: GlycanGT study: class Macro-F1

Taxonomy: 13,209 glycans total across eight levels, 4–1,737 classes per level. Official fixed GlycanML motif-based train/validation/test splits (8:1:1). Section 2.5 reports class-balanced classifiers, train ∪ validation hyperparameter selection by randomized search with 3-fold cross-validation, followed by one evaluation on the held-out test set; complete procedure repeated with three random seeds, reporting mean and standard deviation. GlycanGT large model pretrained with 35% masking provides [Graph] embeddings to SVM/LightGBM; the selected classifier for each S4 row is not identified. Section 2.6 states that all four graph baselines were trained and evaluated on the same datasets/splits; it does not establish that each baseline used the GlycanGT downstream classifier search. SVM search: 10 iterations, RBF/linear, C logU(1e-3,1e2), gamma logU(1e-4,1e-1); LightGBM: 15 randomized iterations. Do not combine with separate original GlycanML-paper protocols.

Aggregation: Not reported

GlycanGT published supplementary archive, Table S4; glycangt: Journal full-text XML · btag147_supplementary_data.zip / Table_S4.xlsx / Sheet1!A72:E72 (mean D72, SD E72)
Pipeline: GlycanGTProtocol: GlycanML domain Accuracy: GlycanGT study: domain Accuracy
Dataset subset: SugarBase taxonomy domain; GlycanML official motif split (GlycanML split)
0.92020311933260801 ± 0.0077199117351442002 accuracy
fraction · higher

Uncertainty: type: standard_deviation; value: 0.0077199117351442002

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · source checked
Methods, coverage and source

GlycanGT on GlycanML domain Accuracy: GlycanGT study: domain Accuracy

Taxonomy: 13,209 glycans total across eight levels, 4–1,737 classes per level. Official fixed GlycanML motif-based train/validation/test splits (8:1:1). Section 2.5 reports class-balanced classifiers, train ∪ validation hyperparameter selection by randomized search with 3-fold cross-validation, followed by one evaluation on the held-out test set; complete procedure repeated with three random seeds, reporting mean and standard deviation. GlycanGT large model pretrained with 35% masking provides [Graph] embeddings to SVM/LightGBM; the selected classifier for each S4 row is not identified. Section 2.6 states that all four graph baselines were trained and evaluated on the same datasets/splits; it does not establish that each baseline used the GlycanGT downstream classifier search. SVM search: 10 iterations, RBF/linear, C logU(1e-3,1e2), gamma logU(1e-4,1e-1); LightGBM: 15 randomized iterations. Do not combine with separate original GlycanML-paper protocols.

Aggregation: Not reported

GlycanGT published supplementary archive, Table S4; glycangt: Journal full-text XML · btag147_supplementary_data.zip / Table_S4.xlsx / Sheet1!A65:E65 (mean D65, SD E65)
Pipeline: GlycanGTProtocol: GlycanML domain Macro-F1: GlycanGT study: domain Macro-F1
Dataset subset: SugarBase taxonomy domain; GlycanML official motif split (GlycanML split)
0.73873655202321897 ± 0.023629706943064301 macro_f1
dimensionless · higher

Uncertainty: type: standard_deviation; value: 0.023629706943064301

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · source checked
Methods, coverage and source

GlycanGT on GlycanML domain Macro-F1: GlycanGT study: domain Macro-F1

Taxonomy: 13,209 glycans total across eight levels, 4–1,737 classes per level. Official fixed GlycanML motif-based train/validation/test splits (8:1:1). Section 2.5 reports class-balanced classifiers, train ∪ validation hyperparameter selection by randomized search with 3-fold cross-validation, followed by one evaluation on the held-out test set; complete procedure repeated with three random seeds, reporting mean and standard deviation. GlycanGT large model pretrained with 35% masking provides [Graph] embeddings to SVM/LightGBM; the selected classifier for each S4 row is not identified. Section 2.6 states that all four graph baselines were trained and evaluated on the same datasets/splits; it does not establish that each baseline used the GlycanGT downstream classifier search. SVM search: 10 iterations, RBF/linear, C logU(1e-3,1e2), gamma logU(1e-4,1e-1); LightGBM: 15 randomized iterations. Do not combine with separate original GlycanML-paper protocols.

Aggregation: Not reported

GlycanGT published supplementary archive, Table S4; glycangt: Journal full-text XML · btag147_supplementary_data.zip / Table_S4.xlsx / Sheet1!A66:E66 (mean D66, SD E66)
Pipeline: GlycanGTProtocol: GlycanML family Accuracy: GlycanGT study: family Accuracy
Dataset subset: SugarBase taxonomy family; GlycanML official motif split (GlycanML split)
0.463184620964816 ± 0.0078717934037760007 accuracy
fraction · higher

Uncertainty: type: standard_deviation; value: 0.0078717934037760007

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · source checked
Methods, coverage and source

GlycanGT on GlycanML family Accuracy: GlycanGT study: family Accuracy

Taxonomy: 13,209 glycans total across eight levels, 4–1,737 classes per level. Official fixed GlycanML motif-based train/validation/test splits (8:1:1). Section 2.5 reports class-balanced classifiers, train ∪ validation hyperparameter selection by randomized search with 3-fold cross-validation, followed by one evaluation on the held-out test set; complete procedure repeated with three random seeds, reporting mean and standard deviation. GlycanGT large model pretrained with 35% masking provides [Graph] embeddings to SVM/LightGBM; the selected classifier for each S4 row is not identified. Section 2.6 states that all four graph baselines were trained and evaluated on the same datasets/splits; it does not establish that each baseline used the GlycanGT downstream classifier search. SVM search: 10 iterations, RBF/linear, C logU(1e-3,1e2), gamma logU(1e-4,1e-1); LightGBM: 15 randomized iterations. Do not combine with separate original GlycanML-paper protocols.

Aggregation: Not reported

GlycanGT published supplementary archive, Table S4; glycangt: Journal full-text XML · btag147_supplementary_data.zip / Table_S4.xlsx / Sheet1!A75:E75 (mean D75, SD E75)
Pipeline: GlycanGTProtocol: GlycanML family Macro-F1: GlycanGT study: family Macro-F1
Dataset subset: SugarBase taxonomy family; GlycanML official motif split (GlycanML split)
0.24638700174167499 ± 0.004794223931708 macro_f1
dimensionless · higher

Uncertainty: type: standard_deviation; value: 0.004794223931708

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · source checked
Methods, coverage and source

GlycanGT on GlycanML family Macro-F1: GlycanGT study: family Macro-F1

Taxonomy: 13,209 glycans total across eight levels, 4–1,737 classes per level. Official fixed GlycanML motif-based train/validation/test splits (8:1:1). Section 2.5 reports class-balanced classifiers, train ∪ validation hyperparameter selection by randomized search with 3-fold cross-validation, followed by one evaluation on the held-out test set; complete procedure repeated with three random seeds, reporting mean and standard deviation. GlycanGT large model pretrained with 35% masking provides [Graph] embeddings to SVM/LightGBM; the selected classifier for each S4 row is not identified. Section 2.6 states that all four graph baselines were trained and evaluated on the same datasets/splits; it does not establish that each baseline used the GlycanGT downstream classifier search. SVM search: 10 iterations, RBF/linear, C logU(1e-3,1e2), gamma logU(1e-4,1e-1); LightGBM: 15 randomized iterations. Do not combine with separate original GlycanML-paper protocols.

Aggregation: Not reported

GlycanGT published supplementary archive, Table S4; glycangt: Journal full-text XML · btag147_supplementary_data.zip / Table_S4.xlsx / Sheet1!A76:E76 (mean D76, SD E76)
Pipeline: GlycanGTProtocol: GlycanML genus Accuracy: GlycanGT study: genus Accuracy
Dataset subset: SugarBase taxonomy genus; GlycanML official motif split (GlycanML split)
0.44577439245556699 ± 0.0045302850913300002 accuracy
fraction · higher

Uncertainty: type: standard_deviation; value: 0.0045302850913300002

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · source checked
Methods, coverage and source

GlycanGT on GlycanML genus Accuracy: GlycanGT study: genus Accuracy

Taxonomy: 13,209 glycans total across eight levels, 4–1,737 classes per level. Official fixed GlycanML motif-based train/validation/test splits (8:1:1). Section 2.5 reports class-balanced classifiers, train ∪ validation hyperparameter selection by randomized search with 3-fold cross-validation, followed by one evaluation on the held-out test set; complete procedure repeated with three random seeds, reporting mean and standard deviation. GlycanGT large model pretrained with 35% masking provides [Graph] embeddings to SVM/LightGBM; the selected classifier for each S4 row is not identified. Section 2.6 states that all four graph baselines were trained and evaluated on the same datasets/splits; it does not establish that each baseline used the GlycanGT downstream classifier search. SVM search: 10 iterations, RBF/linear, C logU(1e-3,1e2), gamma logU(1e-4,1e-1); LightGBM: 15 randomized iterations. Do not combine with separate original GlycanML-paper protocols.

Aggregation: Not reported

GlycanGT published supplementary archive, Table S4; glycangt: Journal full-text XML · btag147_supplementary_data.zip / Table_S4.xlsx / Sheet1!A77:E77 (mean D77, SD E77)
Pipeline: GlycanGTProtocol: GlycanML genus Macro-F1: GlycanGT study: genus Macro-F1
Dataset subset: SugarBase taxonomy genus; GlycanML official motif split (GlycanML split)
0.188457844899896 ± 0.0007473331496721 macro_f1
dimensionless · higher

Uncertainty: type: standard_deviation; value: 0.0007473331496721

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · source checked
Methods, coverage and source

GlycanGT on GlycanML genus Macro-F1: GlycanGT study: genus Macro-F1

Taxonomy: 13,209 glycans total across eight levels, 4–1,737 classes per level. Official fixed GlycanML motif-based train/validation/test splits (8:1:1). Section 2.5 reports class-balanced classifiers, train ∪ validation hyperparameter selection by randomized search with 3-fold cross-validation, followed by one evaluation on the held-out test set; complete procedure repeated with three random seeds, reporting mean and standard deviation. GlycanGT large model pretrained with 35% masking provides [Graph] embeddings to SVM/LightGBM; the selected classifier for each S4 row is not identified. Section 2.6 states that all four graph baselines were trained and evaluated on the same datasets/splits; it does not establish that each baseline used the GlycanGT downstream classifier search. SVM search: 10 iterations, RBF/linear, C logU(1e-3,1e2), gamma logU(1e-4,1e-1); LightGBM: 15 randomized iterations. Do not combine with separate original GlycanML-paper protocols.

Aggregation: Not reported

GlycanGT published supplementary archive, Table S4; glycangt: Journal full-text XML · btag147_supplementary_data.zip / Table_S4.xlsx / Sheet1!A78:E78 (mean D78, SD E78)
Pipeline: GlycanGTProtocol: GlycanML glycosylation Accuracy: GlycanGT study: glycosylation Accuracy
Dataset subset: GlyConnect glycosylation; GlycanML official motif split (GlycanML split)
0.98295454545454497 ± 0 accuracy
fraction · higher

Uncertainty: type: standard_deviation; value: 0

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · source checked
Methods, coverage and source

GlycanGT on GlycanML glycosylation Accuracy: GlycanGT study: glycosylation Accuracy

Glycosylation: 1,683 glycans total; N-linked/O-linked/free. Official fixed GlycanML motif-based train/validation/test splits (8:1:1). Section 2.5 reports class-balanced classifiers, train ∪ validation hyperparameter selection by randomized search with 3-fold cross-validation, followed by one evaluation on the held-out test set; complete procedure repeated with three random seeds, reporting mean and standard deviation. GlycanGT large model pretrained with 35% masking provides [Graph] embeddings to SVM/LightGBM; the selected classifier for each S4 row is not identified. Section 2.6 states that all four graph baselines were trained and evaluated on the same datasets/splits; it does not establish that each baseline used the GlycanGT downstream classifier search. SVM search: 10 iterations, RBF/linear, C logU(1e-3,1e2), gamma logU(1e-4,1e-1); LightGBM: 15 randomized iterations. Do not combine with separate original GlycanML-paper protocols.

Aggregation: Not reported

GlycanGT published supplementary archive, Table S4; glycangt: Journal full-text XML · btag147_supplementary_data.zip / Table_S4.xlsx / Sheet1!A84:E84 (mean D84, SD E84)
Pipeline: GlycanGTProtocol: GlycanML glycosylation Macro-F1: GlycanGT study: glycosylation Macro-F1
Dataset subset: GlyConnect glycosylation; GlycanML official motif split (GlycanML split)
0.932366005641867 ± 0 macro_f1
dimensionless · higher

Uncertainty: type: standard_deviation; value: 0

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · source checked
Methods, coverage and source

GlycanGT on GlycanML glycosylation Macro-F1: GlycanGT study: glycosylation Macro-F1

Glycosylation: 1,683 glycans total; N-linked/O-linked/free. Official fixed GlycanML motif-based train/validation/test splits (8:1:1). Section 2.5 reports class-balanced classifiers, train ∪ validation hyperparameter selection by randomized search with 3-fold cross-validation, followed by one evaluation on the held-out test set; complete procedure repeated with three random seeds, reporting mean and standard deviation. GlycanGT large model pretrained with 35% masking provides [Graph] embeddings to SVM/LightGBM; the selected classifier for each S4 row is not identified. Section 2.6 states that all four graph baselines were trained and evaluated on the same datasets/splits; it does not establish that each baseline used the GlycanGT downstream classifier search. SVM search: 10 iterations, RBF/linear, C logU(1e-3,1e2), gamma logU(1e-4,1e-1); LightGBM: 15 randomized iterations. Do not combine with separate original GlycanML-paper protocols.

Aggregation: Not reported

GlycanGT published supplementary archive, Table S4; glycangt: Journal full-text XML · btag147_supplementary_data.zip / Table_S4.xlsx / Sheet1!A85:E85 (mean D85, SD E85)
Pipeline: GlycanGTProtocol: GlycanML immunogenicity Accuracy: GlycanGT study: immunogenicity Accuracy
Dataset subset: SugarBase immunogenicity; GlycanML official motif split (GlycanML split)
0.93333333333333302 ± 0.031853807955289602 accuracy
fraction · higher

Uncertainty: type: standard_deviation; value: 0.031853807955289602

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · source checked
Methods, coverage and source

GlycanGT on GlycanML immunogenicity Accuracy: GlycanGT study: immunogenicity Accuracy

Immunogenicity: 1,320 glycans total; binary immune activity. Official fixed GlycanML motif-based train/validation/test splits (8:1:1). Section 2.5 reports class-balanced classifiers, train ∪ validation hyperparameter selection by randomized search with 3-fold cross-validation, followed by one evaluation on the held-out test set; complete procedure repeated with three random seeds, reporting mean and standard deviation. GlycanGT large model pretrained with 35% masking provides [Graph] embeddings to SVM/LightGBM; the selected classifier for each S4 row is not identified. Section 2.6 states that all four graph baselines were trained and evaluated on the same datasets/splits; it does not establish that each baseline used the GlycanGT downstream classifier search. SVM search: 10 iterations, RBF/linear, C logU(1e-3,1e2), gamma logU(1e-4,1e-1); LightGBM: 15 randomized iterations. Do not combine with separate original GlycanML-paper protocols.

Aggregation: Not reported

GlycanGT published supplementary archive, Table S4; glycangt: Journal full-text XML · btag147_supplementary_data.zip / Table_S4.xlsx / Sheet1!A81:E81 (mean D81, SD E81)
Pipeline: GlycanGTProtocol: GlycanML immunogenicity AUPRC: GlycanGT study: immunogenicity AUPRC
Dataset subset: SugarBase immunogenicity; GlycanML official motif split (GlycanML split)
0.84419033765294005 ± 0.0036259797661145998 auprc
dimensionless · higher

Uncertainty: type: standard_deviation; value: 0.0036259797661145998

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · source checked
Methods, coverage and source

GlycanGT on GlycanML immunogenicity AUPRC: GlycanGT study: immunogenicity AUPRC

Immunogenicity: 1,320 glycans total; binary immune activity. Official fixed GlycanML motif-based train/validation/test splits (8:1:1). Section 2.5 reports class-balanced classifiers, train ∪ validation hyperparameter selection by randomized search with 3-fold cross-validation, followed by one evaluation on the held-out test set; complete procedure repeated with three random seeds, reporting mean and standard deviation. GlycanGT large model pretrained with 35% masking provides [Graph] embeddings to SVM/LightGBM; the selected classifier for each S4 row is not identified. Section 2.6 states that all four graph baselines were trained and evaluated on the same datasets/splits; it does not establish that each baseline used the GlycanGT downstream classifier search. SVM search: 10 iterations, RBF/linear, C logU(1e-3,1e2), gamma logU(1e-4,1e-1); LightGBM: 15 randomized iterations. Do not combine with separate original GlycanML-paper protocols.

Aggregation: Not reported

GlycanGT published supplementary archive, Table S4; glycangt: Journal full-text XML · btag147_supplementary_data.zip / Table_S4.xlsx / Sheet1!A83:E83 (mean D83, SD E83)
Pipeline: GlycanGTProtocol: GlycanML immunogenicity Macro-F1: GlycanGT study: immunogenicity Macro-F1
Dataset subset: SugarBase immunogenicity; GlycanML official motif split (GlycanML split)
0.87036334625073897 ± 0.065078625040751806 macro_f1
dimensionless · higher

Uncertainty: type: standard_deviation; value: 0.065078625040751806

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · source checked
Methods, coverage and source

GlycanGT on GlycanML immunogenicity Macro-F1: GlycanGT study: immunogenicity Macro-F1

Immunogenicity: 1,320 glycans total; binary immune activity. Official fixed GlycanML motif-based train/validation/test splits (8:1:1). Section 2.5 reports class-balanced classifiers, train ∪ validation hyperparameter selection by randomized search with 3-fold cross-validation, followed by one evaluation on the held-out test set; complete procedure repeated with three random seeds, reporting mean and standard deviation. GlycanGT large model pretrained with 35% masking provides [Graph] embeddings to SVM/LightGBM; the selected classifier for each S4 row is not identified. Section 2.6 states that all four graph baselines were trained and evaluated on the same datasets/splits; it does not establish that each baseline used the GlycanGT downstream classifier search. SVM search: 10 iterations, RBF/linear, C logU(1e-3,1e2), gamma logU(1e-4,1e-1); LightGBM: 15 randomized iterations. Do not combine with separate original GlycanML-paper protocols.

Aggregation: Not reported

GlycanGT published supplementary archive, Table S4; glycangt: Journal full-text XML · btag147_supplementary_data.zip / Table_S4.xlsx / Sheet1!A82:E82 (mean D82, SD E82)
Pipeline: GlycanGTProtocol: GlycanML kingdom Accuracy: GlycanGT study: kingdom Accuracy
Dataset subset: SugarBase taxonomy kingdom; GlycanML official motif split (GlycanML split)
0.90243017772941603 ± 0.017489392022112599 accuracy
fraction · higher

Uncertainty: type: standard_deviation; value: 0.017489392022112599

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · source checked
Methods, coverage and source

GlycanGT on GlycanML kingdom Accuracy: GlycanGT study: kingdom Accuracy

Taxonomy: 13,209 glycans total across eight levels, 4–1,737 classes per level. Official fixed GlycanML motif-based train/validation/test splits (8:1:1). Section 2.5 reports class-balanced classifiers, train ∪ validation hyperparameter selection by randomized search with 3-fold cross-validation, followed by one evaluation on the held-out test set; complete procedure repeated with three random seeds, reporting mean and standard deviation. GlycanGT large model pretrained with 35% masking provides [Graph] embeddings to SVM/LightGBM; the selected classifier for each S4 row is not identified. Section 2.6 states that all four graph baselines were trained and evaluated on the same datasets/splits; it does not establish that each baseline used the GlycanGT downstream classifier search. SVM search: 10 iterations, RBF/linear, C logU(1e-3,1e2), gamma logU(1e-4,1e-1); LightGBM: 15 randomized iterations. Do not combine with separate original GlycanML-paper protocols.

Aggregation: Not reported

GlycanGT published supplementary archive, Table S4; glycangt: Journal full-text XML · btag147_supplementary_data.zip / Table_S4.xlsx / Sheet1!A67:E67 (mean D67, SD E67)
Pipeline: GlycanGTProtocol: GlycanML kingdom Macro-F1: GlycanGT study: kingdom Macro-F1
Dataset subset: SugarBase taxonomy kingdom; GlycanML official motif split (GlycanML split)
0.66958587101112299 ± 0.082979941298309295 macro_f1
dimensionless · higher

Uncertainty: type: standard_deviation; value: 0.082979941298309295

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · source checked
Methods, coverage and source

GlycanGT on GlycanML kingdom Macro-F1: GlycanGT study: kingdom Macro-F1

Taxonomy: 13,209 glycans total across eight levels, 4–1,737 classes per level. Official fixed GlycanML motif-based train/validation/test splits (8:1:1). Section 2.5 reports class-balanced classifiers, train ∪ validation hyperparameter selection by randomized search with 3-fold cross-validation, followed by one evaluation on the held-out test set; complete procedure repeated with three random seeds, reporting mean and standard deviation. GlycanGT large model pretrained with 35% masking provides [Graph] embeddings to SVM/LightGBM; the selected classifier for each S4 row is not identified. Section 2.6 states that all four graph baselines were trained and evaluated on the same datasets/splits; it does not establish that each baseline used the GlycanGT downstream classifier search. SVM search: 10 iterations, RBF/linear, C logU(1e-3,1e2), gamma logU(1e-4,1e-1); LightGBM: 15 randomized iterations. Do not combine with separate original GlycanML-paper protocols.

Aggregation: Not reported

GlycanGT published supplementary archive, Table S4; glycangt: Journal full-text XML · btag147_supplementary_data.zip / Table_S4.xlsx / Sheet1!A68:E68 (mean D68, SD E68)
Pipeline: GlycanGTProtocol: GlycanML order Accuracy: GlycanGT study: order Accuracy
Dataset subset: SugarBase taxonomy order; GlycanML official motif split (GlycanML split)
0.466086325716358 ± 0.020598144163221799 accuracy
fraction · higher

Uncertainty: type: standard_deviation; value: 0.020598144163221799

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · source checked
Methods, coverage and source

GlycanGT on GlycanML order Accuracy: GlycanGT study: order Accuracy

Taxonomy: 13,209 glycans total across eight levels, 4–1,737 classes per level. Official fixed GlycanML motif-based train/validation/test splits (8:1:1). Section 2.5 reports class-balanced classifiers, train ∪ validation hyperparameter selection by randomized search with 3-fold cross-validation, followed by one evaluation on the held-out test set; complete procedure repeated with three random seeds, reporting mean and standard deviation. GlycanGT large model pretrained with 35% masking provides [Graph] embeddings to SVM/LightGBM; the selected classifier for each S4 row is not identified. Section 2.6 states that all four graph baselines were trained and evaluated on the same datasets/splits; it does not establish that each baseline used the GlycanGT downstream classifier search. SVM search: 10 iterations, RBF/linear, C logU(1e-3,1e2), gamma logU(1e-4,1e-1); LightGBM: 15 randomized iterations. Do not combine with separate original GlycanML-paper protocols.

Aggregation: Not reported

GlycanGT published supplementary archive, Table S4; glycangt: Journal full-text XML · btag147_supplementary_data.zip / Table_S4.xlsx / Sheet1!A73:E73 (mean D73, SD E73)
Pipeline: GlycanGTProtocol: GlycanML order Macro-F1: GlycanGT study: order Macro-F1
Dataset subset: SugarBase taxonomy order; GlycanML official motif split (GlycanML split)
0.278530744279541 ± 0.0086310922520820999 macro_f1
dimensionless · higher

Uncertainty: type: standard_deviation; value: 0.0086310922520820999

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · source checked
Methods, coverage and source

GlycanGT on GlycanML order Macro-F1: GlycanGT study: order Macro-F1

Taxonomy: 13,209 glycans total across eight levels, 4–1,737 classes per level. Official fixed GlycanML motif-based train/validation/test splits (8:1:1). Section 2.5 reports class-balanced classifiers, train ∪ validation hyperparameter selection by randomized search with 3-fold cross-validation, followed by one evaluation on the held-out test set; complete procedure repeated with three random seeds, reporting mean and standard deviation. GlycanGT large model pretrained with 35% masking provides [Graph] embeddings to SVM/LightGBM; the selected classifier for each S4 row is not identified. Section 2.6 states that all four graph baselines were trained and evaluated on the same datasets/splits; it does not establish that each baseline used the GlycanGT downstream classifier search. SVM search: 10 iterations, RBF/linear, C logU(1e-3,1e2), gamma logU(1e-4,1e-1); LightGBM: 15 randomized iterations. Do not combine with separate original GlycanML-paper protocols.

Aggregation: Not reported

GlycanGT published supplementary archive, Table S4; glycangt: Journal full-text XML · btag147_supplementary_data.zip / Table_S4.xlsx / Sheet1!A74:E74 (mean D74, SD E74)
Pipeline: GlycanGTProtocol: GlycanML phylum Accuracy: GlycanGT study: phylum Accuracy
Dataset subset: SugarBase taxonomy phylum; GlycanML official motif split (GlycanML split)
0.84149437794704296 ± 0.003324320417088 accuracy
fraction · higher

Uncertainty: type: standard_deviation; value: 0.003324320417088

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · source checked
Methods, coverage and source

GlycanGT on GlycanML phylum Accuracy: GlycanGT study: phylum Accuracy

Taxonomy: 13,209 glycans total across eight levels, 4–1,737 classes per level. Official fixed GlycanML motif-based train/validation/test splits (8:1:1). Section 2.5 reports class-balanced classifiers, train ∪ validation hyperparameter selection by randomized search with 3-fold cross-validation, followed by one evaluation on the held-out test set; complete procedure repeated with three random seeds, reporting mean and standard deviation. GlycanGT large model pretrained with 35% masking provides [Graph] embeddings to SVM/LightGBM; the selected classifier for each S4 row is not identified. Section 2.6 states that all four graph baselines were trained and evaluated on the same datasets/splits; it does not establish that each baseline used the GlycanGT downstream classifier search. SVM search: 10 iterations, RBF/linear, C logU(1e-3,1e2), gamma logU(1e-4,1e-1); LightGBM: 15 randomized iterations. Do not combine with separate original GlycanML-paper protocols.

Aggregation: Not reported

GlycanGT published supplementary archive, Table S4; glycangt: Journal full-text XML · btag147_supplementary_data.zip / Table_S4.xlsx / Sheet1!A69:E69 (mean D69, SD E69)
Pipeline: GlycanGTProtocol: GlycanML phylum Macro-F1: GlycanGT study: phylum Macro-F1
Dataset subset: SugarBase taxonomy phylum; GlycanML official motif split (GlycanML split)
0.49333063993822701 ± 0.080739293560509295 macro_f1
dimensionless · higher

Uncertainty: type: standard_deviation; value: 0.080739293560509295

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · source checked
Methods, coverage and source

GlycanGT on GlycanML phylum Macro-F1: GlycanGT study: phylum Macro-F1

Taxonomy: 13,209 glycans total across eight levels, 4–1,737 classes per level. Official fixed GlycanML motif-based train/validation/test splits (8:1:1). Section 2.5 reports class-balanced classifiers, train ∪ validation hyperparameter selection by randomized search with 3-fold cross-validation, followed by one evaluation on the held-out test set; complete procedure repeated with three random seeds, reporting mean and standard deviation. GlycanGT large model pretrained with 35% masking provides [Graph] embeddings to SVM/LightGBM; the selected classifier for each S4 row is not identified. Section 2.6 states that all four graph baselines were trained and evaluated on the same datasets/splits; it does not establish that each baseline used the GlycanGT downstream classifier search. SVM search: 10 iterations, RBF/linear, C logU(1e-3,1e2), gamma logU(1e-4,1e-1); LightGBM: 15 randomized iterations. Do not combine with separate original GlycanML-paper protocols.

Aggregation: Not reported

GlycanGT published supplementary archive, Table S4; glycangt: Journal full-text XML · btag147_supplementary_data.zip / Table_S4.xlsx / Sheet1!A70:E70 (mean D70, SD E70)
Pipeline: GlycanGTProtocol: GlycanML species Accuracy: GlycanGT study: species Accuracy
Dataset subset: SugarBase taxonomy species; GlycanML official motif split (GlycanML split)
0.39173014145810597 ± 0.021681021594310401 accuracy
fraction · higher

Uncertainty: type: standard_deviation; value: 0.021681021594310401

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · source checked
Methods, coverage and source

GlycanGT on GlycanML species Accuracy: GlycanGT study: species Accuracy

Taxonomy: 13,209 glycans total across eight levels, 4–1,737 classes per level. Official fixed GlycanML motif-based train/validation/test splits (8:1:1). Section 2.5 reports class-balanced classifiers, train ∪ validation hyperparameter selection by randomized search with 3-fold cross-validation, followed by one evaluation on the held-out test set; complete procedure repeated with three random seeds, reporting mean and standard deviation. GlycanGT large model pretrained with 35% masking provides [Graph] embeddings to SVM/LightGBM; the selected classifier for each S4 row is not identified. Section 2.6 states that all four graph baselines were trained and evaluated on the same datasets/splits; it does not establish that each baseline used the GlycanGT downstream classifier search. SVM search: 10 iterations, RBF/linear, C logU(1e-3,1e2), gamma logU(1e-4,1e-1); LightGBM: 15 randomized iterations. Do not combine with separate original GlycanML-paper protocols.

Aggregation: Not reported

GlycanGT published supplementary archive, Table S4; glycangt: Journal full-text XML · btag147_supplementary_data.zip / Table_S4.xlsx / Sheet1!A79:E79 (mean D79, SD E79)
Pipeline: GlycanGTProtocol: GlycanML species Macro-F1: GlycanGT study: species Macro-F1
Dataset subset: SugarBase taxonomy species; GlycanML official motif split (GlycanML split)
0.184571928631623 ± 0.00080738146274129995 macro_f1
dimensionless · higher

Uncertainty: type: standard_deviation; value: 0.00080738146274129995

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · source checked
Methods, coverage and source

GlycanGT on GlycanML species Macro-F1: GlycanGT study: species Macro-F1

Taxonomy: 13,209 glycans total across eight levels, 4–1,737 classes per level. Official fixed GlycanML motif-based train/validation/test splits (8:1:1). Section 2.5 reports class-balanced classifiers, train ∪ validation hyperparameter selection by randomized search with 3-fold cross-validation, followed by one evaluation on the held-out test set; complete procedure repeated with three random seeds, reporting mean and standard deviation. GlycanGT large model pretrained with 35% masking provides [Graph] embeddings to SVM/LightGBM; the selected classifier for each S4 row is not identified. Section 2.6 states that all four graph baselines were trained and evaluated on the same datasets/splits; it does not establish that each baseline used the GlycanGT downstream classifier search. SVM search: 10 iterations, RBF/linear, C logU(1e-3,1e2), gamma logU(1e-4,1e-1); LightGBM: 15 randomized iterations. Do not combine with separate original GlycanML-paper protocols.

Aggregation: Not reported

GlycanGT published supplementary archive, Table S4; glycangt: Journal full-text XML · btag147_supplementary_data.zip / Table_S4.xlsx / Sheet1!A80:E80 (mean D80, SD E80)

Source checking is not independent reproduction. Release 2026-09-23-2b89723c6dd9.

Use this model

How it works, versions and access

Underlying model: GlycanGT. Results on this page belong to this pipeline and its evaluated settings.

Strengths, limitations and unresolved questions

Evidence

Source checking verifies the cited claim or transcription. It does not establish independent reproduction.

Evidence table

Inspect claims, sources and review details

Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.

One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.

2 evidence rows matching the loaded filters

Claims, original sources and review scope · Release 2026-09-23-2b89723c6dd9
Property and statementOriginal source and locationReview and provenance
Relationship: uses model
discovery-model-glycangt
Individual claims
glycangt: Journal full-text XML

Original source ↗

GlycanGT primary article Sections 2.1, 2.5, 2.6 and 3.1–3.2 (PMC13105845), Supplementary Table S4; column B GlycanGT

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: Primary article XML snapshot
Retrieved: 2026-09-16T20:20:56.439056+00:00

source checked

automated source review · 2026-09-23

Audit details

Source-backed evaluated identity only; no independent reproduction.

Field: links:uses_model:discovery-model-glycangt

Claim: glycangt-2026-table-s4-method-glycangt-discovery-model-glycangt-identity-claim

Source artifact SHA-256: 53e89a636c868c0329ee7eb6ae92f1028ec891940bc61730a148981b647fbbe5

Hash scope: SHA-256 of retrieved original artifact bytes

Format: original_artifact

Inspected artifact

Relationship: uses model
discovery-model-glycangt
Individual claims
GlycanGT published supplementary archive, Table S4

Original source ↗

GlycanGT primary article Sections 2.1, 2.5, 2.6 and 3.1–3.2 (PMC13105845), Supplementary Table S4; column B GlycanGT

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: Published Bioinformatics btag147 supplementary archive, retrieved 2026-09-23; Table_S4.xlsx SHA-256 d7c35909bac6bcb78ca8fdb32c0463f05e692f15861a8184e65e415e7216f493
Retrieved: 2026-09-23T11:31:33.520570+00:00

source checked

automated source review · 2026-09-23

Audit details

Source-backed evaluated identity only; no independent reproduction.

Field: links:uses_model:discovery-model-glycangt

Claim: glycangt-2026-table-s4-method-glycangt-discovery-model-glycangt-identity-claim

Source artifact SHA-256: 27748c6c0c0bb07b0105d274fb745fd4dfe702d34b9ee367d1ddacbc71c56ab0

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Sources and history

View linked audit checks and correction history

Release 2026-09-23-2b89723c6dd9 · Record review: source checked

2 source records and release historyDownload this release
Technical metadata and extraction receipts

Stable ID: glycangt-2026-table-s4-method-glycangt

areas
glycans
source locator
GlycanGT primary article Sections 2.1, 2.5, 2.6 and 3.1–3.2 (PMC13105845), Supplementary Table S4; column B GlycanGT
missing metadata
checkpoint revision: unreported; parameters: unextracted
Related records

Suggest a correction