GlycanGT
Large GlycanGT encoder, 35% node/edge masking pretraining; [Graph] embeddings supplied to a trained SVM or LightGBM classifier. Table S4 does not identify the chosen classifier per task. This is a downstream prediction pipeline, not the bare encoder.
Overview
Large GlycanGT encoder, 35% node/edge masking pretraining; [Graph] embeddings supplied to a trained SVM or LightGBM classifier. Table S4 does not identify the chosen classifier per task. This is a downstream prediction pipeline, not the bare encoder.
Consult the linked sources for architecture or protocol details. Missing evidence is not evidence of a missing capability.
Evaluations and results
21 evaluations · 21 metric rows. Different protocols are not a single leaderboard.
Filter evaluations
Applied filters: All linked evaluations
| Tested configuration | Protocol and dataset | Finding | Evidence and details |
|---|---|---|---|
| Pipeline: GlycanGT | Protocol: GlycanML class Accuracy: GlycanGT study: class Accuracy Dataset subset: SugarBase taxonomy class; GlycanML official motif split (GlycanML split) | 0.69350743561842498 ± 0.0072451883770178003 accuracy fraction · higher Uncertainty: type: standard_deviation; value: 0.0072451883770178003 Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · source checkedMethods, coverage and sourceGlycanGT on GlycanML class Accuracy: GlycanGT study: class Accuracy Taxonomy: 13,209 glycans total across eight levels, 4–1,737 classes per level. Official fixed GlycanML motif-based train/validation/test splits (8:1:1). Section 2.5 reports class-balanced classifiers, train ∪ validation hyperparameter selection by randomized search with 3-fold cross-validation, followed by one evaluation on the held-out test set; complete procedure repeated with three random seeds, reporting mean and standard deviation. GlycanGT large model pretrained with 35% masking provides [Graph] embeddings to SVM/LightGBM; the selected classifier for each S4 row is not identified. Section 2.6 states that all four graph baselines were trained and evaluated on the same datasets/splits; it does not establish that each baseline used the GlycanGT downstream classifier search. SVM search: 10 iterations, RBF/linear, C logU(1e-3,1e2), gamma logU(1e-4,1e-1); LightGBM: 15 randomized iterations. Do not combine with separate original GlycanML-paper protocols. Aggregation: Not reported GlycanGT published supplementary archive, Table S4; glycangt: Journal full-text XML · btag147_supplementary_data.zip / Table_S4.xlsx / Sheet1!A71:E71 (mean D71, SD E71) |
| Pipeline: GlycanGT | Protocol: GlycanML class Macro-F1: GlycanGT study: class Macro-F1 Dataset subset: SugarBase taxonomy class; GlycanML official motif split (GlycanML split) | 0.39715349769695701 ± 0.0084870172846663004 macro_f1 dimensionless · higher Uncertainty: type: standard_deviation; value: 0.0084870172846663004 Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · source checkedMethods, coverage and sourceGlycanGT on GlycanML class Macro-F1: GlycanGT study: class Macro-F1 Taxonomy: 13,209 glycans total across eight levels, 4–1,737 classes per level. Official fixed GlycanML motif-based train/validation/test splits (8:1:1). Section 2.5 reports class-balanced classifiers, train ∪ validation hyperparameter selection by randomized search with 3-fold cross-validation, followed by one evaluation on the held-out test set; complete procedure repeated with three random seeds, reporting mean and standard deviation. GlycanGT large model pretrained with 35% masking provides [Graph] embeddings to SVM/LightGBM; the selected classifier for each S4 row is not identified. Section 2.6 states that all four graph baselines were trained and evaluated on the same datasets/splits; it does not establish that each baseline used the GlycanGT downstream classifier search. SVM search: 10 iterations, RBF/linear, C logU(1e-3,1e2), gamma logU(1e-4,1e-1); LightGBM: 15 randomized iterations. Do not combine with separate original GlycanML-paper protocols. Aggregation: Not reported GlycanGT published supplementary archive, Table S4; glycangt: Journal full-text XML · btag147_supplementary_data.zip / Table_S4.xlsx / Sheet1!A72:E72 (mean D72, SD E72) |
| Pipeline: GlycanGT | Protocol: GlycanML domain Accuracy: GlycanGT study: domain Accuracy Dataset subset: SugarBase taxonomy domain; GlycanML official motif split (GlycanML split) | 0.92020311933260801 ± 0.0077199117351442002 accuracy fraction · higher Uncertainty: type: standard_deviation; value: 0.0077199117351442002 Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · source checkedMethods, coverage and sourceGlycanGT on GlycanML domain Accuracy: GlycanGT study: domain Accuracy Taxonomy: 13,209 glycans total across eight levels, 4–1,737 classes per level. Official fixed GlycanML motif-based train/validation/test splits (8:1:1). Section 2.5 reports class-balanced classifiers, train ∪ validation hyperparameter selection by randomized search with 3-fold cross-validation, followed by one evaluation on the held-out test set; complete procedure repeated with three random seeds, reporting mean and standard deviation. GlycanGT large model pretrained with 35% masking provides [Graph] embeddings to SVM/LightGBM; the selected classifier for each S4 row is not identified. Section 2.6 states that all four graph baselines were trained and evaluated on the same datasets/splits; it does not establish that each baseline used the GlycanGT downstream classifier search. SVM search: 10 iterations, RBF/linear, C logU(1e-3,1e2), gamma logU(1e-4,1e-1); LightGBM: 15 randomized iterations. Do not combine with separate original GlycanML-paper protocols. Aggregation: Not reported GlycanGT published supplementary archive, Table S4; glycangt: Journal full-text XML · btag147_supplementary_data.zip / Table_S4.xlsx / Sheet1!A65:E65 (mean D65, SD E65) |
| Pipeline: GlycanGT | Protocol: GlycanML domain Macro-F1: GlycanGT study: domain Macro-F1 Dataset subset: SugarBase taxonomy domain; GlycanML official motif split (GlycanML split) | 0.73873655202321897 ± 0.023629706943064301 macro_f1 dimensionless · higher Uncertainty: type: standard_deviation; value: 0.023629706943064301 Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · source checkedMethods, coverage and sourceGlycanGT on GlycanML domain Macro-F1: GlycanGT study: domain Macro-F1 Taxonomy: 13,209 glycans total across eight levels, 4–1,737 classes per level. Official fixed GlycanML motif-based train/validation/test splits (8:1:1). Section 2.5 reports class-balanced classifiers, train ∪ validation hyperparameter selection by randomized search with 3-fold cross-validation, followed by one evaluation on the held-out test set; complete procedure repeated with three random seeds, reporting mean and standard deviation. GlycanGT large model pretrained with 35% masking provides [Graph] embeddings to SVM/LightGBM; the selected classifier for each S4 row is not identified. Section 2.6 states that all four graph baselines were trained and evaluated on the same datasets/splits; it does not establish that each baseline used the GlycanGT downstream classifier search. SVM search: 10 iterations, RBF/linear, C logU(1e-3,1e2), gamma logU(1e-4,1e-1); LightGBM: 15 randomized iterations. Do not combine with separate original GlycanML-paper protocols. Aggregation: Not reported GlycanGT published supplementary archive, Table S4; glycangt: Journal full-text XML · btag147_supplementary_data.zip / Table_S4.xlsx / Sheet1!A66:E66 (mean D66, SD E66) |
| Pipeline: GlycanGT | Protocol: GlycanML family Accuracy: GlycanGT study: family Accuracy Dataset subset: SugarBase taxonomy family; GlycanML official motif split (GlycanML split) | 0.463184620964816 ± 0.0078717934037760007 accuracy fraction · higher Uncertainty: type: standard_deviation; value: 0.0078717934037760007 Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · source checkedMethods, coverage and sourceGlycanGT on GlycanML family Accuracy: GlycanGT study: family Accuracy Taxonomy: 13,209 glycans total across eight levels, 4–1,737 classes per level. Official fixed GlycanML motif-based train/validation/test splits (8:1:1). Section 2.5 reports class-balanced classifiers, train ∪ validation hyperparameter selection by randomized search with 3-fold cross-validation, followed by one evaluation on the held-out test set; complete procedure repeated with three random seeds, reporting mean and standard deviation. GlycanGT large model pretrained with 35% masking provides [Graph] embeddings to SVM/LightGBM; the selected classifier for each S4 row is not identified. Section 2.6 states that all four graph baselines were trained and evaluated on the same datasets/splits; it does not establish that each baseline used the GlycanGT downstream classifier search. SVM search: 10 iterations, RBF/linear, C logU(1e-3,1e2), gamma logU(1e-4,1e-1); LightGBM: 15 randomized iterations. Do not combine with separate original GlycanML-paper protocols. Aggregation: Not reported GlycanGT published supplementary archive, Table S4; glycangt: Journal full-text XML · btag147_supplementary_data.zip / Table_S4.xlsx / Sheet1!A75:E75 (mean D75, SD E75) |
| Pipeline: GlycanGT | Protocol: GlycanML family Macro-F1: GlycanGT study: family Macro-F1 Dataset subset: SugarBase taxonomy family; GlycanML official motif split (GlycanML split) | 0.24638700174167499 ± 0.004794223931708 macro_f1 dimensionless · higher Uncertainty: type: standard_deviation; value: 0.004794223931708 Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · source checkedMethods, coverage and sourceGlycanGT on GlycanML family Macro-F1: GlycanGT study: family Macro-F1 Taxonomy: 13,209 glycans total across eight levels, 4–1,737 classes per level. Official fixed GlycanML motif-based train/validation/test splits (8:1:1). Section 2.5 reports class-balanced classifiers, train ∪ validation hyperparameter selection by randomized search with 3-fold cross-validation, followed by one evaluation on the held-out test set; complete procedure repeated with three random seeds, reporting mean and standard deviation. GlycanGT large model pretrained with 35% masking provides [Graph] embeddings to SVM/LightGBM; the selected classifier for each S4 row is not identified. Section 2.6 states that all four graph baselines were trained and evaluated on the same datasets/splits; it does not establish that each baseline used the GlycanGT downstream classifier search. SVM search: 10 iterations, RBF/linear, C logU(1e-3,1e2), gamma logU(1e-4,1e-1); LightGBM: 15 randomized iterations. Do not combine with separate original GlycanML-paper protocols. Aggregation: Not reported GlycanGT published supplementary archive, Table S4; glycangt: Journal full-text XML · btag147_supplementary_data.zip / Table_S4.xlsx / Sheet1!A76:E76 (mean D76, SD E76) |
| Pipeline: GlycanGT | Protocol: GlycanML genus Accuracy: GlycanGT study: genus Accuracy Dataset subset: SugarBase taxonomy genus; GlycanML official motif split (GlycanML split) | 0.44577439245556699 ± 0.0045302850913300002 accuracy fraction · higher Uncertainty: type: standard_deviation; value: 0.0045302850913300002 Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · source checkedMethods, coverage and sourceGlycanGT on GlycanML genus Accuracy: GlycanGT study: genus Accuracy Taxonomy: 13,209 glycans total across eight levels, 4–1,737 classes per level. Official fixed GlycanML motif-based train/validation/test splits (8:1:1). Section 2.5 reports class-balanced classifiers, train ∪ validation hyperparameter selection by randomized search with 3-fold cross-validation, followed by one evaluation on the held-out test set; complete procedure repeated with three random seeds, reporting mean and standard deviation. GlycanGT large model pretrained with 35% masking provides [Graph] embeddings to SVM/LightGBM; the selected classifier for each S4 row is not identified. Section 2.6 states that all four graph baselines were trained and evaluated on the same datasets/splits; it does not establish that each baseline used the GlycanGT downstream classifier search. SVM search: 10 iterations, RBF/linear, C logU(1e-3,1e2), gamma logU(1e-4,1e-1); LightGBM: 15 randomized iterations. Do not combine with separate original GlycanML-paper protocols. Aggregation: Not reported GlycanGT published supplementary archive, Table S4; glycangt: Journal full-text XML · btag147_supplementary_data.zip / Table_S4.xlsx / Sheet1!A77:E77 (mean D77, SD E77) |
| Pipeline: GlycanGT | Protocol: GlycanML genus Macro-F1: GlycanGT study: genus Macro-F1 Dataset subset: SugarBase taxonomy genus; GlycanML official motif split (GlycanML split) | 0.188457844899896 ± 0.0007473331496721 macro_f1 dimensionless · higher Uncertainty: type: standard_deviation; value: 0.0007473331496721 Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · source checkedMethods, coverage and sourceGlycanGT on GlycanML genus Macro-F1: GlycanGT study: genus Macro-F1 Taxonomy: 13,209 glycans total across eight levels, 4–1,737 classes per level. Official fixed GlycanML motif-based train/validation/test splits (8:1:1). Section 2.5 reports class-balanced classifiers, train ∪ validation hyperparameter selection by randomized search with 3-fold cross-validation, followed by one evaluation on the held-out test set; complete procedure repeated with three random seeds, reporting mean and standard deviation. GlycanGT large model pretrained with 35% masking provides [Graph] embeddings to SVM/LightGBM; the selected classifier for each S4 row is not identified. Section 2.6 states that all four graph baselines were trained and evaluated on the same datasets/splits; it does not establish that each baseline used the GlycanGT downstream classifier search. SVM search: 10 iterations, RBF/linear, C logU(1e-3,1e2), gamma logU(1e-4,1e-1); LightGBM: 15 randomized iterations. Do not combine with separate original GlycanML-paper protocols. Aggregation: Not reported GlycanGT published supplementary archive, Table S4; glycangt: Journal full-text XML · btag147_supplementary_data.zip / Table_S4.xlsx / Sheet1!A78:E78 (mean D78, SD E78) |
| Pipeline: GlycanGT | Protocol: GlycanML glycosylation Accuracy: GlycanGT study: glycosylation Accuracy Dataset subset: GlyConnect glycosylation; GlycanML official motif split (GlycanML split) | 0.98295454545454497 ± 0 accuracy fraction · higher Uncertainty: type: standard_deviation; value: 0 Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · source checkedMethods, coverage and sourceGlycanGT on GlycanML glycosylation Accuracy: GlycanGT study: glycosylation Accuracy Glycosylation: 1,683 glycans total; N-linked/O-linked/free. Official fixed GlycanML motif-based train/validation/test splits (8:1:1). Section 2.5 reports class-balanced classifiers, train ∪ validation hyperparameter selection by randomized search with 3-fold cross-validation, followed by one evaluation on the held-out test set; complete procedure repeated with three random seeds, reporting mean and standard deviation. GlycanGT large model pretrained with 35% masking provides [Graph] embeddings to SVM/LightGBM; the selected classifier for each S4 row is not identified. Section 2.6 states that all four graph baselines were trained and evaluated on the same datasets/splits; it does not establish that each baseline used the GlycanGT downstream classifier search. SVM search: 10 iterations, RBF/linear, C logU(1e-3,1e2), gamma logU(1e-4,1e-1); LightGBM: 15 randomized iterations. Do not combine with separate original GlycanML-paper protocols. Aggregation: Not reported GlycanGT published supplementary archive, Table S4; glycangt: Journal full-text XML · btag147_supplementary_data.zip / Table_S4.xlsx / Sheet1!A84:E84 (mean D84, SD E84) |
| Pipeline: GlycanGT | Protocol: GlycanML glycosylation Macro-F1: GlycanGT study: glycosylation Macro-F1 Dataset subset: GlyConnect glycosylation; GlycanML official motif split (GlycanML split) | 0.932366005641867 ± 0 macro_f1 dimensionless · higher Uncertainty: type: standard_deviation; value: 0 Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · source checkedMethods, coverage and sourceGlycanGT on GlycanML glycosylation Macro-F1: GlycanGT study: glycosylation Macro-F1 Glycosylation: 1,683 glycans total; N-linked/O-linked/free. Official fixed GlycanML motif-based train/validation/test splits (8:1:1). Section 2.5 reports class-balanced classifiers, train ∪ validation hyperparameter selection by randomized search with 3-fold cross-validation, followed by one evaluation on the held-out test set; complete procedure repeated with three random seeds, reporting mean and standard deviation. GlycanGT large model pretrained with 35% masking provides [Graph] embeddings to SVM/LightGBM; the selected classifier for each S4 row is not identified. Section 2.6 states that all four graph baselines were trained and evaluated on the same datasets/splits; it does not establish that each baseline used the GlycanGT downstream classifier search. SVM search: 10 iterations, RBF/linear, C logU(1e-3,1e2), gamma logU(1e-4,1e-1); LightGBM: 15 randomized iterations. Do not combine with separate original GlycanML-paper protocols. Aggregation: Not reported GlycanGT published supplementary archive, Table S4; glycangt: Journal full-text XML · btag147_supplementary_data.zip / Table_S4.xlsx / Sheet1!A85:E85 (mean D85, SD E85) |
| Pipeline: GlycanGT | Protocol: GlycanML immunogenicity Accuracy: GlycanGT study: immunogenicity Accuracy Dataset subset: SugarBase immunogenicity; GlycanML official motif split (GlycanML split) | 0.93333333333333302 ± 0.031853807955289602 accuracy fraction · higher Uncertainty: type: standard_deviation; value: 0.031853807955289602 Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · source checkedMethods, coverage and sourceGlycanGT on GlycanML immunogenicity Accuracy: GlycanGT study: immunogenicity Accuracy Immunogenicity: 1,320 glycans total; binary immune activity. Official fixed GlycanML motif-based train/validation/test splits (8:1:1). Section 2.5 reports class-balanced classifiers, train ∪ validation hyperparameter selection by randomized search with 3-fold cross-validation, followed by one evaluation on the held-out test set; complete procedure repeated with three random seeds, reporting mean and standard deviation. GlycanGT large model pretrained with 35% masking provides [Graph] embeddings to SVM/LightGBM; the selected classifier for each S4 row is not identified. Section 2.6 states that all four graph baselines were trained and evaluated on the same datasets/splits; it does not establish that each baseline used the GlycanGT downstream classifier search. SVM search: 10 iterations, RBF/linear, C logU(1e-3,1e2), gamma logU(1e-4,1e-1); LightGBM: 15 randomized iterations. Do not combine with separate original GlycanML-paper protocols. Aggregation: Not reported GlycanGT published supplementary archive, Table S4; glycangt: Journal full-text XML · btag147_supplementary_data.zip / Table_S4.xlsx / Sheet1!A81:E81 (mean D81, SD E81) |
| Pipeline: GlycanGT | Protocol: GlycanML immunogenicity AUPRC: GlycanGT study: immunogenicity AUPRC Dataset subset: SugarBase immunogenicity; GlycanML official motif split (GlycanML split) | 0.84419033765294005 ± 0.0036259797661145998 auprc dimensionless · higher Uncertainty: type: standard_deviation; value: 0.0036259797661145998 Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · source checkedMethods, coverage and sourceGlycanGT on GlycanML immunogenicity AUPRC: GlycanGT study: immunogenicity AUPRC Immunogenicity: 1,320 glycans total; binary immune activity. Official fixed GlycanML motif-based train/validation/test splits (8:1:1). Section 2.5 reports class-balanced classifiers, train ∪ validation hyperparameter selection by randomized search with 3-fold cross-validation, followed by one evaluation on the held-out test set; complete procedure repeated with three random seeds, reporting mean and standard deviation. GlycanGT large model pretrained with 35% masking provides [Graph] embeddings to SVM/LightGBM; the selected classifier for each S4 row is not identified. Section 2.6 states that all four graph baselines were trained and evaluated on the same datasets/splits; it does not establish that each baseline used the GlycanGT downstream classifier search. SVM search: 10 iterations, RBF/linear, C logU(1e-3,1e2), gamma logU(1e-4,1e-1); LightGBM: 15 randomized iterations. Do not combine with separate original GlycanML-paper protocols. Aggregation: Not reported GlycanGT published supplementary archive, Table S4; glycangt: Journal full-text XML · btag147_supplementary_data.zip / Table_S4.xlsx / Sheet1!A83:E83 (mean D83, SD E83) |
| Pipeline: GlycanGT | Protocol: GlycanML immunogenicity Macro-F1: GlycanGT study: immunogenicity Macro-F1 Dataset subset: SugarBase immunogenicity; GlycanML official motif split (GlycanML split) | 0.87036334625073897 ± 0.065078625040751806 macro_f1 dimensionless · higher Uncertainty: type: standard_deviation; value: 0.065078625040751806 Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · source checkedMethods, coverage and sourceGlycanGT on GlycanML immunogenicity Macro-F1: GlycanGT study: immunogenicity Macro-F1 Immunogenicity: 1,320 glycans total; binary immune activity. Official fixed GlycanML motif-based train/validation/test splits (8:1:1). Section 2.5 reports class-balanced classifiers, train ∪ validation hyperparameter selection by randomized search with 3-fold cross-validation, followed by one evaluation on the held-out test set; complete procedure repeated with three random seeds, reporting mean and standard deviation. GlycanGT large model pretrained with 35% masking provides [Graph] embeddings to SVM/LightGBM; the selected classifier for each S4 row is not identified. Section 2.6 states that all four graph baselines were trained and evaluated on the same datasets/splits; it does not establish that each baseline used the GlycanGT downstream classifier search. SVM search: 10 iterations, RBF/linear, C logU(1e-3,1e2), gamma logU(1e-4,1e-1); LightGBM: 15 randomized iterations. Do not combine with separate original GlycanML-paper protocols. Aggregation: Not reported GlycanGT published supplementary archive, Table S4; glycangt: Journal full-text XML · btag147_supplementary_data.zip / Table_S4.xlsx / Sheet1!A82:E82 (mean D82, SD E82) |
| Pipeline: GlycanGT | Protocol: GlycanML kingdom Accuracy: GlycanGT study: kingdom Accuracy Dataset subset: SugarBase taxonomy kingdom; GlycanML official motif split (GlycanML split) | 0.90243017772941603 ± 0.017489392022112599 accuracy fraction · higher Uncertainty: type: standard_deviation; value: 0.017489392022112599 Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · source checkedMethods, coverage and sourceGlycanGT on GlycanML kingdom Accuracy: GlycanGT study: kingdom Accuracy Taxonomy: 13,209 glycans total across eight levels, 4–1,737 classes per level. Official fixed GlycanML motif-based train/validation/test splits (8:1:1). Section 2.5 reports class-balanced classifiers, train ∪ validation hyperparameter selection by randomized search with 3-fold cross-validation, followed by one evaluation on the held-out test set; complete procedure repeated with three random seeds, reporting mean and standard deviation. GlycanGT large model pretrained with 35% masking provides [Graph] embeddings to SVM/LightGBM; the selected classifier for each S4 row is not identified. Section 2.6 states that all four graph baselines were trained and evaluated on the same datasets/splits; it does not establish that each baseline used the GlycanGT downstream classifier search. SVM search: 10 iterations, RBF/linear, C logU(1e-3,1e2), gamma logU(1e-4,1e-1); LightGBM: 15 randomized iterations. Do not combine with separate original GlycanML-paper protocols. Aggregation: Not reported GlycanGT published supplementary archive, Table S4; glycangt: Journal full-text XML · btag147_supplementary_data.zip / Table_S4.xlsx / Sheet1!A67:E67 (mean D67, SD E67) |
| Pipeline: GlycanGT | Protocol: GlycanML kingdom Macro-F1: GlycanGT study: kingdom Macro-F1 Dataset subset: SugarBase taxonomy kingdom; GlycanML official motif split (GlycanML split) | 0.66958587101112299 ± 0.082979941298309295 macro_f1 dimensionless · higher Uncertainty: type: standard_deviation; value: 0.082979941298309295 Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · source checkedMethods, coverage and sourceGlycanGT on GlycanML kingdom Macro-F1: GlycanGT study: kingdom Macro-F1 Taxonomy: 13,209 glycans total across eight levels, 4–1,737 classes per level. Official fixed GlycanML motif-based train/validation/test splits (8:1:1). Section 2.5 reports class-balanced classifiers, train ∪ validation hyperparameter selection by randomized search with 3-fold cross-validation, followed by one evaluation on the held-out test set; complete procedure repeated with three random seeds, reporting mean and standard deviation. GlycanGT large model pretrained with 35% masking provides [Graph] embeddings to SVM/LightGBM; the selected classifier for each S4 row is not identified. Section 2.6 states that all four graph baselines were trained and evaluated on the same datasets/splits; it does not establish that each baseline used the GlycanGT downstream classifier search. SVM search: 10 iterations, RBF/linear, C logU(1e-3,1e2), gamma logU(1e-4,1e-1); LightGBM: 15 randomized iterations. Do not combine with separate original GlycanML-paper protocols. Aggregation: Not reported GlycanGT published supplementary archive, Table S4; glycangt: Journal full-text XML · btag147_supplementary_data.zip / Table_S4.xlsx / Sheet1!A68:E68 (mean D68, SD E68) |
| Pipeline: GlycanGT | Protocol: GlycanML order Accuracy: GlycanGT study: order Accuracy Dataset subset: SugarBase taxonomy order; GlycanML official motif split (GlycanML split) | 0.466086325716358 ± 0.020598144163221799 accuracy fraction · higher Uncertainty: type: standard_deviation; value: 0.020598144163221799 Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · source checkedMethods, coverage and sourceGlycanGT on GlycanML order Accuracy: GlycanGT study: order Accuracy Taxonomy: 13,209 glycans total across eight levels, 4–1,737 classes per level. Official fixed GlycanML motif-based train/validation/test splits (8:1:1). Section 2.5 reports class-balanced classifiers, train ∪ validation hyperparameter selection by randomized search with 3-fold cross-validation, followed by one evaluation on the held-out test set; complete procedure repeated with three random seeds, reporting mean and standard deviation. GlycanGT large model pretrained with 35% masking provides [Graph] embeddings to SVM/LightGBM; the selected classifier for each S4 row is not identified. Section 2.6 states that all four graph baselines were trained and evaluated on the same datasets/splits; it does not establish that each baseline used the GlycanGT downstream classifier search. SVM search: 10 iterations, RBF/linear, C logU(1e-3,1e2), gamma logU(1e-4,1e-1); LightGBM: 15 randomized iterations. Do not combine with separate original GlycanML-paper protocols. Aggregation: Not reported GlycanGT published supplementary archive, Table S4; glycangt: Journal full-text XML · btag147_supplementary_data.zip / Table_S4.xlsx / Sheet1!A73:E73 (mean D73, SD E73) |
| Pipeline: GlycanGT | Protocol: GlycanML order Macro-F1: GlycanGT study: order Macro-F1 Dataset subset: SugarBase taxonomy order; GlycanML official motif split (GlycanML split) | 0.278530744279541 ± 0.0086310922520820999 macro_f1 dimensionless · higher Uncertainty: type: standard_deviation; value: 0.0086310922520820999 Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · source checkedMethods, coverage and sourceGlycanGT on GlycanML order Macro-F1: GlycanGT study: order Macro-F1 Taxonomy: 13,209 glycans total across eight levels, 4–1,737 classes per level. Official fixed GlycanML motif-based train/validation/test splits (8:1:1). Section 2.5 reports class-balanced classifiers, train ∪ validation hyperparameter selection by randomized search with 3-fold cross-validation, followed by one evaluation on the held-out test set; complete procedure repeated with three random seeds, reporting mean and standard deviation. GlycanGT large model pretrained with 35% masking provides [Graph] embeddings to SVM/LightGBM; the selected classifier for each S4 row is not identified. Section 2.6 states that all four graph baselines were trained and evaluated on the same datasets/splits; it does not establish that each baseline used the GlycanGT downstream classifier search. SVM search: 10 iterations, RBF/linear, C logU(1e-3,1e2), gamma logU(1e-4,1e-1); LightGBM: 15 randomized iterations. Do not combine with separate original GlycanML-paper protocols. Aggregation: Not reported GlycanGT published supplementary archive, Table S4; glycangt: Journal full-text XML · btag147_supplementary_data.zip / Table_S4.xlsx / Sheet1!A74:E74 (mean D74, SD E74) |
| Pipeline: GlycanGT | Protocol: GlycanML phylum Accuracy: GlycanGT study: phylum Accuracy Dataset subset: SugarBase taxonomy phylum; GlycanML official motif split (GlycanML split) | 0.84149437794704296 ± 0.003324320417088 accuracy fraction · higher Uncertainty: type: standard_deviation; value: 0.003324320417088 Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · source checkedMethods, coverage and sourceGlycanGT on GlycanML phylum Accuracy: GlycanGT study: phylum Accuracy Taxonomy: 13,209 glycans total across eight levels, 4–1,737 classes per level. Official fixed GlycanML motif-based train/validation/test splits (8:1:1). Section 2.5 reports class-balanced classifiers, train ∪ validation hyperparameter selection by randomized search with 3-fold cross-validation, followed by one evaluation on the held-out test set; complete procedure repeated with three random seeds, reporting mean and standard deviation. GlycanGT large model pretrained with 35% masking provides [Graph] embeddings to SVM/LightGBM; the selected classifier for each S4 row is not identified. Section 2.6 states that all four graph baselines were trained and evaluated on the same datasets/splits; it does not establish that each baseline used the GlycanGT downstream classifier search. SVM search: 10 iterations, RBF/linear, C logU(1e-3,1e2), gamma logU(1e-4,1e-1); LightGBM: 15 randomized iterations. Do not combine with separate original GlycanML-paper protocols. Aggregation: Not reported GlycanGT published supplementary archive, Table S4; glycangt: Journal full-text XML · btag147_supplementary_data.zip / Table_S4.xlsx / Sheet1!A69:E69 (mean D69, SD E69) |
| Pipeline: GlycanGT | Protocol: GlycanML phylum Macro-F1: GlycanGT study: phylum Macro-F1 Dataset subset: SugarBase taxonomy phylum; GlycanML official motif split (GlycanML split) | 0.49333063993822701 ± 0.080739293560509295 macro_f1 dimensionless · higher Uncertainty: type: standard_deviation; value: 0.080739293560509295 Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · source checkedMethods, coverage and sourceGlycanGT on GlycanML phylum Macro-F1: GlycanGT study: phylum Macro-F1 Taxonomy: 13,209 glycans total across eight levels, 4–1,737 classes per level. Official fixed GlycanML motif-based train/validation/test splits (8:1:1). Section 2.5 reports class-balanced classifiers, train ∪ validation hyperparameter selection by randomized search with 3-fold cross-validation, followed by one evaluation on the held-out test set; complete procedure repeated with three random seeds, reporting mean and standard deviation. GlycanGT large model pretrained with 35% masking provides [Graph] embeddings to SVM/LightGBM; the selected classifier for each S4 row is not identified. Section 2.6 states that all four graph baselines were trained and evaluated on the same datasets/splits; it does not establish that each baseline used the GlycanGT downstream classifier search. SVM search: 10 iterations, RBF/linear, C logU(1e-3,1e2), gamma logU(1e-4,1e-1); LightGBM: 15 randomized iterations. Do not combine with separate original GlycanML-paper protocols. Aggregation: Not reported GlycanGT published supplementary archive, Table S4; glycangt: Journal full-text XML · btag147_supplementary_data.zip / Table_S4.xlsx / Sheet1!A70:E70 (mean D70, SD E70) |
| Pipeline: GlycanGT | Protocol: GlycanML species Accuracy: GlycanGT study: species Accuracy Dataset subset: SugarBase taxonomy species; GlycanML official motif split (GlycanML split) | 0.39173014145810597 ± 0.021681021594310401 accuracy fraction · higher Uncertainty: type: standard_deviation; value: 0.021681021594310401 Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · source checkedMethods, coverage and sourceGlycanGT on GlycanML species Accuracy: GlycanGT study: species Accuracy Taxonomy: 13,209 glycans total across eight levels, 4–1,737 classes per level. Official fixed GlycanML motif-based train/validation/test splits (8:1:1). Section 2.5 reports class-balanced classifiers, train ∪ validation hyperparameter selection by randomized search with 3-fold cross-validation, followed by one evaluation on the held-out test set; complete procedure repeated with three random seeds, reporting mean and standard deviation. GlycanGT large model pretrained with 35% masking provides [Graph] embeddings to SVM/LightGBM; the selected classifier for each S4 row is not identified. Section 2.6 states that all four graph baselines were trained and evaluated on the same datasets/splits; it does not establish that each baseline used the GlycanGT downstream classifier search. SVM search: 10 iterations, RBF/linear, C logU(1e-3,1e2), gamma logU(1e-4,1e-1); LightGBM: 15 randomized iterations. Do not combine with separate original GlycanML-paper protocols. Aggregation: Not reported GlycanGT published supplementary archive, Table S4; glycangt: Journal full-text XML · btag147_supplementary_data.zip / Table_S4.xlsx / Sheet1!A79:E79 (mean D79, SD E79) |
| Pipeline: GlycanGT | Protocol: GlycanML species Macro-F1: GlycanGT study: species Macro-F1 Dataset subset: SugarBase taxonomy species; GlycanML official motif split (GlycanML split) | 0.184571928631623 ± 0.00080738146274129995 macro_f1 dimensionless · higher Uncertainty: type: standard_deviation; value: 0.00080738146274129995 Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · source checkedMethods, coverage and sourceGlycanGT on GlycanML species Macro-F1: GlycanGT study: species Macro-F1 Taxonomy: 13,209 glycans total across eight levels, 4–1,737 classes per level. Official fixed GlycanML motif-based train/validation/test splits (8:1:1). Section 2.5 reports class-balanced classifiers, train ∪ validation hyperparameter selection by randomized search with 3-fold cross-validation, followed by one evaluation on the held-out test set; complete procedure repeated with three random seeds, reporting mean and standard deviation. GlycanGT large model pretrained with 35% masking provides [Graph] embeddings to SVM/LightGBM; the selected classifier for each S4 row is not identified. Section 2.6 states that all four graph baselines were trained and evaluated on the same datasets/splits; it does not establish that each baseline used the GlycanGT downstream classifier search. SVM search: 10 iterations, RBF/linear, C logU(1e-3,1e2), gamma logU(1e-4,1e-1); LightGBM: 15 randomized iterations. Do not combine with separate original GlycanML-paper protocols. Aggregation: Not reported GlycanGT published supplementary archive, Table S4; glycangt: Journal full-text XML · btag147_supplementary_data.zip / Table_S4.xlsx / Sheet1!A80:E80 (mean D80, SD E80) |
Source checking is not independent reproduction. Release 2026-09-23-2b89723c6dd9.
Use this model
How it works, versions and access
Underlying model: GlycanGT. Results on this page belong to this pipeline and its evaluated settings.
Strengths, limitations and unresolved questions
Evidence
Source checking verifies the cited claim or transcription. It does not establish independent reproduction.
Evidence table
Inspect claims, sources and review details
Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.
One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.
2 evidence rows matching the loaded filters
| Property and statement | Original source and location | Review and provenance |
|---|---|---|
| Relationship: uses model discovery-model-glycangt Individual claims | glycangt: Journal full-text XML GlycanGT primary article Sections 2.1, 2.5, 2.6 and 3.1–3.2 (PMC13105845), Supplementary Table S4; column B GlycanGT Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: Primary article XML snapshot | source checked automated source review · 2026-09-23 Audit detailsSource-backed evaluated identity only; no independent reproduction. Field: Claim: glycangt-2026-table-s4-method-glycangt-discovery-model-glycangt-identity-claim Source artifact SHA-256: Hash scope: SHA-256 of retrieved original artifact bytes Format: original_artifact |
| Relationship: uses model discovery-model-glycangt Individual claims | GlycanGT published supplementary archive, Table S4 GlycanGT primary article Sections 2.1, 2.5, 2.6 and 3.1–3.2 (PMC13105845), Supplementary Table S4; column B GlycanGT Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: Published Bioinformatics btag147 supplementary archive, retrieved 2026-09-23; Table_S4.xlsx SHA-256 d7c35909bac6bcb78ca8fdb32c0463f05e692f15861a8184e65e415e7216f493 | source checked automated source review · 2026-09-23 Audit detailsSource-backed evaluated identity only; no independent reproduction. Field: Claim: glycangt-2026-table-s4-method-glycangt-discovery-model-glycangt-identity-claim Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
Sources and history
View linked audit checks and correction history
Release 2026-09-23-2b89723c6dd9 · Record review: source checked
2 source records and release history
- GlycanGT published supplementary archive, Table S4 · Original source · Published Bioinformatics btag147 supplementary archive, retrieved 2026-09-23; Table_S4.xlsx SHA-256 d7c35909bac6bcb78ca8fdb32c0463f05e692f15861a8184e65e415e7216f493
- glycangt: Journal full-text XML · Original source · Primary article XML snapshot
Technical metadata and extraction receipts
Stable ID: glycangt-2026-table-s4-method-glycangt
- areas
- glycans
- source locator
- GlycanGT primary article Sections 2.1, 2.5, 2.6 and 3.1–3.2 (PMC13105845), Supplementary Table S4; column B GlycanGT
- missing metadata
- checkpoint revision: unreported; parameters: unextracted
Related records
- uses model: GlycanGT
- model: GlycanGT on GlycanML class Accuracy: GlycanGT study: class Accuracy
- model: GlycanGT on GlycanML class Macro-F1: GlycanGT study: class Macro-F1
- model: GlycanGT on GlycanML domain Accuracy: GlycanGT study: domain Accuracy
- model: GlycanGT on GlycanML domain Macro-F1: GlycanGT study: domain Macro-F1
- model: GlycanGT on GlycanML family Accuracy: GlycanGT study: family Accuracy
- model: GlycanGT on GlycanML family Macro-F1: GlycanGT study: family Macro-F1
- model: GlycanGT on GlycanML genus Accuracy: GlycanGT study: genus Accuracy
- model: GlycanGT on GlycanML genus Macro-F1: GlycanGT study: genus Macro-F1
- model: GlycanGT on GlycanML glycosylation Accuracy: GlycanGT study: glycosylation Accuracy
- model: GlycanGT on GlycanML glycosylation Macro-F1: GlycanGT study: glycosylation Macro-F1
- model: GlycanGT on GlycanML immunogenicity Accuracy: GlycanGT study: immunogenicity Accuracy
- model: GlycanGT on GlycanML immunogenicity AUPRC: GlycanGT study: immunogenicity AUPRC
- model: GlycanGT on GlycanML immunogenicity Macro-F1: GlycanGT study: immunogenicity Macro-F1
- model: GlycanGT on GlycanML kingdom Accuracy: GlycanGT study: kingdom Accuracy
- model: GlycanGT on GlycanML kingdom Macro-F1: GlycanGT study: kingdom Macro-F1
- model: GlycanGT on GlycanML order Accuracy: GlycanGT study: order Accuracy
- model: GlycanGT on GlycanML order Macro-F1: GlycanGT study: order Macro-F1
- model: GlycanGT on GlycanML phylum Accuracy: GlycanGT study: phylum Accuracy
- model: GlycanGT on GlycanML phylum Macro-F1: GlycanGT study: phylum Macro-F1
- model: GlycanGT on GlycanML species Accuracy: GlycanGT study: species Accuracy
- model: GlycanGT on GlycanML species Macro-F1: GlycanGT study: species Macro-F1
- subject: GlycanGT: uses model discovery-model-glycangt