rewire.it
Task

clathrin protein classification

Clathrin classification uses cross-validation and multiple independent benchmark datasets with different redundancy filters.

SourcesAdvancing the accuracy of clathrin protein prediction through multi-source protein language models · Methods: Dataset construction; Performance evaluation; Results: independent test datasets; cached text lines 9–10, 24–25, 54

14 evaluations · 79 metric rows

At a glance

Inputs, training, access and other details

Explanatory profile: limited source coverage · Automated source review, 2026-09-16. Review applies to the cited claims; unresolved fields are listed below. Numerical results retain their own review status.

Data, procedure and scoring
PropertyDescription and evidence
DatasetsLe2019 and Zhang2020-derived clathrin/non-clathrin sequence collections.
SourcesAdvancing the accuracy of clathrin protein prediction through multi-source protein language models · Methods: Dataset construction; Performance evaluation; Results: independent test datasets; cached text lines 9–10, 24–25, 54
SplitsTen-fold cross-validation on training data plus named independent CLA test sets.
SourcesAdvancing the accuracy of clathrin protein prediction through multi-source protein language models · Methods: Dataset construction; Performance evaluation; Results: independent test datasets; cached text lines 9–10, 24–25, 54
MetricsAccuracy, AUC, Matthews correlation coefficient, F1, sensitivity and specificity are reported. The Performance evaluation subsection describes ten-fold evaluation and early stopping but does not specify a universal probability threshold for every comparator.
SourcesAdvancing the accuracy of clathrin protein prediction through multi-source protein language models · Materials and methods: Performance evaluation
Baselinesdeep-clathrin includes literature-reported scores; DeepCLA is reimplemented for comparison.
SourcesAdvancing the accuracy of clathrin protein prediction through multi-source protein language models · Methods: Dataset construction; Performance evaluation; Results: independent test datasets; cached text lines 9–10, 24–25, 54
Leakage controlsThe Zhang2020-derived collection uses BLAST redundancy filtering at 0.7, and the new dataset uses CD-HIT at 0.6. These dataset filters do not establish absence from the pretrained protein encoders’ corpora. Feature subsets are evaluated on both cross-validation and independent-test performance, which limits treating that test as untouched model selection.
SourcesAdvancing the accuracy of clathrin protein prediction through multi-source protein language models · Materials and methods: Dataset construction and Overall framework of PLM-CLA; Results: The effect of feature selection methods on the predictive performance
UncertaintyThe cited text-accessible evaluation sections give no confidence-interval, resampling or repeat-run error-bar specification. Image-only tables and uninspected supplements are outside this absence claim. · Not reported in inspected sources
SourcesAdvancing the accuracy of clathrin protein prediction through multi-source protein language models · Methods: Dataset construction; Performance evaluation; Results: independent test datasets; cached text lines 9–10, 24–25, 54
Entity typePaper-specific computational evaluation protocol.
SourcesAdvancing the accuracy of clathrin protein prediction through multi-source protein language models · Methods: Dataset construction; Performance evaluation; Results: independent test datasets; cached text lines 9–10, 24–25, 54
OrganismsThe paper does not enumerate organisms. Its pinned released CSVs contain sequence identifiers and amino-acid sequences, without organism columns; one dataset uses synthetic row identifiers. Consequently, a complete source-species inventory is not established by the supplied benchmark metadata. · Not reported in inspected sources
Sources (5)Advancing the accuracy of clathrin protein prediction through multi-source protein language models; clathrin__Dataset__Clathrin0.6.csv; clathrin__Dataset__Clathrin0.7.csv; clathrin__Dataset__Clathrin1.0.csv; clathrin__README.md · Dataset construction; pinned repository README and Dataset/Clathrin0.6.csv, Clathrin0.7.csv and Clathrin1.0.csv headers
AssaysClathrin/non-clathrin sequence labels.
SourcesAdvancing the accuracy of clathrin protein prediction through multi-source protein language models · Methods: Dataset construction; Performance evaluation; Results: independent test datasets; cached text lines 9–10, 24–25, 54
Allowed inputsProtein amino-acid sequences represented by protein language models.
SourcesAdvancing the accuracy of clathrin protein prediction through multi-source protein language models · Methods: Dataset construction; Performance evaluation; Results: independent test datasets; cached text lines 9–10, 24–25, 54
AdaptationSupervised classification with cross-validation and independent test collections.
SourcesAdvancing the accuracy of clathrin protein prediction through multi-source protein language models · Methods: Dataset construction; Performance evaluation; Results: independent test datasets; cached text lines 9–10, 24–25, 54

How it works

How it worksComputational evaluation flow
Computational evaluation flow1. Input: Protein amino-acid sequences represented by protein language models.. Then: 2. Evaluation: Ten-fold cross-validation on training data plus named independent CLA test sets.. Then: 3. Readout: Accuracy, AUC, Matthews correlation coefficient, F1, sensitivity and specificity are reported. The Performance evaluation subsection describes ten-fold evaluation and early stopping but does not specify a universal probability threshold for every comparator.Computational evaluation flow1. Input: Protein amino-acid sequences represented by protein language models.. Then: 2. Evaluation: Ten-fold cross-validation on training data plus named independent CLA test sets.. Then: 3. Readout: Accuracy, AUC, Matthews correlation coefficient, F1, sensitivity and specificity are reported. The Performance evaluation subsection describes ten-fold evaluation and early stopping but does not specify a universal probability threshold for every comparator.Computational evaluation flow1. Input: Protein amino-acid sequences represented by protein language models.. Then: 2. Evaluation: Ten-fold cross-validation on training data plus named independent CLA test sets.. Then: 3. Readout: Accuracy, AUC, Matthews correlation coefficient, F1, sensitivity and specificity are reported. The Performance evaluation subsection describes ten-fold evaluation and early stopping but does not specify a universal probability threshold for every comparator.

Conceptual summary of the cited evaluation; exact task configuration and source version remain part of the protocol.

SourcesAdvancing the accuracy of clathrin protein prediction through multi-source protein language models · Methods: Dataset construction; Performance evaluation; Results: independent test datasets; cached text lines 9–10, 24–25, 54; Materials and methods: Performance evaluation
Evaluation methodology

Le2019 and Zhang2020-derived clathrin/non-clathrin sequence collections. Ten-fold cross-validation on training data plus named independent CLA test sets. Accuracy, AUC, Matthews correlation coefficient, F1, sensitivity and specificity are reported. The Performance evaluation subsection describes ten-fold evaluation and early stopping but does not specify a universal probability threshold for every comparator. deep-clathrin includes literature-reported scores; DeepCLA is reimplemented for comparison. The Zhang2020-derived collection uses BLAST redundancy filtering at 0.7, and the new dataset uses CD-HIT at 0.6. These dataset filters do not establish absence from the pretrained protein encoders’ corpora. Feature subsets are evaluated on both cross-validation and independent-test performance, which limits treating that test as untouched model selection.

SourcesAdvancing the accuracy of clathrin protein prediction through multi-source protein language models · Methods: Dataset construction; Performance evaluation; Results: independent test datasets; cached text lines 9–10, 24–25, 54; Materials and methods: Performance evaluation; Materials and methods: Dataset construction and Overall framework of PLM-CLA; Results: The effect of feature selection methods on the predictive performance

Evaluation design

Benchmarks bring together tasks and protocols. A task describes the biological question; a protocol defines a particular test.

These source-backed links do not make different protocols or scores interchangeable.

Recorded evaluations

Each evaluation records what was tested and under which conditions.

Published comparisons

Explore the results reported under one evaluation protocol. Each figure keeps its source, dataset and metric together; it is not a ranking across studies.

Clathrin independent test: selected-embedding classifiers · Table 3

ACC (fraction) · Higher values are better for this metric.

Conventional classifiers and PLM-CLA compared using paper-selected features. CLA-IND 0.6 independent test from Shoombuatong 2024 dataset; source Table 1 carries cohort counts.

Evaluation protocol · Clathrin independent test: selected-embedding classifiers

  1. DT · Configuration · Author-reported evaluation0.715
  2. NB · Configuration · Author-reported evaluation0.827
  3. PLS · Configuration · Author-reported evaluation0.849
  4. ADA · Configuration · Author-reported evaluation0.877
  5. LDA · Configuration · Author-reported evaluation0.866
  6. RF · Configuration · Author-reported evaluation0.894
  7. LR · Configuration · Author-reported evaluation0.877
  8. KNN · Configuration · Author-reported evaluation0.916
  9. ET · Configuration · Author-reported evaluation0.888
  10. XGB · Configuration · Author-reported evaluation0.877
  11. MLP · Configuration · Author-reported evaluation0.933
  12. SVM · Configuration · Author-reported evaluation0.927
  13. PLM-CLA · Configuration · Author-reported evaluation0.961

Source order is preserved. Plotted marks show point estimates; uncertainty, where reported, is retained in the printed values and table. Differences do not establish statistical significance.

Advancing the accuracy of clathrin protein prediction through multi-source protein language models · Table 3: ACC, Clathrin independent test: selected-embedding classifiers
Values, uncertainty and evidence
ACC: original source values
Tested entityPrinted valueUncertaintyEvidence
DT · Configuration0.715 fractionNot reportedAuthor-reported evaluation · source checkedAdvancing the accuracy of clathrin protein prediction through multi-source protein language models · Table 3, row DT, column ACC; XML row2 column2
NB · Configuration0.827 fractionNot reportedAuthor-reported evaluation · source checkedAdvancing the accuracy of clathrin protein prediction through multi-source protein language models · Table 3, row NB, column ACC; XML row3 column2
PLS · Configuration0.849 fractionNot reportedAuthor-reported evaluation · source checkedAdvancing the accuracy of clathrin protein prediction through multi-source protein language models · Table 3, row PLS, column ACC; XML row4 column2
ADA · Configuration0.877 fractionNot reportedAuthor-reported evaluation · source checkedAdvancing the accuracy of clathrin protein prediction through multi-source protein language models · Table 3, row ADA, column ACC; XML row5 column2
LDA · Configuration0.866 fractionNot reportedAuthor-reported evaluation · source checkedAdvancing the accuracy of clathrin protein prediction through multi-source protein language models · Table 3, row LDA, column ACC; XML row6 column2
RF · Configuration0.894 fractionNot reportedAuthor-reported evaluation · source checkedAdvancing the accuracy of clathrin protein prediction through multi-source protein language models · Table 3, row RF, column ACC; XML row7 column2
LR · Configuration0.877 fractionNot reportedAuthor-reported evaluation · source checkedAdvancing the accuracy of clathrin protein prediction through multi-source protein language models · Table 3, row LR, column ACC; XML row8 column2
KNN · Configuration0.916 fractionNot reportedAuthor-reported evaluation · source checkedAdvancing the accuracy of clathrin protein prediction through multi-source protein language models · Table 3, row KNN, column ACC; XML row9 column2
ET · Configuration0.888 fractionNot reportedAuthor-reported evaluation · source checkedAdvancing the accuracy of clathrin protein prediction through multi-source protein language models · Table 3, row ET, column ACC; XML row10 column2
XGB · Configuration0.877 fractionNot reportedAuthor-reported evaluation · source checkedAdvancing the accuracy of clathrin protein prediction through multi-source protein language models · Table 3, row XGB, column ACC; XML row11 column2
MLP · Configuration0.933 fractionNot reportedAuthor-reported evaluation · source checkedAdvancing the accuracy of clathrin protein prediction through multi-source protein language models · Table 3, row MLP, column ACC; XML row12 column2
SVM · Configuration0.927 fractionNot reportedAuthor-reported evaluation · source checkedAdvancing the accuracy of clathrin protein prediction through multi-source protein language models · Table 3, row SVM, column ACC; XML row13 column2
PLM-CLA · Configuration0.961 fractionNot reportedAuthor-reported evaluation · source checkedAdvancing the accuracy of clathrin protein prediction through multi-source protein language models · Table 3, row PLM-CLA, column ACC; XML row14 column2
Scope and limitations
  • Source prose has inconsistent dataset sizes; exact cohort counts left unresolved, not inferred.
  • Feature-selection and evaluation leakage cannot be excluded by this table alone.
  • No interval assigned unless printed in source cell.

Source transcription and grouping reviewed by automated source review on 2026-09-17. These experiments were not independently reproduced by rewire.

Tested entities and results

Release 2026-09-17-d277315f7d76 · 14 evaluations · 79 metric rows. Different protocols are not a single leaderboard.

Results grouped by the exact reported evaluation
Metric and findingCoverage and uncertaintyEvidence
ESM-2 embedding + paper classifier: clathrin protein classification

single-feature ESM-2 embedding comparison

Independent external evaluation · Evaluation metadata: needs review

0.916 accuracy

Unit: fraction · Direction: unknown

Uncertainty: not reported in legacy extract

Scored: Not reported · Eligible: Not reported

source checkedAdvancing the accuracy of clathrin protein prediction through multi-source protein language models · Table 2, Independent test / ESM-2 row, ACC column

Source checking is not independent reproduction.

DT: Clathrin independent test: selected-embedding classifiers

Conventional classifiers and PLM-CLA compared using paper-selected features. CLA-IND 0.6 independent test from Shoombuatong 2024 dataset; source Table 1 carries cohort counts.

Author-reported evaluation · Evaluation metadata: needs review

0.715 ACC

Unit: fraction · Direction: higher

Uncertainty: unreported

Scored: Not reported · Eligible: Not reported

source checkedAdvancing the accuracy of clathrin protein prediction through multi-source protein language models · Table 3, row DT, column ACC; XML row2 column2

Source checking is not independent reproduction.

0.551 SP

Unit: fraction · Direction: higher

Uncertainty: unreported

Scored: Not reported · Eligible: Not reported

source checkedAdvancing the accuracy of clathrin protein prediction through multi-source protein language models · Table 3, row DT, column SP; XML row2 column4

Source checking is not independent reproduction.

0.726 AUC

Unit: fraction · Direction: higher

Uncertainty: unreported

Scored: Not reported · Eligible: Not reported

source checkedAdvancing the accuracy of clathrin protein prediction through multi-source protein language models · Table 3, row DT, column AUC; XML row2 column7

Source checking is not independent reproduction.

RF: Clathrin independent test: selected-embedding classifiers

Conventional classifiers and PLM-CLA compared using paper-selected features. CLA-IND 0.6 independent test from Shoombuatong 2024 dataset; source Table 1 carries cohort counts.

Author-reported evaluation · Evaluation metadata: needs review

0.884 SP

Unit: fraction · Direction: higher

Uncertainty: unreported

Scored: Not reported · Eligible: Not reported

source checkedAdvancing the accuracy of clathrin protein prediction through multi-source protein language models · Table 3, row RF, column SP; XML row7 column4

Source checking is not independent reproduction.

0.912 F1

Unit: fraction · Direction: higher

Uncertainty: unreported

Scored: Not reported · Eligible: Not reported

source checkedAdvancing the accuracy of clathrin protein prediction through multi-source protein language models · Table 3, row RF, column F1; XML row7 column6

Source checking is not independent reproduction.

0.894 ACC

Unit: fraction · Direction: higher

Uncertainty: unreported

Scored: Not reported · Eligible: Not reported

source checkedAdvancing the accuracy of clathrin protein prediction through multi-source protein language models · Table 3, row RF, column ACC; XML row7 column2

Source checking is not independent reproduction.

0.959 AUC

Unit: fraction · Direction: higher

Uncertainty: unreported

Scored: Not reported · Eligible: Not reported

source checkedAdvancing the accuracy of clathrin protein prediction through multi-source protein language models · Table 3, row RF, column AUC; XML row7 column7

Source checking is not independent reproduction.

PLM-CLA: Clathrin independent test: selected-embedding classifiers

Conventional classifiers and PLM-CLA compared using paper-selected features. CLA-IND 0.6 independent test from Shoombuatong 2024 dataset; source Table 1 carries cohort counts.

Author-reported evaluation · Evaluation metadata: needs review

0.917 MCC

Unit: dimensionless · Direction: higher

Uncertainty: unreported

Scored: Not reported · Eligible: Not reported

source checkedAdvancing the accuracy of clathrin protein prediction through multi-source protein language models · Table 3, row PLM-CLA, column MCC; XML row14 column5

Source checking is not independent reproduction.

0.949 F1

Unit: fraction · Direction: higher

Uncertainty: unreported

Scored: Not reported · Eligible: Not reported

source checkedAdvancing the accuracy of clathrin protein prediction through multi-source protein language models · Table 3, row PLM-CLA, column F1; XML row14 column6

Source checking is not independent reproduction.

0.961 ACC

Unit: fraction · Direction: higher

Uncertainty: unreported

Scored: Not reported · Eligible: Not reported

source checkedAdvancing the accuracy of clathrin protein prediction through multi-source protein language models · Table 3, row PLM-CLA, column ACC; XML row14 column2

Source checking is not independent reproduction.

PLS: Clathrin independent test: selected-embedding classifiers

Conventional classifiers and PLM-CLA compared using paper-selected features. CLA-IND 0.6 independent test from Shoombuatong 2024 dataset; source Table 1 carries cohort counts.

Author-reported evaluation · Evaluation metadata: needs review

0.690 MCC

Unit: dimensionless · Direction: higher

Uncertainty: unreported

Scored: Not reported · Eligible: Not reported

source checkedAdvancing the accuracy of clathrin protein prediction through multi-source protein language models · Table 3, row PLS, column MCC; XML row4 column5

Source checking is not independent reproduction.

0.845 SN

Unit: fraction · Direction: higher

Uncertainty: unreported

Scored: Not reported · Eligible: Not reported

source checkedAdvancing the accuracy of clathrin protein prediction through multi-source protein language models · Table 3, row PLS, column SN; XML row4 column3

Source checking is not independent reproduction.

LR: Clathrin independent test: selected-embedding classifiers

Conventional classifiers and PLM-CLA compared using paper-selected features. CLA-IND 0.6 independent test from Shoombuatong 2024 dataset; source Table 1 carries cohort counts.

Author-reported evaluation · Evaluation metadata: needs review

0.747 MCC

Unit: dimensionless · Direction: higher

Uncertainty: unreported

Scored: Not reported · Eligible: Not reported

source checkedAdvancing the accuracy of clathrin protein prediction through multi-source protein language models · Table 3, row LR, column MCC; XML row8 column5

Source checking is not independent reproduction.

0.884 SP

Unit: fraction · Direction: higher

Uncertainty: unreported

Scored: Not reported · Eligible: Not reported

source checkedAdvancing the accuracy of clathrin protein prediction through multi-source protein language models · Table 3, row LR, column SP; XML row8 column4

Source checking is not independent reproduction.

ADA: Clathrin independent test: selected-embedding classifiers

Conventional classifiers and PLM-CLA compared using paper-selected features. CLA-IND 0.6 independent test from Shoombuatong 2024 dataset; source Table 1 carries cohort counts.

Author-reported evaluation · Evaluation metadata: needs review

0.877 ACC

Unit: fraction · Direction: higher

Uncertainty: unreported

Scored: Not reported · Eligible: Not reported

source checkedAdvancing the accuracy of clathrin protein prediction through multi-source protein language models · Table 3, row ADA, column ACC; XML row5 column2

Source checking is not independent reproduction.

0.739 MCC

Unit: dimensionless · Direction: higher

Uncertainty: unreported

Scored: Not reported · Eligible: Not reported

source checkedAdvancing the accuracy of clathrin protein prediction through multi-source protein language models · Table 3, row ADA, column MCC; XML row5 column5

Source checking is not independent reproduction.

0.918 SN

Unit: fraction · Direction: higher

Uncertainty: unreported

Scored: Not reported · Eligible: Not reported

source checkedAdvancing the accuracy of clathrin protein prediction through multi-source protein language models · Table 3, row ADA, column SN; XML row5 column3

Source checking is not independent reproduction.

NB: Clathrin independent test: selected-embedding classifiers

Conventional classifiers and PLM-CLA compared using paper-selected features. CLA-IND 0.6 independent test from Shoombuatong 2024 dataset; source Table 1 carries cohort counts.

Author-reported evaluation · Evaluation metadata: needs review

0.773 SN

Unit: fraction · Direction: higher

Uncertainty: unreported

Scored: Not reported · Eligible: Not reported

source checkedAdvancing the accuracy of clathrin protein prediction through multi-source protein language models · Table 3, row NB, column SN; XML row3 column3

Source checking is not independent reproduction.

0.868 AUC

Unit: fraction · Direction: higher

Uncertainty: unreported

Scored: Not reported · Eligible: Not reported

source checkedAdvancing the accuracy of clathrin protein prediction through multi-source protein language models · Table 3, row NB, column AUC; XML row3 column7

Source checking is not independent reproduction.

SVM: Clathrin independent test: selected-embedding classifiers

Conventional classifiers and PLM-CLA compared using paper-selected features. CLA-IND 0.6 independent test from Shoombuatong 2024 dataset; source Table 1 carries cohort counts.

Author-reported evaluation · Evaluation metadata: needs review

0.942 F1

Unit: fraction · Direction: higher

Uncertainty: unreported

Scored: Not reported · Eligible: Not reported

source checkedAdvancing the accuracy of clathrin protein prediction through multi-source protein language models · Table 3, row SVM, column F1; XML row13 column6

Source checking is not independent reproduction.

0.964 SN

Unit: fraction · Direction: higher

Uncertainty: unreported

Scored: Not reported · Eligible: Not reported

source checkedAdvancing the accuracy of clathrin protein prediction through multi-source protein language models · Table 3, row SVM, column SN; XML row13 column3

Source checking is not independent reproduction.

MLP: Clathrin independent test: selected-embedding classifiers

Conventional classifiers and PLM-CLA compared using paper-selected features. CLA-IND 0.6 independent test from Shoombuatong 2024 dataset; source Table 1 carries cohort counts.

Author-reported evaluation · Evaluation metadata: needs review

0.946 F1

Unit: fraction · Direction: higher

Uncertainty: unreported

Scored: Not reported · Eligible: Not reported

source checkedAdvancing the accuracy of clathrin protein prediction through multi-source protein language models · Table 3, row MLP, column F1; XML row12 column6

Source checking is not independent reproduction.

ET: Clathrin independent test: selected-embedding classifiers

Conventional classifiers and PLM-CLA compared using paper-selected features. CLA-IND 0.6 independent test from Shoombuatong 2024 dataset; source Table 1 carries cohort counts.

Author-reported evaluation · Evaluation metadata: needs review

0.764 MCC

Unit: dimensionless · Direction: higher

Uncertainty: unreported

Scored: Not reported · Eligible: Not reported

source checkedAdvancing the accuracy of clathrin protein prediction through multi-source protein language models · Table 3, row ET, column MCC; XML row10 column5

Source checking is not independent reproduction.

XGB: Clathrin independent test: selected-embedding classifiers

Conventional classifiers and PLM-CLA compared using paper-selected features. CLA-IND 0.6 independent test from Shoombuatong 2024 dataset; source Table 1 carries cohort counts.

Author-reported evaluation · Evaluation metadata: needs review

0.877 ACC

Unit: fraction · Direction: higher

Uncertainty: unreported

Scored: Not reported · Eligible: Not reported

source checkedAdvancing the accuracy of clathrin protein prediction through multi-source protein language models · Table 3, row XGB, column ACC; XML row11 column2

Source checking is not independent reproduction.

Papers and result coverage

Last literature check: 2026-09-17. Primary-source discovery and table/protocol screening; source checked is not independently reproduced. Raw acquisitions not automatically numerical publication approval.

Paper or primary resourceVersionReference
Advancing the accuracy of clathrin protein prediction through multi-source protein language modelsjournal full text in PMCRead source
DOI: 10.1038/s41598-025-08510-4

What is still missing

  • Independent batch review before import; preserve existing observation identities.
Search and extraction details

complete comparison tables extracted pending publication review

Searches

  • Advancing the accuracy of clathrin protein prediction through multi-source protein language models primary paper benchmark results

Evidence locations

  • Tables2–5; independent-test and dataset-construction sections

Strengths and limitations

Strengths and considerations

No source-reviewed explanatory claims are recorded here yet.

Profile review details

Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.

Stable record: reported-task-786c09824e9bf5

Evidence table

Inspect claims, sources and review details

Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.

One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.

21 evidence rows matching the loaded filters

Claims, original sources and review scope · Release 2026-09-17-d277315f7d76
Property and statementOriginal source and locationReview and provenance
Diagram caption

Conceptual summary of the cited evaluation; exact task configuration and source version remain part of the protocol.

Individual claims
Advancing the accuracy of clathrin protein prediction through multi-source protein language models

Original source ↗

Methods: Dataset construction; Performance evaluation; Results: independent test datasets; cached text lines 9–10, 24–25, 54; Materials and methods: Performance evaluation

Version: journal full text in PMC
Retrieved: 2026-09-16T10:38:57.558194+00:00

source checked

automated source review · 2026-09-16

Audit details

Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.

Field: attributes.profile.diagram.caption

Source artifact SHA-256: 2edc86b25707c1b737d26117093ce8d856e79cc5d0b335f27c1c341f887f1c7e

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Diagram steps

["Input: Protein amino-acid sequences represented by protein language models.","Evaluation: Ten-fold cross-validation on training data plus named independent CLA test sets.","Readout: Accuracy, AUC, Matthews correlation coefficient, F1, sensitivity and specificity are reported. The Performance evaluation subsection describes ten-fold evaluation and early stopping but does not specify a universal probability threshold for every comparator."]

Individual claims
Advancing the accuracy of clathrin protein prediction through multi-source protein language models

Original source ↗

Methods: Dataset construction; Performance evaluation; Results: independent test datasets; cached text lines 9–10, 24–25, 54; Materials and methods: Performance evaluation

Version: journal full text in PMC
Retrieved: 2026-09-16T10:38:57.558194+00:00

source checked

automated source review · 2026-09-16

Audit details

Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.

Field: attributes.profile.diagram.steps

Source artifact SHA-256: 2edc86b25707c1b737d26117093ce8d856e79cc5d0b335f27c1c341f887f1c7e

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Diagram title

Computational evaluation flow

Individual claims
Advancing the accuracy of clathrin protein prediction through multi-source protein language models

Original source ↗

Methods: Dataset construction; Performance evaluation; Results: independent test datasets; cached text lines 9–10, 24–25, 54; Materials and methods: Performance evaluation

Version: journal full text in PMC
Retrieved: 2026-09-16T10:38:57.558194+00:00

source checked

automated source review · 2026-09-16

Audit details

Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.

Field: attributes.profile.diagram.title

Source artifact SHA-256: 2edc86b25707c1b737d26117093ce8d856e79cc5d0b335f27c1c341f887f1c7e

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Datasets

Le2019 and Zhang2020-derived clathrin/non-clathrin sequence collections.

Individual claims
Advancing the accuracy of clathrin protein prediction through multi-source protein language models

Original source ↗

Methods: Dataset construction; Performance evaluation; Results: independent test datasets; cached text lines 9–10, 24–25, 54

Version: journal full text in PMC
Retrieved: 2026-09-16T10:38:57.558194+00:00

source checked

automated source review · 2026-09-16

Audit details

Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.

Field: attributes.profile.facts.0.value

Source artifact SHA-256: 2edc86b25707c1b737d26117093ce8d856e79cc5d0b335f27c1c341f887f1c7e

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Splits

Ten-fold cross-validation on training data plus named independent CLA test sets.

Individual claims
Advancing the accuracy of clathrin protein prediction through multi-source protein language models

Original source ↗

Methods: Dataset construction; Performance evaluation; Results: independent test datasets; cached text lines 9–10, 24–25, 54

Version: journal full text in PMC
Retrieved: 2026-09-16T10:38:57.558194+00:00

source checked

automated source review · 2026-09-16

Audit details

Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.

Field: attributes.profile.facts.1.value

Source artifact SHA-256: 2edc86b25707c1b737d26117093ce8d856e79cc5d0b335f27c1c341f887f1c7e

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Adaptation

Supervised classification with cross-validation and independent test collections.

Individual claims
Advancing the accuracy of clathrin protein prediction through multi-source protein language models

Original source ↗

Methods: Dataset construction; Performance evaluation; Results: independent test datasets; cached text lines 9–10, 24–25, 54

Version: journal full text in PMC
Retrieved: 2026-09-16T10:38:57.558194+00:00

source checked

automated source review · 2026-09-16

Audit details

Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.

Field: attributes.profile.facts.10.value

Source artifact SHA-256: 2edc86b25707c1b737d26117093ce8d856e79cc5d0b335f27c1c341f887f1c7e

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Metrics

Accuracy, AUC, Matthews correlation coefficient, F1, sensitivity and specificity are reported. The Performance evaluation subsection describes ten-fold evaluation and early stopping but does not specify a universal probability threshold for every comparator.

Individual claims
Advancing the accuracy of clathrin protein prediction through multi-source protein language models

Original source ↗

Materials and methods: Performance evaluation

Version: journal full text in PMC
Retrieved: 2026-09-16T10:38:57.558194+00:00

source checked

automated source review · 2026-09-16

Audit details

Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.

Field: attributes.profile.facts.2.value

Source artifact SHA-256: 2edc86b25707c1b737d26117093ce8d856e79cc5d0b335f27c1c341f887f1c7e

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Baselines

deep-clathrin includes literature-reported scores; DeepCLA is reimplemented for comparison.

Individual claims
Advancing the accuracy of clathrin protein prediction through multi-source protein language models

Original source ↗

Methods: Dataset construction; Performance evaluation; Results: independent test datasets; cached text lines 9–10, 24–25, 54

Version: journal full text in PMC
Retrieved: 2026-09-16T10:38:57.558194+00:00

source checked

automated source review · 2026-09-16

Audit details

Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.

Field: attributes.profile.facts.3.value

Source artifact SHA-256: 2edc86b25707c1b737d26117093ce8d856e79cc5d0b335f27c1c341f887f1c7e

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Leakage controls

The Zhang2020-derived collection uses BLAST redundancy filtering at 0.7, and the new dataset uses CD-HIT at 0.6. These dataset filters do not establish absence from the pretrained protein encoders’ corpora. Feature subsets are evaluated on both cross-validation and independent-test performance, which limits treating that test as untouched model selection.

Individual claims
Advancing the accuracy of clathrin protein prediction through multi-source protein language models

Original source ↗

Materials and methods: Dataset construction and Overall framework of PLM-CLA; Results: The effect of feature selection methods on the predictive performance

Version: journal full text in PMC
Retrieved: 2026-09-16T10:38:57.558194+00:00

source checked

automated source review · 2026-09-16

Audit details

Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.

Field: attributes.profile.facts.4.value

Source artifact SHA-256: 2edc86b25707c1b737d26117093ce8d856e79cc5d0b335f27c1c341f887f1c7e

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Uncertainty

The cited text-accessible evaluation sections give no confidence-interval, resampling or repeat-run error-bar specification. Image-only tables and uninspected supplements are outside this absence claim.

Individual claims
Advancing the accuracy of clathrin protein prediction through multi-source protein language models

Original source ↗

Methods: Dataset construction; Performance evaluation; Results: independent test datasets; cached text lines 9–10, 24–25, 54

Version: journal full text in PMC
Retrieved: 2026-09-16T10:38:57.558194+00:00

unreported

automated source review · 2026-09-16

Audit details

Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.

Field: attributes.profile.facts.5.value

Source artifact SHA-256: 2edc86b25707c1b737d26117093ce8d856e79cc5d0b335f27c1c341f887f1c7e

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Sources and history

Release 2026-09-17-d277315f7d76 · Record review: needs review

6 source records and release historyDownload this release
Technical metadata and extraction receipts

Stable ID: reported-task-786c09824e9bf5

areas
proteins-complexes
tasks
clathrin protein classification
entity level
task
version
Not reported
task
clathrin protein classification
scope note
Paper-specific evaluation task; protocol completeness requires further extraction.
comparison panels
id: part2-clathrin-plm-2025-Tab3-7ef65949fa; title: Clathrin independent test: selected-embedding classifiers · Table 3; protocol id: paper-protocol-bbffa94da73852557b; dataset id: paper-dataset-a1836de18515f3ab55; metric: ACC; unit: fraction; direction: higher; result ids: paper-result-0730029c96e20b98de; paper-result-a55f9924436b4f2946; paper-result-6d26a3281ddec02ba6; paper-result-199f6ba07b9c097f4f; paper-result-644ae2f042fb2f5b23; paper-result-36ae69af478636d799; paper-result-7ed01313e3f4c28c40; paper-result-7195f65c3684098379; paper-result-a969b52eed01ab25da; paper-result-4cf9ff8fa7d95ed4ce; paper-result-e677c183db2fc0a6fe; paper-result-f444ed165adbeb9c85; paper-result-48b062f1f39640cb79; source ids: part2-clathrin-plm-2025; source locator: Table 3: ACC, Clathrin independent test: selected-embedding classifiers; context: Conventional classifiers and PLM-CLA compared using paper-selected features. CLA-IND 0.6 independent test from Shoombuatong 2024 dataset; source Table 1 carries cohort counts.; caveats: Source prose has inconsistent dataset sizes; exact cohort counts left unresolved, not inferred.; Feature-selection and evaluation leakage cannot be excluded by this table alone.; No interval assigned unless printed in source cell.; review: method: automated_source_review; date: 2026-09-17; id: part2-clathrin-plm-2025-Tab3-9059dcb8eb; title: Clathrin independent test: selected-embedding classifiers · Table 3; protocol id: paper-protocol-bbffa94da73852557b; dataset id: paper-dataset-a1836de18515f3ab55; metric: SN; unit: fraction; direction: higher; result ids: paper-result-7e844e0a18dfd2a35b; paper-result-1d8775900e322544d5; paper-result-14af9ff7e7145d4d29; paper-result-2691b7b496a2b70805; paper-result-8a421ada6b63ee530d; paper-result-9824c5bbad9176ec8e; paper-result-5667de2e439b758d52; paper-result-c5b318cd60d0cdb4a7; paper-result-ac6a602d275b3fea3c; paper-result-5a7ce3c38ddae8961d; paper-result-a2a2ddd6bf8f13daef; paper-result-387a61f6d26366959d; paper-result-d62014e2f2c00f0b19; source ids: part2-clathrin-plm-2025; source locator: Table 3: SN, Clathrin independent test: selected-embedding classifiers; context: Conventional classifiers and PLM-CLA compared using paper-selected features. CLA-IND 0.6 independent test from Shoombuatong 2024 dataset; source Table 1 carries cohort counts.; caveats: Source prose has inconsistent dataset sizes; exact cohort counts left unresolved, not inferred.; Feature-selection and evaluation leakage cannot be excluded by this table alone.; No interval assigned unless printed in source cell.; review: method: automated_source_review; date: 2026-09-17; id: part2-clathrin-plm-2025-Tab3-1637766aec; title: Clathrin independent test: selected-embedding classifiers · Table 3; protocol id: paper-protocol-bbffa94da73852557b; dataset id: paper-dataset-a1836de18515f3ab55; metric: SP; unit: fraction; direction: higher; result ids: paper-result-21feace8cb2ce5e3eb; paper-result-e33d8392c76daadbc3; paper-result-97cc71a9c8b17a8bc1; paper-result-667c1dd0ce262e830c; paper-result-f1e415b2cfba651012; paper-result-0ce650036f050c2762; paper-result-21c44a4182b8225eb1; paper-result-a2d81dd3a62697406b; paper-result-7c339a10051c86a7c0; paper-result-ea67a0b7b24f3bb052; paper-result-9bdbc84d434b285e04; paper-result-87b08711535c28932a; paper-result-d04ecf04da1004cff8; source ids: part2-clathrin-plm-2025; source locator: Table 3: SP, Clathrin independent test: selected-embedding classifiers; context: Conventional classifiers and PLM-CLA compared using paper-selected features. CLA-IND 0.6 independent test from Shoombuatong 2024 dataset; source Table 1 carries cohort counts.; caveats: Source prose has inconsistent dataset sizes; exact cohort counts left unresolved, not inferred.; Feature-selection and evaluation leakage cannot be excluded by this table alone.; No interval assigned unless printed in source cell.; review: method: automated_source_review; date: 2026-09-17; id: part2-clathrin-plm-2025-Tab3-657b713876; title: Clathrin independent test: selected-embedding classifiers · Table 3; protocol id: paper-protocol-bbffa94da73852557b; dataset id: paper-dataset-a1836de18515f3ab55; metric: MCC; unit: dimensionless; direction: higher; result ids: paper-result-a4a56a5352e1f4b25f; paper-result-655c72140d737d23c3; paper-result-1482c92aa694d2f027; paper-result-250b4aaac43175427d; paper-result-ddc5d5b4d806f5d683; paper-result-92d64dfea940c8e704; paper-result-1671d2e2b268f272a5; paper-result-b30374085233a5ee59; paper-result-40e463b7ca62e3c241; paper-result-6644f2850c0d5eb339; paper-result-c9dfd2f1b60d02b3df; paper-result-930a9b4bcf61fbe236; paper-result-12c7518a23c15c58d6; source ids: part2-clathrin-plm-2025; source locator: Table 3: MCC, Clathrin independent test: selected-embedding classifiers; context: Conventional classifiers and PLM-CLA compared using paper-selected features. CLA-IND 0.6 independent test from Shoombuatong 2024 dataset; source Table 1 carries cohort counts.; caveats: Source prose has inconsistent dataset sizes; exact cohort counts left unresolved, not inferred.; Feature-selection and evaluation leakage cannot be excluded by this table alone.; No interval assigned unless printed in source cell.; review: method: automated_source_review; date: 2026-09-17; id: part2-clathrin-plm-2025-Tab3-a89d53b495; title: Clathrin independent test: selected-embedding classifiers · Table 3; protocol id: paper-protocol-bbffa94da73852557b; dataset id: paper-dataset-a1836de18515f3ab55; metric: F1; unit: fraction; direction: higher; result ids: paper-result-f92d473cd400e9e621; paper-result-96f846d80898fae95b; paper-result-9b9e600ad2446e697f; paper-result-af17b84609852bd79e; paper-result-5a6bc1834630e501e8; paper-result-2d7b6ae84bb35ea892; paper-result-a5bbea217db4b4a850; paper-result-d35402f96201b16bf8; paper-result-b30dbb80fd831aad99; paper-result-798a21dedf23053ac0; paper-result-3e7012cff605fa9f7f; paper-result-3431e8053e5b0bad25; paper-result-2b3a797967cdef7a15; source ids: part2-clathrin-plm-2025; source locator: Table 3: F1, Clathrin independent test: selected-embedding classifiers; context: Conventional classifiers and PLM-CLA compared using paper-selected features. CLA-IND 0.6 independent test from Shoombuatong 2024 dataset; source Table 1 carries cohort counts.; caveats: Source prose has inconsistent dataset sizes; exact cohort counts left unresolved, not inferred.; Feature-selection and evaluation leakage cannot be excluded by this table alone.; No interval assigned unless printed in source cell.; review: method: automated_source_review; date: 2026-09-17; id: part2-clathrin-plm-2025-Tab3-b6e9f189ae; title: Clathrin independent test: selected-embedding classifiers · Table 3; protocol id: paper-protocol-bbffa94da73852557b; dataset id: paper-dataset-a1836de18515f3ab55; metric: AUC; unit: fraction; direction: higher; result ids: paper-result-2eaf94777fc7f857fc; paper-result-47d350f89848982a43; paper-result-f1a86ea15cd7479d5b; paper-result-e274031e8e17dc0933; paper-result-ad1e11f4f3b25e62ed; paper-result-3741a0c6925f2947e2; paper-result-ca68d167787e314805; paper-result-7ef5e7bcd2b4a87ef9; paper-result-9565fcdb350e8db673; paper-result-5250d6a1217d2812cf; paper-result-e0a382c956e9add325; paper-result-9c2aeff72feeb82f45; paper-result-ab057b14f1c421f3ca; source ids: part2-clathrin-plm-2025; source locator: Table 3: AUC, Clathrin independent test: selected-embedding classifiers; context: Conventional classifiers and PLM-CLA compared using paper-selected features. CLA-IND 0.6 independent test from Shoombuatong 2024 dataset; source Table 1 carries cohort counts.; caveats: Source prose has inconsistent dataset sizes; exact cohort counts left unresolved, not inferred.; Feature-selection and evaluation leakage cannot be excluded by this table alone.; No interval assigned unless printed in source cell.; review: method: automated_source_review; date: 2026-09-17
benchmark research
review date: 2026-09-17; status: complete_comparison_tables_extracted_pending_publication_review; primary sources: part2-clathrin-plm-2025; inspected locators: Tables2–5; independent-test and dataset-construction sections; searched queries: Advancing the accuracy of clathrin protein prediction through multi-source protein language models primary paper benchmark results; gaps: Independent batch review before import; preserve existing observation identities.; claim scope: Primary-source discovery and table/protocol screening; source checked is not independently reproduced. Raw acquisitions not automatically numerical publication approval.
historical missing metadata
protocol version: not_reported_in_legacy_extract
metadata review scope
historical_missing_metadata preserves the original discovery state. Current descriptive evidence and missingness are recorded in profile.facts; numerical-result review is separate.
legacy kinds
benchmark
entity classification
review date: 2026-09-17; rationale: This source-scoped record identifies the biological prediction task and holds its paper context. Preserve the existing task identity; exact split, model adaptation and scoring remain in linked evaluations or separate protocol records.; source ids: clathrin-plm-2025; source locator: Methods: Dataset construction; Performance evaluation; Results: independent test datasets; cached text lines 9–10, 24–25, 54; ambiguities: A paper- or suite-specific task may constrain some inputs or metrics; that alone does not make it interchangeable with a complete versioned protocol. No protocol equivalence is inferred.; Some legacy profile Entity type facts use the generic phrase computational evaluation protocol. That boilerplate is not sufficient to establish a single fixed protocol identity or to merge this task with another protocol record.
Related records

Suggest a correction