rewire.it
Task

Enzyme functional identity prediction

Enzyme functional-identity classification predicts whether a protein pair shares its annotated reaction function.

SourcesEnhanced prediction of protein functional identity through the integration of sequence and structural features · Methods: Dataset construction; classification and metrics; Results: model comparison; cached text lines 10–11, 25, 28, 52, 55; uncertainty/repeat-run/statistical-comparison passages

5 evaluations · 40 metric rows

At a glance

Inputs, training, access and other details

Explanatory profile: limited source coverage · Automated source review, 2026-09-16. Review applies to the cited claims; unresolved fields are listed below. Numerical results retain their own review status.

Data, procedure and scoring
PropertyDescription and evidence
DatasetsSwiss-Prot reaction annotations and AlphaFold Database structures; balanced same-function and different-function pairs.
SourcesEnhanced prediction of protein functional identity through the integration of sequence and structural features · Methods: Dataset construction; classification and metrics; Results: model comparison; cached text lines 10–11, 25, 28, 52, 55; uncertainty/repeat-run/statistical-comparison passages
SplitsRandom pair partitions provide training, validation and test subsets; model hyperparameters are tuned using cross-validation.
SourcesEnhanced prediction of protein functional identity through the integration of sequence and structural features · Methods: Dataset construction; classification and metrics; Results: model comparison; cached text lines 10–11, 25, 28, 52, 55; uncertainty/repeat-run/statistical-comparison passages
MetricsAccuracy, false-positive rate, MCC, precision, recall, F1, AUROC and AUPR.
SourcesEnhanced prediction of protein functional identity through the integration of sequence and structural features · Methods: Dataset construction; classification and metrics; Results: model comparison; cached text lines 10–11, 25, 28, 52, 55; uncertainty/repeat-run/statistical-comparison passages
BaselinesLightGBM and multiple conventional classifiers.
SourcesEnhanced prediction of protein functional identity through the integration of sequence and structural features · Methods: Dataset construction; classification and metrics; Results: model comparison; cached text lines 10–11, 25, 28, 52, 55; uncertainty/repeat-run/statistical-comparison passages
Leakage controlsThe low-sequence-similarity evaluation excludes pairs from the original dataset; random splitting of the main pair dataset does not establish protein-identity holdout.
SourcesEnhanced prediction of protein functional identity through the integration of sequence and structural features · Methods: Dataset construction; classification and metrics; Results: model comparison; cached text lines 10–11, 25, 28, 52, 55; uncertainty/repeat-run/statistical-comparison passages
UncertaintyThe paper assesses stability using repeated bootstrap iterations; its sampling unit must remain attached to the reported interval.
SourcesEnhanced prediction of protein functional identity through the integration of sequence and structural features · Methods: Dataset construction; classification and metrics; Results: model comparison; cached text lines 10–11, 25, 28, 52, 55; uncertainty/repeat-run/statistical-comparison passages
Entity typePaper-specific computational evaluation protocol.
SourcesEnhanced prediction of protein functional identity through the integration of sequence and structural features · Methods: Dataset construction; classification and metrics; Results: model comparison; cached text lines 10–11, 25, 28, 52, 55; uncertainty/repeat-run/statistical-comparison passages
OrganismsProteins were selected through Swiss-Prot release 2022_04 entries with Rhea reaction annotations and AlphaFold DB v4 structures. Dataset construction does not enumerate organism frequencies for the sampled 100,000 protein pairs, so a species-restricted population cannot be assigned. · Not reported in inspected sources
SourcesEnhanced prediction of protein functional identity through the integration of sequence and structural features · Materials and methods: Dataset construction
AssaysSwiss-Prot reaction/function annotations with AlphaFold structures.
SourcesEnhanced prediction of protein functional identity through the integration of sequence and structural features · Methods: Dataset construction; classification and metrics; Results: model comparison; cached text lines 10–11, 25, 28, 52, 55; uncertainty/repeat-run/statistical-comparison passages
Allowed inputsPairs of enzyme representations.
SourcesEnhanced prediction of protein functional identity through the integration of sequence and structural features · Methods: Dataset construction; classification and metrics; Results: model comparison; cached text lines 10–11, 25, 28, 52, 55; uncertainty/repeat-run/statistical-comparison passages
AdaptationSupervised same-function classification; hyperparameters use cross-validation, with algorithm choice additionally compared on test data.
SourcesEnhanced prediction of protein functional identity through the integration of sequence and structural features · Methods: Dataset construction; classification and metrics; Results: model comparison; cached text lines 10–11, 25, 28, 52, 55; uncertainty/repeat-run/statistical-comparison passages

How it works

How it worksComputational evaluation flow
Computational evaluation flow1. Input: Pairs of enzyme representations.. Then: 2. Evaluation: Random pair partitions provide training, validation and test subsets; model hyperparameters are tuned using cross-validation.. Then: 3. Readout: Accuracy, false-positive rate, MCC, precision, recall, F1, AUROC and AUPR.Computational evaluation flow1. Input: Pairs of enzyme representations.. Then: 2. Evaluation: Random pair partitions provide training, validation and test subsets; model hyperparameters are tuned using cross-validation.. Then: 3. Readout: Accuracy, false-positive rate, MCC, precision, recall, F1, AUROC and AUPR.Computational evaluation flow1. Input: Pairs of enzyme representations.. Then: 2. Evaluation: Random pair partitions provide training, validation and test subsets; model hyperparameters are tuned using cross-validation.. Then: 3. Readout: Accuracy, false-positive rate, MCC, precision, recall, F1, AUROC and AUPR.

Conceptual summary of the cited evaluation; exact task configuration and source version remain part of the protocol.

SourcesEnhanced prediction of protein functional identity through the integration of sequence and structural features · Methods: Dataset construction; classification and metrics; Results: model comparison; cached text lines 10–11, 25, 28, 52, 55; uncertainty/repeat-run/statistical-comparison passages
Evaluation methodology

Swiss-Prot reaction annotations and AlphaFold Database structures; balanced same-function and different-function pairs. Random pair partitions provide training, validation and test subsets; model hyperparameters are tuned using cross-validation. Accuracy, false-positive rate, MCC, precision, recall, F1, AUROC and AUPR. LightGBM and multiple conventional classifiers. The low-sequence-similarity evaluation excludes pairs from the original dataset; random splitting of the main pair dataset does not establish protein-identity holdout.

SourcesEnhanced prediction of protein functional identity through the integration of sequence and structural features · Methods: Dataset construction; classification and metrics; Results: model comparison; cached text lines 10–11, 25, 28, 52, 55; uncertainty/repeat-run/statistical-comparison passages

Evaluation design

Benchmarks bring together tasks and protocols. A task describes the biological question; a protocol defines a particular test.

These source-backed links do not make different protocols or scores interchangeable.

Published comparisons

Explore the results reported under one evaluation protocol. Each figure keeps its source, dataset and metric together; it is not a ranking across studies.

Enzyme-pair functional identity: original held-out test · Table 1

ACC (%) (percent) · Higher values are better for this metric.

LightGBM FUJISAN and comparison methods; predictions thresholded to maximizeF 1. Protein-pair splitting is not evidence of disjoint proteins or families. 41,600 protein pairs; balanced functional-identity classes; random 56.25/18.75/25% training/validation/test split.

Evaluation protocol · FUJISAN test sub-dataset

  1. FUJISAN · Configuration · Author-reported evaluation87.05
  2. E-value · Configuration · Author-reported evaluation81.91
  3. DeepFRI · Configuration · Author-reported evaluation80.88
  4. ESM2 · Configuration · Independent external evaluation71.33
  5. Pfam · Configuration · Author-reported evaluation61.99

Source order is preserved. Plotted marks show point estimates; uncertainty, where reported, is retained in the printed values and table. Differences do not establish statistical significance.

Enhanced prediction of protein functional identity through the integration of sequence and structural features · Table 1: ACC (%), Enzyme-pair functional identity: original held-out test
Values, uncertainty and evidence
ACC (%): original source values
Tested entityPrinted valueUncertaintyEvidence
FUJISAN · Configuration87.05 percentNot reportedAuthor-reported evaluation · source checkedEnhanced prediction of protein functional identity through the integration of sequence and structural features · Table 1, row FUJISAN, column ACC (%); XML row2 column2
E-value · Configuration81.91 percentNot reportedAuthor-reported evaluation · source checkedEnhanced prediction of protein functional identity through the integration of sequence and structural features · Table 1, row E-value, column ACC (%); XML row3 column2
DeepFRI · Configuration80.88 percentNot reportedAuthor-reported evaluation · source checkedEnhanced prediction of protein functional identity through the integration of sequence and structural features · Table 1, row DeepFRI, column ACC (%); XML row4 column2
ESM2 · Configuration71.33 percentNot reportedIndependent external evaluation · source checkedEnhanced prediction of protein functional identity through the integration of sequence and structural features · Table 1, row ESM2, column ACC (%); XML row5 column2
Pfam · Configuration61.99 percentNot reportedAuthor-reported evaluation · source checkedEnhanced prediction of protein functional identity through the integration of sequence and structural features · Table 1, row Pfam, column ACC (%); XML row6 column2
Scope and limitations
  • Input information differs: sequence similarity, protein embeddings, structure-aware descriptors and domain annotations.
  • Table1 reports point values;50bootstrap iterations described elsewhere do not establish a Table1 interval.
  • No interval assigned unless printed in source cell.

Source transcription and grouping reviewed by automated source review on 2026-09-17. These experiments were not independently reproduced by rewire.

Tested entities and results

Release 2026-09-17-d277315f7d76 · 5 evaluations · 40 metric rows. Different protocols are not a single leaderboard.

Results grouped by the exact reported evaluation
Metric and findingCoverage and uncertaintyEvidence
FUJISAN: Enzyme functional identity prediction

LightGBM FUJISAN and comparison methods; predictions thresholded to maximizeF 1. Protein-pair splitting is not evidence of disjoint proteins or families. 41,600 protein pairs; balanced functional-identity classes; random 56.25/18.75/25% training/validation/test split.

Author-reported evaluation · Evaluation metadata: needs review

0.9427 AUROC

Unit: unitless · Direction: higher

Uncertainty: not reported in legacy extract

Scored: Not reported · Eligible: Not reported

source checkedEnhanced prediction of protein functional identity through the integration of sequence and structural features; Enhanced prediction of protein functional identity through the integration of sequence and structural features · Table 1, FUJISAN row, AUROC column

Source checking is not independent reproduction.

87.05% REC (%)

Unit: percent · Direction: higher

Uncertainty: unreported

Scored: Not reported · Eligible: Not reported

source checkedEnhanced prediction of protein functional identity through the integration of sequence and structural features · Table 1, row FUJISAN, column REC (%); XML row2 column4

Source checking is not independent reproduction.

0.7421 MCC

Unit: dimensionless · Direction: higher

Uncertainty: unreported

Scored: Not reported · Eligible: Not reported

source checkedEnhanced prediction of protein functional identity through the integration of sequence and structural features · Table 1, row FUJISAN, column MCC; XML row2 column7

Source checking is not independent reproduction.

87.24% PRE (%)

Unit: percent · Direction: higher

Uncertainty: unreported

Scored: Not reported · Eligible: Not reported

source checkedEnhanced prediction of protein functional identity through the integration of sequence and structural features · Table 1, row FUJISAN, column PRE (%); XML row2 column3

Source checking is not independent reproduction.

ESM2: Enzyme functional identity prediction

LightGBM FUJISAN and comparison methods; predictions thresholded to maximizeF 1. Protein-pair splitting is not evidence of disjoint proteins or families. 41,600 protein pairs; balanced functional-identity classes; random 56.25/18.75/25% training/validation/test split.

Independent external evaluation · Evaluation metadata: needs review

0.7991 AUROC

Unit: unitless · Direction: higher

Uncertainty: not reported in legacy extract

Scored: Not reported · Eligible: Not reported

source checkedEnhanced prediction of protein functional identity through the integration of sequence and structural features; Enhanced prediction of protein functional identity through the integration of sequence and structural features · Table 1, ESM2 row, AUROC column

Source checking is not independent reproduction.

79.33% REC (%)

Unit: percent · Direction: higher

Uncertainty: unreported

Scored: Not reported · Eligible: Not reported

source checkedEnhanced prediction of protein functional identity through the integration of sequence and structural features · Table 1, row ESM2, column REC (%); XML row5 column4

Source checking is not independent reproduction.

0.8147 AUPR

Unit: dimensionless · Direction: higher

Uncertainty: unreported

Scored: Not reported · Eligible: Not reported

source checkedEnhanced prediction of protein functional identity through the integration of sequence and structural features · Table 1, row ESM2, column AUPR; XML row5 column9

Source checking is not independent reproduction.

0.4321 MCC

Unit: dimensionless · Direction: higher

Uncertainty: unreported

Scored: Not reported · Eligible: Not reported

source checkedEnhanced prediction of protein functional identity through the integration of sequence and structural features · Table 1, row ESM2, column MCC; XML row5 column7

Source checking is not independent reproduction.

68.39% PRE (%)

Unit: percent · Direction: higher

Uncertainty: unreported

Scored: Not reported · Eligible: Not reported

source checkedEnhanced prediction of protein functional identity through the integration of sequence and structural features · Table 1, row ESM2, column PRE (%); XML row5 column3

Source checking is not independent reproduction.

E-value: Enzyme-pair functional identity: original held-out test

LightGBM FUJISAN and comparison methods; predictions thresholded to maximizeF 1. Protein-pair splitting is not evidence of disjoint proteins or families. 41,600 protein pairs; balanced functional-identity classes; random 56.25/18.75/25% training/validation/test split.

Author-reported evaluation · Evaluation metadata: needs review

78.58% PRE (%)

Unit: percent · Direction: higher

Uncertainty: unreported

Scored: Not reported · Eligible: Not reported

source checkedEnhanced prediction of protein functional identity through the integration of sequence and structural features · Table 1, row E-value, column PRE (%); XML row3 column3

Source checking is not independent reproduction.

0.8291 F1

Unit: dimensionless · Direction: higher

Uncertainty: unreported

Scored: Not reported · Eligible: Not reported

source checkedEnhanced prediction of protein functional identity through the integration of sequence and structural features · Table 1, row E-value, column F1; XML row3 column6

Source checking is not independent reproduction.

0.6427 MCC

Unit: dimensionless · Direction: higher

Uncertainty: unreported

Scored: Not reported · Eligible: Not reported

source checkedEnhanced prediction of protein functional identity through the integration of sequence and structural features · Table 1, row E-value, column MCC; XML row3 column7

Source checking is not independent reproduction.

23.92% FPR (%)

Unit: percent · Direction: lower

Uncertainty: unreported

Scored: Not reported · Eligible: Not reported

source checkedEnhanced prediction of protein functional identity through the integration of sequence and structural features · Table 1, row E-value, column FPR (%); XML row3 column5

Source checking is not independent reproduction.

87.75% REC (%)

Unit: percent · Direction: higher

Uncertainty: unreported

Scored: Not reported · Eligible: Not reported

source checkedEnhanced prediction of protein functional identity through the integration of sequence and structural features · Table 1, row E-value, column REC (%); XML row3 column4

Source checking is not independent reproduction.

81.91% ACC (%)

Unit: percent · Direction: higher

Uncertainty: unreported

Scored: Not reported · Eligible: Not reported

source checkedEnhanced prediction of protein functional identity through the integration of sequence and structural features · Table 1, row E-value, column ACC (%); XML row3 column2

Source checking is not independent reproduction.

DeepFRI: Enzyme-pair functional identity: original held-out test

LightGBM FUJISAN and comparison methods; predictions thresholded to maximizeF 1. Protein-pair splitting is not evidence of disjoint proteins or families. 41,600 protein pairs; balanced functional-identity classes; random 56.25/18.75/25% training/validation/test split.

Author-reported evaluation · Evaluation metadata: needs review

83.16% REC (%)

Unit: percent · Direction: higher

Uncertainty: unreported

Scored: Not reported · Eligible: Not reported

source checkedEnhanced prediction of protein functional identity through the integration of sequence and structural features · Table 1, row DeepFRI, column REC (%); XML row4 column4

Source checking is not independent reproduction.

80.88% ACC (%)

Unit: percent · Direction: higher

Uncertainty: unreported

Scored: Not reported · Eligible: Not reported

source checkedEnhanced prediction of protein functional identity through the integration of sequence and structural features · Table 1, row DeepFRI, column ACC (%); XML row4 column2

Source checking is not independent reproduction.

0.6790 MCC

Unit: dimensionless · Direction: higher

Uncertainty: unreported

Scored: Not reported · Eligible: Not reported

source checkedEnhanced prediction of protein functional identity through the integration of sequence and structural features · Table 1, row DeepFRI, column MCC; XML row4 column7

Source checking is not independent reproduction.

0.8143 F1

Unit: dimensionless · Direction: higher

Uncertainty: unreported

Scored: Not reported · Eligible: Not reported

source checkedEnhanced prediction of protein functional identity through the integration of sequence and structural features · Table 1, row DeepFRI, column F1; XML row4 column6

Source checking is not independent reproduction.

Pfam: Enzyme-pair functional identity: original held-out test

LightGBM FUJISAN and comparison methods; predictions thresholded to maximizeF 1. Protein-pair splitting is not evidence of disjoint proteins or families. 41,600 protein pairs; balanced functional-identity classes; random 56.25/18.75/25% training/validation/test split.

Author-reported evaluation · Evaluation metadata: needs review

61.99% ACC (%)

Unit: percent · Direction: higher

Uncertainty: unreported

Scored: Not reported · Eligible: Not reported

source checkedEnhanced prediction of protein functional identity through the integration of sequence and structural features · Table 1, row Pfam, column ACC (%); XML row6 column2

Source checking is not independent reproduction.

0.3520 MCC

Unit: dimensionless · Direction: higher

Uncertainty: unreported

Scored: Not reported · Eligible: Not reported

source checkedEnhanced prediction of protein functional identity through the integration of sequence and structural features · Table 1, row Pfam, column MCC; XML row6 column7

Source checking is not independent reproduction.

56.92% PRE (%)

Unit: percent · Direction: higher

Uncertainty: unreported

Scored: Not reported · Eligible: Not reported

source checkedEnhanced prediction of protein functional identity through the integration of sequence and structural features · Table 1, row Pfam, column PRE (%); XML row6 column3

Source checking is not independent reproduction.

0.7217 F1

Unit: dimensionless · Direction: higher

Uncertainty: unreported

Scored: Not reported · Eligible: Not reported

source checkedEnhanced prediction of protein functional identity through the integration of sequence and structural features · Table 1, row Pfam, column F1; XML row6 column6

Source checking is not independent reproduction.

74.62% FPR (%)

Unit: percent · Direction: lower

Uncertainty: unreported

Scored: Not reported · Eligible: Not reported

source checkedEnhanced prediction of protein functional identity through the integration of sequence and structural features · Table 1, row Pfam, column FPR (%); XML row6 column5

Source checking is not independent reproduction.

98.60% REC (%)

Unit: percent · Direction: higher

Uncertainty: unreported

Scored: Not reported · Eligible: Not reported

source checkedEnhanced prediction of protein functional identity through the integration of sequence and structural features · Table 1, row Pfam, column REC (%); XML row6 column4

Source checking is not independent reproduction.

Papers and result coverage

Last literature check: 2026-09-17. Primary-source discovery and table/protocol screening; source checked is not independently reproduced. Raw acquisitions not automatically numerical publication approval.

Paper or primary resourceVersionReference
Enhanced prediction of protein functional identity through the integration of sequence and structural featuresPMC11609699.1Read source
DOI: 10.1016/j.csbj.2024.11.028

What is still missing

  • Independent batch review before import; preserve existing observation identities.
Search and extraction details

complete comparison tables extracted pending publication review

Searches

  • Enhanced prediction of protein functional identity through the integration of sequence and structural features primary paper benchmark results

Evidence locations

  • Table1; Model training and hyperparameter optimization; Performance assessment

Strengths and limitations

Strengths and considerations

No source-reviewed explanatory claims are recorded here yet.

Limitations and conditions

Profile review details

Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.

Stable record: reported-task-1ebf9b408517f9

Evidence table

Inspect claims, sources and review details

Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.

One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.

17 evidence rows matching the loaded filters

Claims, original sources and review scope · Release 2026-09-17-d277315f7d76
Property and statementOriginal source and locationReview and provenance
Diagram caption

Conceptual summary of the cited evaluation; exact task configuration and source version remain part of the protocol.

Individual claims
Enhanced prediction of protein functional identity through the integration of sequence and structural features

Original source ↗

Methods: Dataset construction; classification and metrics; Results: model comparison; cached text lines 10–11, 25, 28, 52, 55; uncertainty/repeat-run/statistical-comparison passages

Version: PMC11609699.1
Retrieved: 2026-09-16T10:33:35.728Z

source checked

automated source review · 2026-09-16

Audit details

Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.

Field: attributes.profile.diagram.caption

Source artifact SHA-256: db33e0542005ffae00cd644dfe697185b94c8823d5aee2768620a6db0c48e56f

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Diagram steps

["Input: Pairs of enzyme representations.","Evaluation: Random pair partitions provide training, validation and test subsets; model hyperparameters are tuned using cross-validation.","Readout: Accuracy, false-positive rate, MCC, precision, recall, F1, AUROC and AUPR."]

Individual claims
Enhanced prediction of protein functional identity through the integration of sequence and structural features

Original source ↗

Methods: Dataset construction; classification and metrics; Results: model comparison; cached text lines 10–11, 25, 28, 52, 55; uncertainty/repeat-run/statistical-comparison passages

Version: PMC11609699.1
Retrieved: 2026-09-16T10:33:35.728Z

source checked

automated source review · 2026-09-16

Audit details

Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.

Field: attributes.profile.diagram.steps

Source artifact SHA-256: db33e0542005ffae00cd644dfe697185b94c8823d5aee2768620a6db0c48e56f

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Diagram title

Computational evaluation flow

Individual claims
Enhanced prediction of protein functional identity through the integration of sequence and structural features

Original source ↗

Methods: Dataset construction; classification and metrics; Results: model comparison; cached text lines 10–11, 25, 28, 52, 55; uncertainty/repeat-run/statistical-comparison passages

Version: PMC11609699.1
Retrieved: 2026-09-16T10:33:35.728Z

source checked

automated source review · 2026-09-16

Audit details

Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.

Field: attributes.profile.diagram.title

Source artifact SHA-256: db33e0542005ffae00cd644dfe697185b94c8823d5aee2768620a6db0c48e56f

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Datasets

Swiss-Prot reaction annotations and AlphaFold Database structures; balanced same-function and different-function pairs.

Individual claims
Enhanced prediction of protein functional identity through the integration of sequence and structural features

Original source ↗

Methods: Dataset construction; classification and metrics; Results: model comparison; cached text lines 10–11, 25, 28, 52, 55; uncertainty/repeat-run/statistical-comparison passages

Version: PMC11609699.1
Retrieved: 2026-09-16T10:33:35.728Z

source checked

automated source review · 2026-09-16

Audit details

Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.

Field: attributes.profile.facts.0.value

Source artifact SHA-256: db33e0542005ffae00cd644dfe697185b94c8823d5aee2768620a6db0c48e56f

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Splits

Random pair partitions provide training, validation and test subsets; model hyperparameters are tuned using cross-validation.

Individual claims
Enhanced prediction of protein functional identity through the integration of sequence and structural features

Original source ↗

Methods: Dataset construction; classification and metrics; Results: model comparison; cached text lines 10–11, 25, 28, 52, 55; uncertainty/repeat-run/statistical-comparison passages

Version: PMC11609699.1
Retrieved: 2026-09-16T10:33:35.728Z

source checked

automated source review · 2026-09-16

Audit details

Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.

Field: attributes.profile.facts.1.value

Source artifact SHA-256: db33e0542005ffae00cd644dfe697185b94c8823d5aee2768620a6db0c48e56f

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Adaptation

Supervised same-function classification; hyperparameters use cross-validation, with algorithm choice additionally compared on test data.

Individual claims
Enhanced prediction of protein functional identity through the integration of sequence and structural features

Original source ↗

Methods: Dataset construction; classification and metrics; Results: model comparison; cached text lines 10–11, 25, 28, 52, 55; uncertainty/repeat-run/statistical-comparison passages

Version: PMC11609699.1
Retrieved: 2026-09-16T10:33:35.728Z

source checked

automated source review · 2026-09-16

Audit details

Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.

Field: attributes.profile.facts.10.value

Source artifact SHA-256: db33e0542005ffae00cd644dfe697185b94c8823d5aee2768620a6db0c48e56f

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Metrics

Accuracy, false-positive rate, MCC, precision, recall, F1, AUROC and AUPR.

Individual claims
Enhanced prediction of protein functional identity through the integration of sequence and structural features

Original source ↗

Methods: Dataset construction; classification and metrics; Results: model comparison; cached text lines 10–11, 25, 28, 52, 55; uncertainty/repeat-run/statistical-comparison passages

Version: PMC11609699.1
Retrieved: 2026-09-16T10:33:35.728Z

source checked

automated source review · 2026-09-16

Audit details

Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.

Field: attributes.profile.facts.2.value

Source artifact SHA-256: db33e0542005ffae00cd644dfe697185b94c8823d5aee2768620a6db0c48e56f

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Baselines

LightGBM and multiple conventional classifiers.

Individual claims
Enhanced prediction of protein functional identity through the integration of sequence and structural features

Original source ↗

Methods: Dataset construction; classification and metrics; Results: model comparison; cached text lines 10–11, 25, 28, 52, 55; uncertainty/repeat-run/statistical-comparison passages

Version: PMC11609699.1
Retrieved: 2026-09-16T10:33:35.728Z

source checked

automated source review · 2026-09-16

Audit details

Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.

Field: attributes.profile.facts.3.value

Source artifact SHA-256: db33e0542005ffae00cd644dfe697185b94c8823d5aee2768620a6db0c48e56f

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Leakage controls

The low-sequence-similarity evaluation excludes pairs from the original dataset; random splitting of the main pair dataset does not establish protein-identity holdout.

Individual claims
Enhanced prediction of protein functional identity through the integration of sequence and structural features

Original source ↗

Methods: Dataset construction; classification and metrics; Results: model comparison; cached text lines 10–11, 25, 28, 52, 55; uncertainty/repeat-run/statistical-comparison passages

Version: PMC11609699.1
Retrieved: 2026-09-16T10:33:35.728Z

source checked

automated source review · 2026-09-16

Audit details

Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.

Field: attributes.profile.facts.4.value

Source artifact SHA-256: db33e0542005ffae00cd644dfe697185b94c8823d5aee2768620a6db0c48e56f

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Uncertainty

The paper assesses stability using repeated bootstrap iterations; its sampling unit must remain attached to the reported interval.

Individual claims
Enhanced prediction of protein functional identity through the integration of sequence and structural features

Original source ↗

Methods: Dataset construction; classification and metrics; Results: model comparison; cached text lines 10–11, 25, 28, 52, 55; uncertainty/repeat-run/statistical-comparison passages

Version: PMC11609699.1
Retrieved: 2026-09-16T10:33:35.728Z

source checked

automated source review · 2026-09-16

Audit details

Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.

Field: attributes.profile.facts.5.value

Source artifact SHA-256: db33e0542005ffae00cd644dfe697185b94c8823d5aee2768620a6db0c48e56f

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Sources and history

Release 2026-09-17-d277315f7d76 · Record review: needs review

2 source records and release historyDownload this release
Technical metadata and extraction receipts

Stable ID: reported-task-1ebf9b408517f9

areas
proteins-complexes
tasks
Enzyme functional identity prediction
entity level
task
version
Not reported
task
Enzyme functional identity prediction
scope note
Paper-specific evaluation task; protocol completeness requires further extraction.
comparison panels
id: part2-fujisan-2024-tbl0005-bac8a6d22a; title: Enzyme-pair functional identity: original held-out test · Table 1; protocol id: paper-protocol-7088bcc2f0033b8a29; dataset id: reported-dataset-5197cca532f89d; metric: ACC (%); unit: percent; direction: higher; result ids: paper-result-eaf1d6c59a5f36f4ce; paper-result-7a33f77a25091433d2; paper-result-357d8f3b2935e06e3d; paper-result-b621ec294f38d3f19c; paper-result-23786f6a1a92179f21; source ids: part2-fujisan-2024; source locator: Table 1: ACC (%), Enzyme-pair functional identity: original held-out test; context: LightGBM FUJISAN and comparison methods; predictions thresholded to maximizeF 1. Protein-pair splitting is not evidence of disjoint proteins or families. 41,600 protein pairs; balanced functional-identity classes; random 56.25/18.75/25% training/validation/test split.; caveats: Input information differs: sequence similarity, protein embeddings, structure-aware descriptors and domain annotations.; Table1 reports point values;50bootstrap iterations described elsewhere do not establish a Table1 interval.; No interval assigned unless printed in source cell.; review: method: automated_source_review; date: 2026-09-17; id: part2-fujisan-2024-tbl0005-8cc77c9aa4; title: Enzyme-pair functional identity: original held-out test · Table 1; protocol id: paper-protocol-7088bcc2f0033b8a29; dataset id: reported-dataset-5197cca532f89d; metric: PRE (%); unit: percent; direction: higher; result ids: paper-result-989a4675188b8141c2; paper-result-003d7035891e68f095; paper-result-b6bde99a5fd8fcf312; paper-result-a18d074f2077ce82f8; paper-result-5c36a73bb7912f46bb; source ids: part2-fujisan-2024; source locator: Table 1: PRE (%), Enzyme-pair functional identity: original held-out test; context: LightGBM FUJISAN and comparison methods; predictions thresholded to maximizeF 1. Protein-pair splitting is not evidence of disjoint proteins or families. 41,600 protein pairs; balanced functional-identity classes; random 56.25/18.75/25% training/validation/test split.; caveats: Input information differs: sequence similarity, protein embeddings, structure-aware descriptors and domain annotations.; Table1 reports point values;50bootstrap iterations described elsewhere do not establish a Table1 interval.; No interval assigned unless printed in source cell.; review: method: automated_source_review; date: 2026-09-17; id: part2-fujisan-2024-tbl0005-3714094960; title: Enzyme-pair functional identity: original held-out test · Table 1; protocol id: paper-protocol-7088bcc2f0033b8a29; dataset id: reported-dataset-5197cca532f89d; metric: REC (%); unit: percent; direction: higher; result ids: paper-result-0081b12c9bb122c13a; paper-result-71abecaef7ab30d471; paper-result-21390f8f107e3d5169; paper-result-0d3498cab7dbd41b99; paper-result-9907ce062be9dbd488; source ids: part2-fujisan-2024; source locator: Table 1: REC (%), Enzyme-pair functional identity: original held-out test; context: LightGBM FUJISAN and comparison methods; predictions thresholded to maximizeF 1. Protein-pair splitting is not evidence of disjoint proteins or families. 41,600 protein pairs; balanced functional-identity classes; random 56.25/18.75/25% training/validation/test split.; caveats: Input information differs: sequence similarity, protein embeddings, structure-aware descriptors and domain annotations.; Table1 reports point values;50bootstrap iterations described elsewhere do not establish a Table1 interval.; No interval assigned unless printed in source cell.; review: method: automated_source_review; date: 2026-09-17; id: part2-fujisan-2024-tbl0005-a129b9c665; title: Enzyme-pair functional identity: original held-out test · Table 1; protocol id: paper-protocol-7088bcc2f0033b8a29; dataset id: reported-dataset-5197cca532f89d; metric: FPR (%); unit: percent; direction: lower; result ids: paper-result-ff3b35eac39b7fa2ed; paper-result-6e46bf4f05ccae3dff; paper-result-a36c2ff5211b482ac8; paper-result-c478bf21db79f9c0fa; paper-result-95850ea0b3b65d0a9f; source ids: part2-fujisan-2024; source locator: Table 1: FPR (%), Enzyme-pair functional identity: original held-out test; context: LightGBM FUJISAN and comparison methods; predictions thresholded to maximizeF 1. Protein-pair splitting is not evidence of disjoint proteins or families. 41,600 protein pairs; balanced functional-identity classes; random 56.25/18.75/25% training/validation/test split.; caveats: Input information differs: sequence similarity, protein embeddings, structure-aware descriptors and domain annotations.; Table1 reports point values;50bootstrap iterations described elsewhere do not establish a Table1 interval.; No interval assigned unless printed in source cell.; review: method: automated_source_review; date: 2026-09-17; id: part2-fujisan-2024-tbl0005-0b4696dd22; title: Enzyme-pair functional identity: original held-out test · Table 1; protocol id: paper-protocol-7088bcc2f0033b8a29; dataset id: reported-dataset-5197cca532f89d; metric: F1; unit: dimensionless; direction: higher; result ids: paper-result-ef58d0787535e7e9a5; paper-result-113e4214c6249dc3c8; paper-result-895a3e38bb00f10f4a; paper-result-b5897e7fb2d9f17e8c; paper-result-73d717fd51d9b2682a; source ids: part2-fujisan-2024; source locator: Table 1: F1, Enzyme-pair functional identity: original held-out test; context: LightGBM FUJISAN and comparison methods; predictions thresholded to maximizeF 1. Protein-pair splitting is not evidence of disjoint proteins or families. 41,600 protein pairs; balanced functional-identity classes; random 56.25/18.75/25% training/validation/test split.; caveats: Input information differs: sequence similarity, protein embeddings, structure-aware descriptors and domain annotations.; Table1 reports point values;50bootstrap iterations described elsewhere do not establish a Table1 interval.; No interval assigned unless printed in source cell.; review: method: automated_source_review; date: 2026-09-17; id: part2-fujisan-2024-tbl0005-d782fec4a4; title: Enzyme-pair functional identity: original held-out test · Table 1; protocol id: paper-protocol-7088bcc2f0033b8a29; dataset id: reported-dataset-5197cca532f89d; metric: MCC; unit: dimensionless; direction: higher; result ids: paper-result-1ebedcb0adb7517e54; paper-result-3559edcbf93633cc0d; paper-result-58038afdc07c1aada5; paper-result-46452dcfe78a4f3d04; paper-result-5ac77bcc4aed51f516; source ids: part2-fujisan-2024; source locator: Table 1: MCC, Enzyme-pair functional identity: original held-out test; context: LightGBM FUJISAN and comparison methods; predictions thresholded to maximizeF 1. Protein-pair splitting is not evidence of disjoint proteins or families. 41,600 protein pairs; balanced functional-identity classes; random 56.25/18.75/25% training/validation/test split.; caveats: Input information differs: sequence similarity, protein embeddings, structure-aware descriptors and domain annotations.; Table1 reports point values;50bootstrap iterations described elsewhere do not establish a Table1 interval.; No interval assigned unless printed in source cell.; review: method: automated_source_review; date: 2026-09-17; id: part2-fujisan-2024-tbl0005-d1fced1a9c; title: Enzyme-pair functional identity: original held-out test · Table 1; protocol id: paper-protocol-7088bcc2f0033b8a29; dataset id: reported-dataset-5197cca532f89d; metric: AUROC; unit: unitless; direction: higher; result ids: lit-019; paper-result-d2f325ea0c890b170d; paper-result-b0dd968f641a7e8f16; lit-020; paper-result-b29f561ff4e41c3215; source ids: part2-fujisan-2024; source locator: Table 1: AUROC, Enzyme-pair functional identity: original held-out test; context: LightGBM FUJISAN and comparison methods; predictions thresholded to maximizeF 1. Protein-pair splitting is not evidence of disjoint proteins or families. 41,600 protein pairs; balanced functional-identity classes; random 56.25/18.75/25% training/validation/test split.; caveats: Input information differs: sequence similarity, protein embeddings, structure-aware descriptors and domain annotations.; Table1 reports point values;50bootstrap iterations described elsewhere do not establish a Table1 interval.; No interval assigned unless printed in source cell.; review: method: automated_source_review; date: 2026-09-17; id: part2-fujisan-2024-tbl0005-f8b070cff8; title: Enzyme-pair functional identity: original held-out test · Table 1; protocol id: paper-protocol-7088bcc2f0033b8a29; dataset id: reported-dataset-5197cca532f89d; metric: AUPR; unit: dimensionless; direction: higher; result ids: paper-result-baae3f4bd4e7d2eb06; paper-result-c9d11b0a9461bd483c; paper-result-ddfb253a50cc5308a8; paper-result-1dc6f9c6ff21f4263a; paper-result-f9f63e4d253b1af323; source ids: part2-fujisan-2024; source locator: Table 1: AUPR, Enzyme-pair functional identity: original held-out test; context: LightGBM FUJISAN and comparison methods; predictions thresholded to maximizeF 1. Protein-pair splitting is not evidence of disjoint proteins or families. 41,600 protein pairs; balanced functional-identity classes; random 56.25/18.75/25% training/validation/test split.; caveats: Input information differs: sequence similarity, protein embeddings, structure-aware descriptors and domain annotations.; Table1 reports point values;50bootstrap iterations described elsewhere do not establish a Table1 interval.; No interval assigned unless printed in source cell.; review: method: automated_source_review; date: 2026-09-17
benchmark research
review date: 2026-09-17; status: complete_comparison_tables_extracted_pending_publication_review; primary sources: part2-fujisan-2024; inspected locators: Table1; Model training and hyperparameter optimization; Performance assessment; searched queries: Enhanced prediction of protein functional identity through the integration of sequence and structural features primary paper benchmark results; gaps: Independent batch review before import; preserve existing observation identities.; claim scope: Primary-source discovery and table/protocol screening; source checked is not independently reproduced. Raw acquisitions not automatically numerical publication approval.
historical missing metadata
protocol version: not_reported_in_legacy_extract; split: not_reported_in_legacy_extract
metadata review scope
historical_missing_metadata preserves the original discovery state. Current descriptive evidence and missingness are recorded in profile.facts; numerical-result review is separate.
legacy kinds
benchmark
entity classification
review date: 2026-09-17; rationale: This source-scoped record identifies the biological prediction task and holds its paper context. Preserve the existing task identity; exact split, model adaptation and scoring remain in linked evaluations or separate protocol records.; source ids: fujisan-2024; source locator: Methods: Dataset construction; classification and metrics; Results: model comparison; cached text lines 10–11, 25, 28, 52, 55; uncertainty/repeat-run/statistical-comparison passages; ambiguities: A paper- or suite-specific task may constrain some inputs or metrics; that alone does not make it interchangeable with a complete versioned protocol. No protocol equivalence is inferred.; Some legacy profile Entity type facts use the generic phrase computational evaluation protocol. That boilerplate is not sufficient to establish a single fixed protocol identity or to merge this task with another protocol record.
Related records

Suggest a correction