Strengths supported by sources
No source-reviewed explanatory claims are recorded here yet.
Polyadenylation-site classification compares two background definitions and two adaptation regimes.
Explanatory profile: source reviewed · Automated source review, 2026-09-16. Review applies to the cited claims; unresolved fields are listed below. Numerical results retain their own review status.
| Property | Description and evidence |
|---|---|
| Datasets | GENCODE v38 human annotations with balanced Gene-Gene and Gene-Intergene datasets; a separate mouse transfer evaluation.SourcesPolyA-GLM: A comprehensive framework for De novo polyadenylation site prediction using genome language models · Methods: dataset construction and evaluation; Results Table 1; cached text lines 12–15, 37, 49–51 |
| Splits | Five-fold evaluation with 60:20:20 training/validation/test proportions per fold.SourcesPolyA-GLM: A comprehensive framework for De novo polyadenylation site prediction using genome language models · Methods: dataset construction and evaluation; Results Table 1; cached text lines 12–15, 37, 49–51 |
| Metrics | Accuracy, precision, recall, F1 and AUC, averaged across folds.SourcesPolyA-GLM: A comprehensive framework for De novo polyadenylation site prediction using genome language models · Methods: dataset construction and evaluation; Results Table 1; cached text lines 12–15, 37, 49–51 |
| Baselines | DNABERT-2, Nucleotide Transformer and HyenaDNA under few-shot and fine-tuned conditions.SourcesPolyA-GLM: A comprehensive framework for De novo polyadenylation site prediction using genome language models · Methods: dataset construction and evaluation; Results Table 1; cached text lines 12–15, 37, 49–51 |
| Leakage controls | Negative examples exclude genic regions overlapping GENCODE v38 polyadenylation annotations. This is a label-construction control; the evaluation description does not establish chromosome or homologous-sequence separation between the five-fold training and test partitions.SourcesPolyA-GLM: A comprehensive framework for De novo polyadenylation site prediction using genome language models · Methods: dataset construction and evaluation framework |
| Uncertainty | Table 1 reports averages over five folds. For few-shot predictions, ten random prototype pairs per fold are averaged. No confidence interval or standard-error estimator for the table’s aggregate metrics is specified there.SourcesPolyA-GLM: A comprehensive framework for De novo polyadenylation site prediction using genome language models · Methods: few-shot evaluation; Table 1 caption |
| Entity type | Paper-specific computational evaluation protocol.SourcesPolyA-GLM: A comprehensive framework for De novo polyadenylation site prediction using genome language models · Methods: dataset construction and evaluation; Results Table 1; cached text lines 12–15, 37, 49–51 |
| Organisms | Human; separate transfer evaluation in mouse.SourcesPolyA-GLM: A comprehensive framework for De novo polyadenylation site prediction using genome language models · Methods: dataset construction and evaluation; Results Table 1; cached text lines 12–15, 37, 49–51 |
| Assays | GENCODE polyadenylation annotations.SourcesPolyA-GLM: A comprehensive framework for De novo polyadenylation site prediction using genome language models · Methods: dataset construction and evaluation; Results Table 1; cached text lines 12–15, 37, 49–51 |
| Allowed inputs | DNA/RNA sequence context for candidate polyadenylation sites.SourcesPolyA-GLM: A comprehensive framework for De novo polyadenylation site prediction using genome language models · Methods: dataset construction and evaluation; Results Table 1; cached text lines 12–15, 37, 49–51 |
| Adaptation | Few-shot and task fine-tuning settings are compared separately.SourcesPolyA-GLM: A comprehensive framework for De novo polyadenylation site prediction using genome language models · Methods: dataset construction and evaluation; Results Table 1; cached text lines 12–15, 37, 49–51 |
Conceptual summary of the cited evaluation; exact task configuration and source version remain part of the protocol.
GENCODE v38 human annotations with balanced Gene-Gene and Gene-Intergene datasets; a separate mouse transfer evaluation. Five-fold evaluation with 60:20:20 training/validation/test proportions per fold. Accuracy, precision, recall, F1 and AUC, averaged across folds. DNABERT-2, Nucleotide Transformer and HyenaDNA under few-shot and fine-tuned conditions. Negative examples exclude genic regions overlapping GENCODE v38 polyadenylation annotations. This is a label-construction control; the evaluation description does not establish chromosome or homologous-sequence separation between the five-fold training and test partitions.
Each evaluation records what was tested and under which conditions.
Release 2026-09-17-d277315f7d76 · 1 evaluation · 1 metric row. Different protocols are not a single leaderboard.
| Metric and finding | Coverage and uncertainty | Evidence |
|---|---|---|
| HyenaDNA: polyadenylation site detection few-shot Gene-Gene negative-set comparison Independent external evaluation · Evaluation metadata: needs review | ||
| 0.7510 AUC Unit: fraction · Direction: unknown | Uncertainty: not reported in legacy extract Scored: Not reported · Eligible: Not reported | source checkedPolyA-GLM: A comprehensive framework for De novo polyadenylation site prediction using genome language models · Table 1, Few-shot HyenaDNA row, G-G AUC column Source checking is not independent reproduction. |
Last literature check: 2026-09-17. Primary-source discovery and table/protocol screening; source checked is not independently reproduced. Raw acquisitions not automatically numerical publication approval.
| Paper or primary resource | Version | Reference |
|---|---|---|
| PolyA-GLM: A comprehensive framework for De novo polyadenylation site prediction using genome language models | journal full text in PMC | Read source DOI: 10.1016/j.csbj.2025.12.011 |
source found structured extraction pending
No source-reviewed explanatory claims are recorded here yet.
Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.
Stable record: reported-task-13dfe6b33e71edTrace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.
One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.
17 evidence rows matching the loaded filters
| Property and statement | Original source and location | Review and provenance |
|---|---|---|
| Diagram caption Conceptual summary of the cited evaluation; exact task configuration and source version remain part of the protocol. Individual claims | PolyA-GLM: A comprehensive framework for De novo polyadenylation site prediction using genome language models Methods: dataset construction and evaluation; Results Table 1; cached text lines 12–15, 37, 49–51 Version: journal full text in PMC | source checked automated source review · 2026-09-16 Audit detailsTargeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Diagram steps ["Input: DNA/RNA sequence context for candidate polyadenylation sites.","Evaluation: Five-fold evaluation with 60:20:20 training/validation/test proportions per fold.","Readout: Accuracy, precision, recall, F1 and AUC, averaged across folds."] Individual claims | PolyA-GLM: A comprehensive framework for De novo polyadenylation site prediction using genome language models Methods: dataset construction and evaluation; Results Table 1; cached text lines 12–15, 37, 49–51 Version: journal full text in PMC | source checked automated source review · 2026-09-16 Audit detailsTargeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Diagram title Computational evaluation flow Individual claims | PolyA-GLM: A comprehensive framework for De novo polyadenylation site prediction using genome language models Methods: dataset construction and evaluation; Results Table 1; cached text lines 12–15, 37, 49–51 Version: journal full text in PMC | source checked automated source review · 2026-09-16 Audit detailsTargeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Datasets GENCODE v38 human annotations with balanced Gene-Gene and Gene-Intergene datasets; a separate mouse transfer evaluation. Individual claims | PolyA-GLM: A comprehensive framework for De novo polyadenylation site prediction using genome language models Methods: dataset construction and evaluation; Results Table 1; cached text lines 12–15, 37, 49–51 Version: journal full text in PMC | source checked automated source review · 2026-09-16 Audit detailsTargeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Splits Five-fold evaluation with 60:20:20 training/validation/test proportions per fold. Individual claims | PolyA-GLM: A comprehensive framework for De novo polyadenylation site prediction using genome language models Methods: dataset construction and evaluation; Results Table 1; cached text lines 12–15, 37, 49–51 Version: journal full text in PMC | source checked automated source review · 2026-09-16 Audit detailsTargeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Adaptation Few-shot and task fine-tuning settings are compared separately. Individual claims | PolyA-GLM: A comprehensive framework for De novo polyadenylation site prediction using genome language models Methods: dataset construction and evaluation; Results Table 1; cached text lines 12–15, 37, 49–51 Version: journal full text in PMC | source checked automated source review · 2026-09-16 Audit detailsTargeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Metrics Accuracy, precision, recall, F1 and AUC, averaged across folds. Individual claims | PolyA-GLM: A comprehensive framework for De novo polyadenylation site prediction using genome language models Methods: dataset construction and evaluation; Results Table 1; cached text lines 12–15, 37, 49–51 Version: journal full text in PMC | source checked automated source review · 2026-09-16 Audit detailsTargeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Baselines DNABERT-2, Nucleotide Transformer and HyenaDNA under few-shot and fine-tuned conditions. Individual claims | PolyA-GLM: A comprehensive framework for De novo polyadenylation site prediction using genome language models Methods: dataset construction and evaluation; Results Table 1; cached text lines 12–15, 37, 49–51 Version: journal full text in PMC | source checked automated source review · 2026-09-16 Audit detailsTargeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Leakage controls Negative examples exclude genic regions overlapping GENCODE v38 polyadenylation annotations. This is a label-construction control; the evaluation description does not establish chromosome or homologous-sequence separation between the five-fold training and test partitions. Individual claims | PolyA-GLM: A comprehensive framework for De novo polyadenylation site prediction using genome language models Methods: dataset construction and evaluation framework Version: journal full text in PMC | source checked automated source review · 2026-09-16 Audit detailsTargeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Uncertainty Table 1 reports averages over five folds. For few-shot predictions, ten random prototype pairs per fold are averaged. No confidence interval or standard-error estimator for the table’s aggregate metrics is specified there. Individual claims | PolyA-GLM: A comprehensive framework for De novo polyadenylation site prediction using genome language models Methods: few-shot evaluation; Table 1 caption Version: journal full text in PMC | source checked automated source review · 2026-09-16 Audit detailsTargeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
Release 2026-09-17-d277315f7d76 · Record review: needs review
Stable ID: reported-task-13dfe6b33e71ed