rewire.it
Task

polyadenylation site detection

Polyadenylation-site classification compares two background definitions and two adaptation regimes.

SourcesPolyA-GLM: A comprehensive framework for De novo polyadenylation site prediction using genome language models · Methods: dataset construction and evaluation; Results Table 1; cached text lines 12–15, 37, 49–51

1 evaluation · 1 metric row

At a glance

Inputs, training, access and other details

Explanatory profile: source reviewed · Automated source review, 2026-09-16. Review applies to the cited claims; unresolved fields are listed below. Numerical results retain their own review status.

Data, procedure and scoring
PropertyDescription and evidence
DatasetsGENCODE v38 human annotations with balanced Gene-Gene and Gene-Intergene datasets; a separate mouse transfer evaluation.
SourcesPolyA-GLM: A comprehensive framework for De novo polyadenylation site prediction using genome language models · Methods: dataset construction and evaluation; Results Table 1; cached text lines 12–15, 37, 49–51
SplitsFive-fold evaluation with 60:20:20 training/validation/test proportions per fold.
SourcesPolyA-GLM: A comprehensive framework for De novo polyadenylation site prediction using genome language models · Methods: dataset construction and evaluation; Results Table 1; cached text lines 12–15, 37, 49–51
MetricsAccuracy, precision, recall, F1 and AUC, averaged across folds.
SourcesPolyA-GLM: A comprehensive framework for De novo polyadenylation site prediction using genome language models · Methods: dataset construction and evaluation; Results Table 1; cached text lines 12–15, 37, 49–51
BaselinesDNABERT-2, Nucleotide Transformer and HyenaDNA under few-shot and fine-tuned conditions.
SourcesPolyA-GLM: A comprehensive framework for De novo polyadenylation site prediction using genome language models · Methods: dataset construction and evaluation; Results Table 1; cached text lines 12–15, 37, 49–51
Leakage controlsNegative examples exclude genic regions overlapping GENCODE v38 polyadenylation annotations. This is a label-construction control; the evaluation description does not establish chromosome or homologous-sequence separation between the five-fold training and test partitions.
SourcesPolyA-GLM: A comprehensive framework for De novo polyadenylation site prediction using genome language models · Methods: dataset construction and evaluation framework
UncertaintyTable 1 reports averages over five folds. For few-shot predictions, ten random prototype pairs per fold are averaged. No confidence interval or standard-error estimator for the table’s aggregate metrics is specified there.
SourcesPolyA-GLM: A comprehensive framework for De novo polyadenylation site prediction using genome language models · Methods: few-shot evaluation; Table 1 caption
Entity typePaper-specific computational evaluation protocol.
SourcesPolyA-GLM: A comprehensive framework for De novo polyadenylation site prediction using genome language models · Methods: dataset construction and evaluation; Results Table 1; cached text lines 12–15, 37, 49–51
OrganismsHuman; separate transfer evaluation in mouse.
SourcesPolyA-GLM: A comprehensive framework for De novo polyadenylation site prediction using genome language models · Methods: dataset construction and evaluation; Results Table 1; cached text lines 12–15, 37, 49–51
AssaysGENCODE polyadenylation annotations.
SourcesPolyA-GLM: A comprehensive framework for De novo polyadenylation site prediction using genome language models · Methods: dataset construction and evaluation; Results Table 1; cached text lines 12–15, 37, 49–51
Allowed inputsDNA/RNA sequence context for candidate polyadenylation sites.
SourcesPolyA-GLM: A comprehensive framework for De novo polyadenylation site prediction using genome language models · Methods: dataset construction and evaluation; Results Table 1; cached text lines 12–15, 37, 49–51
AdaptationFew-shot and task fine-tuning settings are compared separately.
SourcesPolyA-GLM: A comprehensive framework for De novo polyadenylation site prediction using genome language models · Methods: dataset construction and evaluation; Results Table 1; cached text lines 12–15, 37, 49–51

How it works

How it worksComputational evaluation flow
Computational evaluation flow1. Input: DNA/RNA sequence context for candidate polyadenylation sites.. Then: 2. Evaluation: Five-fold evaluation with 60:20:20 training/validation/test proportions per fold.. Then: 3. Readout: Accuracy, precision, recall, F1 and AUC, averaged across folds.Computational evaluation flow1. Input: DNA/RNA sequence context for candidate polyadenylation sites.. Then: 2. Evaluation: Five-fold evaluation with 60:20:20 training/validation/test proportions per fold.. Then: 3. Readout: Accuracy, precision, recall, F1 and AUC, averaged across folds.Computational evaluation flow1. Input: DNA/RNA sequence context for candidate polyadenylation sites.. Then: 2. Evaluation: Five-fold evaluation with 60:20:20 training/validation/test proportions per fold.. Then: 3. Readout: Accuracy, precision, recall, F1 and AUC, averaged across folds.

Conceptual summary of the cited evaluation; exact task configuration and source version remain part of the protocol.

SourcesPolyA-GLM: A comprehensive framework for De novo polyadenylation site prediction using genome language models · Methods: dataset construction and evaluation; Results Table 1; cached text lines 12–15, 37, 49–51
Evaluation methodology

GENCODE v38 human annotations with balanced Gene-Gene and Gene-Intergene datasets; a separate mouse transfer evaluation. Five-fold evaluation with 60:20:20 training/validation/test proportions per fold. Accuracy, precision, recall, F1 and AUC, averaged across folds. DNABERT-2, Nucleotide Transformer and HyenaDNA under few-shot and fine-tuned conditions. Negative examples exclude genic regions overlapping GENCODE v38 polyadenylation annotations. This is a label-construction control; the evaluation description does not establish chromosome or homologous-sequence separation between the five-fold training and test partitions.

SourcesPolyA-GLM: A comprehensive framework for De novo polyadenylation site prediction using genome language models · Methods: dataset construction and evaluation; Results Table 1; cached text lines 12–15, 37, 49–51; Methods: dataset construction and evaluation framework

Recorded evaluations

Each evaluation records what was tested and under which conditions.

Tested entities and results

Release 2026-09-17-d277315f7d76 · 1 evaluation · 1 metric row. Different protocols are not a single leaderboard.

Results grouped by the exact reported evaluation
Metric and findingCoverage and uncertaintyEvidence
HyenaDNA: polyadenylation site detection

few-shot Gene-Gene negative-set comparison

Independent external evaluation · Evaluation metadata: needs review

0.7510 AUC

Unit: fraction · Direction: unknown

Uncertainty: not reported in legacy extract

Scored: Not reported · Eligible: Not reported

source checkedPolyA-GLM: A comprehensive framework for De novo polyadenylation site prediction using genome language models · Table 1, Few-shot HyenaDNA row, G-G AUC column

Source checking is not independent reproduction.

Papers and result coverage

Last literature check: 2026-09-17. Primary-source discovery and table/protocol screening; source checked is not independently reproduced. Raw acquisitions not automatically numerical publication approval.

Paper or primary resourceVersionReference
PolyA-GLM: A comprehensive framework for De novo polyadenylation site prediction using genome language modelsjournal full text in PMCRead source
DOI: 10.1016/j.csbj.2025.12.011

What is still missing

  • Raw complete comparison acquired. Few-shot versus fine-tuning and Gene-Gene versus Intergenic-Gene conditions cannot share a chart group; NT100M and 500M are different configurations. Structured extraction pending.
Search and extraction details

source found structured extraction pending

Searches

  • PolyA-GLM: A comprehensive framework for De novo polyadenylation site prediction using genome language models primary paper benchmark results

Evidence locations

  • Table1 two negative-set conditions and few-shot/fine-tuning row groups

Strengths and limitations

Strengths supported by sources

No source-reviewed explanatory claims are recorded here yet.

Profile review details

Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.

Stable record: reported-task-13dfe6b33e71ed

Evidence table

Inspect claims, sources and review details

Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.

One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.

17 evidence rows matching the loaded filters

Claims, original sources and review scope · Release 2026-09-17-d277315f7d76
Property and statementOriginal source and locationReview and provenance
Diagram caption

Conceptual summary of the cited evaluation; exact task configuration and source version remain part of the protocol.

Individual claims
PolyA-GLM: A comprehensive framework for De novo polyadenylation site prediction using genome language models

Original source ↗

Methods: dataset construction and evaluation; Results Table 1; cached text lines 12–15, 37, 49–51

Version: journal full text in PMC
Retrieved: 2026-09-16T10:38:57.558220+00:00

source checked

automated source review · 2026-09-16

Audit details

Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.

Field: attributes.profile.diagram.caption

Source artifact SHA-256: e9ebd53d88837ad8d457881ffee918d2734dcae87d3c5cd03135947b6cf5dbde

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Diagram steps

["Input: DNA/RNA sequence context for candidate polyadenylation sites.","Evaluation: Five-fold evaluation with 60:20:20 training/validation/test proportions per fold.","Readout: Accuracy, precision, recall, F1 and AUC, averaged across folds."]

Individual claims
PolyA-GLM: A comprehensive framework for De novo polyadenylation site prediction using genome language models

Original source ↗

Methods: dataset construction and evaluation; Results Table 1; cached text lines 12–15, 37, 49–51

Version: journal full text in PMC
Retrieved: 2026-09-16T10:38:57.558220+00:00

source checked

automated source review · 2026-09-16

Audit details

Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.

Field: attributes.profile.diagram.steps

Source artifact SHA-256: e9ebd53d88837ad8d457881ffee918d2734dcae87d3c5cd03135947b6cf5dbde

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Diagram title

Computational evaluation flow

Individual claims
PolyA-GLM: A comprehensive framework for De novo polyadenylation site prediction using genome language models

Original source ↗

Methods: dataset construction and evaluation; Results Table 1; cached text lines 12–15, 37, 49–51

Version: journal full text in PMC
Retrieved: 2026-09-16T10:38:57.558220+00:00

source checked

automated source review · 2026-09-16

Audit details

Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.

Field: attributes.profile.diagram.title

Source artifact SHA-256: e9ebd53d88837ad8d457881ffee918d2734dcae87d3c5cd03135947b6cf5dbde

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Datasets

GENCODE v38 human annotations with balanced Gene-Gene and Gene-Intergene datasets; a separate mouse transfer evaluation.

Individual claims
PolyA-GLM: A comprehensive framework for De novo polyadenylation site prediction using genome language models

Original source ↗

Methods: dataset construction and evaluation; Results Table 1; cached text lines 12–15, 37, 49–51

Version: journal full text in PMC
Retrieved: 2026-09-16T10:38:57.558220+00:00

source checked

automated source review · 2026-09-16

Audit details

Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.

Field: attributes.profile.facts.0.value

Source artifact SHA-256: e9ebd53d88837ad8d457881ffee918d2734dcae87d3c5cd03135947b6cf5dbde

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Splits

Five-fold evaluation with 60:20:20 training/validation/test proportions per fold.

Individual claims
PolyA-GLM: A comprehensive framework for De novo polyadenylation site prediction using genome language models

Original source ↗

Methods: dataset construction and evaluation; Results Table 1; cached text lines 12–15, 37, 49–51

Version: journal full text in PMC
Retrieved: 2026-09-16T10:38:57.558220+00:00

source checked

automated source review · 2026-09-16

Audit details

Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.

Field: attributes.profile.facts.1.value

Source artifact SHA-256: e9ebd53d88837ad8d457881ffee918d2734dcae87d3c5cd03135947b6cf5dbde

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Adaptation

Few-shot and task fine-tuning settings are compared separately.

Individual claims
PolyA-GLM: A comprehensive framework for De novo polyadenylation site prediction using genome language models

Original source ↗

Methods: dataset construction and evaluation; Results Table 1; cached text lines 12–15, 37, 49–51

Version: journal full text in PMC
Retrieved: 2026-09-16T10:38:57.558220+00:00

source checked

automated source review · 2026-09-16

Audit details

Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.

Field: attributes.profile.facts.10.value

Source artifact SHA-256: e9ebd53d88837ad8d457881ffee918d2734dcae87d3c5cd03135947b6cf5dbde

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Metrics

Accuracy, precision, recall, F1 and AUC, averaged across folds.

Individual claims
PolyA-GLM: A comprehensive framework for De novo polyadenylation site prediction using genome language models

Original source ↗

Methods: dataset construction and evaluation; Results Table 1; cached text lines 12–15, 37, 49–51

Version: journal full text in PMC
Retrieved: 2026-09-16T10:38:57.558220+00:00

source checked

automated source review · 2026-09-16

Audit details

Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.

Field: attributes.profile.facts.2.value

Source artifact SHA-256: e9ebd53d88837ad8d457881ffee918d2734dcae87d3c5cd03135947b6cf5dbde

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Baselines

DNABERT-2, Nucleotide Transformer and HyenaDNA under few-shot and fine-tuned conditions.

Individual claims
PolyA-GLM: A comprehensive framework for De novo polyadenylation site prediction using genome language models

Original source ↗

Methods: dataset construction and evaluation; Results Table 1; cached text lines 12–15, 37, 49–51

Version: journal full text in PMC
Retrieved: 2026-09-16T10:38:57.558220+00:00

source checked

automated source review · 2026-09-16

Audit details

Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.

Field: attributes.profile.facts.3.value

Source artifact SHA-256: e9ebd53d88837ad8d457881ffee918d2734dcae87d3c5cd03135947b6cf5dbde

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Leakage controls

Negative examples exclude genic regions overlapping GENCODE v38 polyadenylation annotations. This is a label-construction control; the evaluation description does not establish chromosome or homologous-sequence separation between the five-fold training and test partitions.

Individual claims
PolyA-GLM: A comprehensive framework for De novo polyadenylation site prediction using genome language models

Original source ↗

Methods: dataset construction and evaluation framework

Version: journal full text in PMC
Retrieved: 2026-09-16T10:38:57.558220+00:00

source checked

automated source review · 2026-09-16

Audit details

Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.

Field: attributes.profile.facts.4.value

Source artifact SHA-256: e9ebd53d88837ad8d457881ffee918d2734dcae87d3c5cd03135947b6cf5dbde

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Uncertainty

Table 1 reports averages over five folds. For few-shot predictions, ten random prototype pairs per fold are averaged. No confidence interval or standard-error estimator for the table’s aggregate metrics is specified there.

Individual claims
PolyA-GLM: A comprehensive framework for De novo polyadenylation site prediction using genome language models

Original source ↗

Methods: few-shot evaluation; Table 1 caption

Version: journal full text in PMC
Retrieved: 2026-09-16T10:38:57.558220+00:00

source checked

automated source review · 2026-09-16

Audit details

Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.

Field: attributes.profile.facts.5.value

Source artifact SHA-256: e9ebd53d88837ad8d457881ffee918d2734dcae87d3c5cd03135947b6cf5dbde

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Sources and history

Release 2026-09-17-d277315f7d76 · Record review: needs review

2 source records and release historyDownload this release
Technical metadata and extraction receipts

Stable ID: reported-task-13dfe6b33e71ed

areas
dna-genomes
tasks
polyadenylation site detection
entity level
task
version
Not reported
task
polyadenylation site detection
scope note
Paper-specific evaluation task; protocol completeness requires further extraction.
benchmark research
review date: 2026-09-17; status: source_found_structured_extraction_pending; primary sources: evidence-expansion-p2-polya-glm-2025-e9ebd53d8883; inspected locators: Table1 two negative-set conditions and few-shot/fine-tuning row groups; searched queries: PolyA-GLM: A comprehensive framework for De novo polyadenylation site prediction using genome language models primary paper benchmark results; gaps: Raw complete comparison acquired. Few-shot versus fine-tuning and Gene-Gene versus Intergenic-Gene conditions cannot share a chart group; NT100M and 500M are different configurations. Structured extraction pending.; claim scope: Primary-source discovery and table/protocol screening; source checked is not independently reproduced. Raw acquisitions not automatically numerical publication approval.
historical missing metadata
protocol version: not_reported_in_legacy_extract
metadata review scope
historical_missing_metadata preserves the original discovery state. Current descriptive evidence and missingness are recorded in profile.facts; numerical-result review is separate.
legacy kinds
benchmark
entity classification
review date: 2026-09-17; rationale: This source-scoped record identifies the biological prediction task and holds its paper context. Preserve the existing task identity; exact split, model adaptation and scoring remain in linked evaluations or separate protocol records.; source ids: polya-glm-2025; source locator: Methods: dataset construction and evaluation; Results Table 1; cached text lines 12–15, 37, 49–51; ambiguities: A paper- or suite-specific task may constrain some inputs or metrics; that alone does not make it interchangeable with a complete versioned protocol. No protocol equivalence is inferred.; Some legacy profile Entity type facts use the generic phrase computational evaluation protocol. That boilerplate is not sufficient to establish a single fixed protocol identity or to merge this task with another protocol record.
Related records

Suggest a correction