Strengths and considerations
No source-reviewed explanatory claims are recorded here yet.
Enhancer classification is evaluated within two published genomic sequence benchmark collections.
Explanatory profile: limited source coverage · Automated source review, 2026-09-16. Review applies to the cited claims; unresolved fields are listed below. Numerical results retain their own review status.
| Property | Description and evidence |
|---|---|
| Datasets | Genomic Benchmarks includes human enhancer datasets; Nucleotide Transformer tasks include enhancer/non-enhancer and enhancer-strength labels.SourcesUnderstanding the natural language of DNA using encoder–decoder foundation models with byte-level precision · Results §1.2.1; Methods §§2.5.1–2.5.2; cached text lines 18–19, 70–74; matching task comparison table/ablation captions |
| Splits | For the Nucleotide Transformer enhancer tasks, Supplementary Table 1 lists 14,968 training sequences and 400 test sequences of length 200. This set combines the original strong/weak/non-enhancer collection with 6,000 synthetic enhancers and 6,000 synthetic non-enhancers. Those counts do not describe the separate Genomic Benchmarks enhancer datasets.Sources (2)Understanding the natural language of DNA using encoder–decoder foundation models with byte-level precision; enbed-2024__vbae117_supplementary_data.pdf · Supplement: Nucleotide Transformer benchmark dataset description and Table 1 |
| Metrics | Classification accuracy is described for the benchmark comparison.SourcesUnderstanding the natural language of DNA using encoder–decoder foundation models with byte-level precision · Results §1.2.1; Methods §§2.5.1–2.5.2; cached text lines 18–19, 70–74; matching task comparison table/ablation captions |
| Baselines | Enformer, DNABERT-2, Nucleotide Transformer and HyenaDNA in the enhancer-containing NT task comparison.SourcesUnderstanding the natural language of DNA using encoder–decoder foundation models with byte-level precision · Results §1.2.1; Methods §§2.5.1–2.5.2; cached text lines 18–19, 70–74; matching task comparison table/ablation captions |
| Leakage controls | The enhancer-specific methods and supplement do not specify a chromosome holdout, homology threshold or synthetic-parent separation across these splits. The clade split and 95% similarity threshold described later in the supplement apply to mutation generation, not this enhancer evaluation. · Not reported in inspected sourcesSources (2)Understanding the natural language of DNA using encoder–decoder foundation models with byte-level precision; enbed-2024__vbae117_supplementary_data.pdf · Supplement: Nucleotide Transformer benchmark description and Table 1; separate Mutation generation section |
| Uncertainty | The cited text-accessible evaluation sections give no confidence-interval, resampling or repeat-run error-bar specification. Image-only tables and uninspected supplements are outside this absence claim. · Not reported in inspected sourcesSourcesUnderstanding the natural language of DNA using encoder–decoder foundation models with byte-level precision · Results §1.2.1; Methods §§2.5.1–2.5.2; cached text lines 18–19, 70–74; matching task comparison table/ablation captions |
| Entity type | Paper-specific computational evaluation protocol.SourcesUnderstanding the natural language of DNA using encoder–decoder foundation models with byte-level precision · Results §1.2.1; Methods §§2.5.1–2.5.2; cached text lines 18–19, 70–74; matching task comparison table/ablation captions |
| Organisms | Human enhancer datasets within Genomic Benchmarks and Nucleotide Transformer tasks.SourcesUnderstanding the natural language of DNA using encoder–decoder foundation models with byte-level precision · Results §1.2.1; Methods §§2.5.1–2.5.2; cached text lines 18–19, 70–74; matching task comparison table/ablation captions |
| Assays | Enhancer identity/strength annotations.SourcesUnderstanding the natural language of DNA using encoder–decoder foundation models with byte-level precision · Results §1.2.1; Methods §§2.5.1–2.5.2; cached text lines 18–19, 70–74; matching task comparison table/ablation captions |
| Allowed inputs | Genomic sequence.SourcesUnderstanding the natural language of DNA using encoder–decoder foundation models with byte-level precision · Results §1.2.1; Methods §§2.5.1–2.5.2; cached text lines 18–19, 70–74; matching task comparison table/ablation captions |
| Adaptation | Task-level enhancer classification; exact task configuration must be selected from the evaluated benchmark collection.SourcesUnderstanding the natural language of DNA using encoder–decoder foundation models with byte-level precision · Results §1.2.1; Methods §§2.5.1–2.5.2; cached text lines 18–19, 70–74; matching task comparison table/ablation captions |
Conceptual summary of the cited evaluation; exact task configuration and source version remain part of the protocol.
Genomic Benchmarks includes human enhancer datasets; Nucleotide Transformer tasks include enhancer/non-enhancer and enhancer-strength labels. For the Nucleotide Transformer enhancer tasks, Supplementary Table 1 lists 14,968 training sequences and 400 test sequences of length 200. This set combines the original strong/weak/non-enhancer collection with 6,000 synthetic enhancers and 6,000 synthetic non-enhancers. Those counts do not describe the separate Genomic Benchmarks enhancer datasets. Classification accuracy is described for the benchmark comparison. Enformer, DNABERT-2, Nucleotide Transformer and HyenaDNA in the enhancer-containing NT task comparison. The enhancer-specific methods and supplement do not specify a chromosome holdout, homology threshold or synthetic-parent separation across these splits. The clade split and 95% similarity threshold described later in the supplement apply to mutation generation, not this enhancer evaluation.
Each evaluation records what was tested and under which conditions.
Release 2026-09-17-d277315f7d76 · 2 evaluations · 2 metric rows. Different protocols are not a single leaderboard.
| Metric and finding | Coverage and uncertainty | Evidence |
|---|---|---|
| ENBED: Enhancer classification Reported Genomic Benchmarks classification accuracy. Author-reported evaluation · Evaluation metadata: needs review | ||
| 90.3 Accuracy Unit: % · Direction: unknown | Uncertainty: not reported in legacy extract Scored: Not reported · Eligible: Not reported | source checkedUnderstanding the natural language of DNA using encoder–decoder foundation models with byte-level precision · Table 2, Mouse Enhancers row, ENBED column Source checking is not independent reproduction. |
| ENBED (GRCh38): Enhancer classification Configuration: ENBED (GRCh38)Task: Enhancer classificationDataset: Genomic Benchmarks Mouse Enhancers ENBED trained on GRCh38; reported Genomic Benchmarks classification accuracy. Author-reported evaluation · Evaluation metadata: needs review | ||
| 81.1 Accuracy Unit: % · Direction: unknown | Uncertainty: not reported in legacy extract Scored: Not reported · Eligible: Not reported | source checkedUnderstanding the natural language of DNA using encoder–decoder foundation models with byte-level precision · Table 2, Mouse Enhancers row, ENBED (GRCh38) column Source checking is not independent reproduction. |
Last literature check: 2026-09-17. Dated primary-source discovery and protocol/table screening. Source checking does not mean experimental reproduction. Only separately extracted and independently reviewed numeric batches are publishable.
| Paper or primary resource | Version | Reference |
|---|---|---|
| Understanding the natural language of DNA using encoder–decoder foundation models with byte-level precision | version of record | Read source |
primary comparison tables located
No source-reviewed explanatory claims are recorded here yet.
Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.
Stable record: reported-task-132da895d4c381Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.
One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.
24 evidence rows matching the loaded filters
| Property and statement | Original source and location | Review and provenance |
|---|---|---|
| Diagram caption Conceptual summary of the cited evaluation; exact task configuration and source version remain part of the protocol. Individual claims | Understanding the natural language of DNA using encoder–decoder foundation models with byte-level precision Results §1.2.1; Methods §§2.5.1–2.5.2; cached text lines 18–19, 70–74; matching task comparison table/ablation captions; Supplement: Nucleotide Transformer benchmark dataset description and Table 1 Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: version of record | source checked automated source review · 2026-09-16 Audit detailsTargeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Diagram caption Conceptual summary of the cited evaluation; exact task configuration and source version remain part of the protocol. Individual claims | enbed-2024__vbae117_supplementary_data.pdf Results §1.2.1; Methods §§2.5.1–2.5.2; cached text lines 18–19, 70–74; matching task comparison table/ablation captions; Supplement: Nucleotide Transformer benchmark dataset description and Table 1 Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: Retrieved 2026-09-16; sha256:a169bbebba2298e9c98c33c28053c1a0b42d0c3e88e5a8795934725c4408cc32 | source checked automated source review · 2026-09-16 Audit detailsTargeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record Archive member: vbae117_supplementary_data.pdf |
| Diagram steps ["Input: Genomic sequence.","Evaluation: For the Nucleotide Transformer enhancer tasks, Supplementary Table 1 lists 14,968 training sequences and 400 test sequences of length 200. This set combines the original strong/weak/non-enhancer collection with 6,000 synthetic enhancers and 6,000 synthetic non-enhancers. Those counts do not describe the separate Genomic Benchmarks enhancer datasets.","Readout: Classification accuracy is described for the benchmark comparison."] Individual claims | Understanding the natural language of DNA using encoder–decoder foundation models with byte-level precision Results §1.2.1; Methods §§2.5.1–2.5.2; cached text lines 18–19, 70–74; matching task comparison table/ablation captions; Supplement: Nucleotide Transformer benchmark dataset description and Table 1 Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: version of record | source checked automated source review · 2026-09-16 Audit detailsTargeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Diagram steps ["Input: Genomic sequence.","Evaluation: For the Nucleotide Transformer enhancer tasks, Supplementary Table 1 lists 14,968 training sequences and 400 test sequences of length 200. This set combines the original strong/weak/non-enhancer collection with 6,000 synthetic enhancers and 6,000 synthetic non-enhancers. Those counts do not describe the separate Genomic Benchmarks enhancer datasets.","Readout: Classification accuracy is described for the benchmark comparison."] Individual claims | enbed-2024__vbae117_supplementary_data.pdf Results §1.2.1; Methods §§2.5.1–2.5.2; cached text lines 18–19, 70–74; matching task comparison table/ablation captions; Supplement: Nucleotide Transformer benchmark dataset description and Table 1 Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: Retrieved 2026-09-16; sha256:a169bbebba2298e9c98c33c28053c1a0b42d0c3e88e5a8795934725c4408cc32 | source checked automated source review · 2026-09-16 Audit detailsTargeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record Archive member: vbae117_supplementary_data.pdf |
| Diagram title Computational evaluation flow Individual claims | Understanding the natural language of DNA using encoder–decoder foundation models with byte-level precision Results §1.2.1; Methods §§2.5.1–2.5.2; cached text lines 18–19, 70–74; matching task comparison table/ablation captions; Supplement: Nucleotide Transformer benchmark dataset description and Table 1 Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: version of record | source checked automated source review · 2026-09-16 Audit detailsTargeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Diagram title Computational evaluation flow Individual claims | enbed-2024__vbae117_supplementary_data.pdf Results §1.2.1; Methods §§2.5.1–2.5.2; cached text lines 18–19, 70–74; matching task comparison table/ablation captions; Supplement: Nucleotide Transformer benchmark dataset description and Table 1 Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: Retrieved 2026-09-16; sha256:a169bbebba2298e9c98c33c28053c1a0b42d0c3e88e5a8795934725c4408cc32 | source checked automated source review · 2026-09-16 Audit detailsTargeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record Archive member: vbae117_supplementary_data.pdf |
| Datasets Genomic Benchmarks includes human enhancer datasets; Nucleotide Transformer tasks include enhancer/non-enhancer and enhancer-strength labels. Individual claims | Understanding the natural language of DNA using encoder–decoder foundation models with byte-level precision Results §1.2.1; Methods §§2.5.1–2.5.2; cached text lines 18–19, 70–74; matching task comparison table/ablation captions Version: version of record | source checked automated source review · 2026-09-16 Audit detailsTargeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Splits For the Nucleotide Transformer enhancer tasks, Supplementary Table 1 lists 14,968 training sequences and 400 test sequences of length 200. This set combines the original strong/weak/non-enhancer collection with 6,000 synthetic enhancers and 6,000 synthetic non-enhancers. Those counts do not describe the separate Genomic Benchmarks enhancer datasets. Individual claims | Understanding the natural language of DNA using encoder–decoder foundation models with byte-level precision Supplement: Nucleotide Transformer benchmark dataset description and Table 1 Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: version of record | source checked automated source review · 2026-09-16 Audit detailsTargeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Splits For the Nucleotide Transformer enhancer tasks, Supplementary Table 1 lists 14,968 training sequences and 400 test sequences of length 200. This set combines the original strong/weak/non-enhancer collection with 6,000 synthetic enhancers and 6,000 synthetic non-enhancers. Those counts do not describe the separate Genomic Benchmarks enhancer datasets. Individual claims | enbed-2024__vbae117_supplementary_data.pdf Supplement: Nucleotide Transformer benchmark dataset description and Table 1 Shared locator for this statement’s cited sources; not a separate locator for each citation. Version: Retrieved 2026-09-16; sha256:a169bbebba2298e9c98c33c28053c1a0b42d0c3e88e5a8795934725c4408cc32 | source checked automated source review · 2026-09-16 Audit detailsTargeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record Archive member: vbae117_supplementary_data.pdf |
| Adaptation Task-level enhancer classification; exact task configuration must be selected from the evaluated benchmark collection. Individual claims | Understanding the natural language of DNA using encoder–decoder foundation models with byte-level precision Results §1.2.1; Methods §§2.5.1–2.5.2; cached text lines 18–19, 70–74; matching task comparison table/ablation captions Version: version of record | source checked automated source review · 2026-09-16 Audit detailsTargeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
Release 2026-09-17-d277315f7d76 · Record review: needs review
Stable ID: reported-task-132da895d4c381