rewire.it
Task

Enhancer classification

Enhancer classification is evaluated within two published genomic sequence benchmark collections.

SourcesUnderstanding the natural language of DNA using encoder–decoder foundation models with byte-level precision · Results §1.2.1; Methods §§2.5.1–2.5.2; cached text lines 18–19, 70–74; matching task comparison table/ablation captions

2 evaluations · 2 metric rows

At a glance

Inputs, training, access and other details

Explanatory profile: limited source coverage · Automated source review, 2026-09-16. Review applies to the cited claims; unresolved fields are listed below. Numerical results retain their own review status.

Data, procedure and scoring
PropertyDescription and evidence
DatasetsGenomic Benchmarks includes human enhancer datasets; Nucleotide Transformer tasks include enhancer/non-enhancer and enhancer-strength labels.
SourcesUnderstanding the natural language of DNA using encoder–decoder foundation models with byte-level precision · Results §1.2.1; Methods §§2.5.1–2.5.2; cached text lines 18–19, 70–74; matching task comparison table/ablation captions
SplitsFor the Nucleotide Transformer enhancer tasks, Supplementary Table 1 lists 14,968 training sequences and 400 test sequences of length 200. This set combines the original strong/weak/non-enhancer collection with 6,000 synthetic enhancers and 6,000 synthetic non-enhancers. Those counts do not describe the separate Genomic Benchmarks enhancer datasets.
Sources (2)Understanding the natural language of DNA using encoder–decoder foundation models with byte-level precision; enbed-2024__vbae117_supplementary_data.pdf · Supplement: Nucleotide Transformer benchmark dataset description and Table 1
MetricsClassification accuracy is described for the benchmark comparison.
SourcesUnderstanding the natural language of DNA using encoder–decoder foundation models with byte-level precision · Results §1.2.1; Methods §§2.5.1–2.5.2; cached text lines 18–19, 70–74; matching task comparison table/ablation captions
BaselinesEnformer, DNABERT-2, Nucleotide Transformer and HyenaDNA in the enhancer-containing NT task comparison.
SourcesUnderstanding the natural language of DNA using encoder–decoder foundation models with byte-level precision · Results §1.2.1; Methods §§2.5.1–2.5.2; cached text lines 18–19, 70–74; matching task comparison table/ablation captions
Leakage controlsThe enhancer-specific methods and supplement do not specify a chromosome holdout, homology threshold or synthetic-parent separation across these splits. The clade split and 95% similarity threshold described later in the supplement apply to mutation generation, not this enhancer evaluation. · Not reported in inspected sources
Sources (2)Understanding the natural language of DNA using encoder–decoder foundation models with byte-level precision; enbed-2024__vbae117_supplementary_data.pdf · Supplement: Nucleotide Transformer benchmark description and Table 1; separate Mutation generation section
UncertaintyThe cited text-accessible evaluation sections give no confidence-interval, resampling or repeat-run error-bar specification. Image-only tables and uninspected supplements are outside this absence claim. · Not reported in inspected sources
SourcesUnderstanding the natural language of DNA using encoder–decoder foundation models with byte-level precision · Results §1.2.1; Methods §§2.5.1–2.5.2; cached text lines 18–19, 70–74; matching task comparison table/ablation captions
Entity typePaper-specific computational evaluation protocol.
SourcesUnderstanding the natural language of DNA using encoder–decoder foundation models with byte-level precision · Results §1.2.1; Methods §§2.5.1–2.5.2; cached text lines 18–19, 70–74; matching task comparison table/ablation captions
OrganismsHuman enhancer datasets within Genomic Benchmarks and Nucleotide Transformer tasks.
SourcesUnderstanding the natural language of DNA using encoder–decoder foundation models with byte-level precision · Results §1.2.1; Methods §§2.5.1–2.5.2; cached text lines 18–19, 70–74; matching task comparison table/ablation captions
AssaysEnhancer identity/strength annotations.
SourcesUnderstanding the natural language of DNA using encoder–decoder foundation models with byte-level precision · Results §1.2.1; Methods §§2.5.1–2.5.2; cached text lines 18–19, 70–74; matching task comparison table/ablation captions
Allowed inputsGenomic sequence.
SourcesUnderstanding the natural language of DNA using encoder–decoder foundation models with byte-level precision · Results §1.2.1; Methods §§2.5.1–2.5.2; cached text lines 18–19, 70–74; matching task comparison table/ablation captions
AdaptationTask-level enhancer classification; exact task configuration must be selected from the evaluated benchmark collection.
SourcesUnderstanding the natural language of DNA using encoder–decoder foundation models with byte-level precision · Results §1.2.1; Methods §§2.5.1–2.5.2; cached text lines 18–19, 70–74; matching task comparison table/ablation captions

How it works

How it worksComputational evaluation flow
Computational evaluation flow1. Input: Genomic sequence.. Then: 2. Evaluation: For the Nucleotide Transformer enhancer tasks, Supplementary Table 1 lists 14,968 training sequences and 400 test sequences of length 200. This set combines the original strong/weak/non-enhancer collection with 6,000 synthetic enhancers and 6,000 synthetic non-enhancers. Those counts do not describe the separate Genomic Benchmarks enhancer datasets.. Then: 3. Readout: Classification accuracy is described for the benchmark comparison.Computational evaluation flow1. Input: Genomic sequence.. Then: 2. Evaluation: For the Nucleotide Transformer enhancer tasks, Supplementary Table 1 lists 14,968 training sequences and 400 test sequences of length 200. This set combines the original strong/weak/non-enhancer collection with 6,000 synthetic enhancers and 6,000 synthetic non-enhancers. Those counts do not describe the separate Genomic Benchmarks enhancer datasets.. Then: 3. Readout: Classification accuracy is described for the benchmark comparison.Computational evaluation flow1. Input: Genomic sequence.. Then: 2. Evaluation: For the Nucleotide Transformer enhancer tasks, Supplementary Table 1 lists 14,968 training sequences and 400 test sequences of length 200. This set combines the original strong/weak/non-enhancer collection with 6,000 synthetic enhancers and 6,000 synthetic non-enhancers. Those counts do not describe the separate Genomic Benchmarks enhancer datasets.. Then: 3. Readout: Classification accuracy is described for the benchmark comparison.

Conceptual summary of the cited evaluation; exact task configuration and source version remain part of the protocol.

Sources (2)Understanding the natural language of DNA using encoder–decoder foundation models with byte-level precision; enbed-2024__vbae117_supplementary_data.pdf · Results §1.2.1; Methods §§2.5.1–2.5.2; cached text lines 18–19, 70–74; matching task comparison table/ablation captions; Supplement: Nucleotide Transformer benchmark dataset description and Table 1
Evaluation methodology

Genomic Benchmarks includes human enhancer datasets; Nucleotide Transformer tasks include enhancer/non-enhancer and enhancer-strength labels. For the Nucleotide Transformer enhancer tasks, Supplementary Table 1 lists 14,968 training sequences and 400 test sequences of length 200. This set combines the original strong/weak/non-enhancer collection with 6,000 synthetic enhancers and 6,000 synthetic non-enhancers. Those counts do not describe the separate Genomic Benchmarks enhancer datasets. Classification accuracy is described for the benchmark comparison. Enformer, DNABERT-2, Nucleotide Transformer and HyenaDNA in the enhancer-containing NT task comparison. The enhancer-specific methods and supplement do not specify a chromosome holdout, homology threshold or synthetic-parent separation across these splits. The clade split and 95% similarity threshold described later in the supplement apply to mutation generation, not this enhancer evaluation.

Sources (2)Understanding the natural language of DNA using encoder–decoder foundation models with byte-level precision; enbed-2024__vbae117_supplementary_data.pdf · Results §1.2.1; Methods §§2.5.1–2.5.2; cached text lines 18–19, 70–74; matching task comparison table/ablation captions; Supplement: Nucleotide Transformer benchmark dataset description and Table 1; Supplement: Nucleotide Transformer benchmark description and Table 1; separate Mutation generation section

Recorded evaluations

Each evaluation records what was tested and under which conditions.

Tested entities and results

Release 2026-09-17-d277315f7d76 · 2 evaluations · 2 metric rows. Different protocols are not a single leaderboard.

Results grouped by the exact reported evaluation
Metric and findingCoverage and uncertaintyEvidence
ENBED: Enhancer classification

Reported Genomic Benchmarks classification accuracy.

Author-reported evaluation · Evaluation metadata: needs review

90.3 Accuracy

Unit: % · Direction: unknown

Uncertainty: not reported in legacy extract

Scored: Not reported · Eligible: Not reported

source checkedUnderstanding the natural language of DNA using encoder–decoder foundation models with byte-level precision · Table 2, Mouse Enhancers row, ENBED column

Source checking is not independent reproduction.

ENBED (GRCh38): Enhancer classification

ENBED trained on GRCh38; reported Genomic Benchmarks classification accuracy.

Author-reported evaluation · Evaluation metadata: needs review

81.1 Accuracy

Unit: % · Direction: unknown

Uncertainty: not reported in legacy extract

Scored: Not reported · Eligible: Not reported

source checkedUnderstanding the natural language of DNA using encoder–decoder foundation models with byte-level precision · Table 2, Mouse Enhancers row, ENBED (GRCh38) column

Source checking is not independent reproduction.

Papers and result coverage

Last literature check: 2026-09-17. Dated primary-source discovery and protocol/table screening. Source checking does not mean experimental reproduction. Only separately extracted and independently reviewed numeric batches are publishable.

What is still missing

  • complete numerical transcription and independent cell review: Full primary artifact and table inventory preserved; no new numeric row is published from this audit alone.
  • exact checkpoint hashes and per-method scored denominators: Table labels alone do not establish these fields; do not infer checkpoint or scored count from model name or dataset size.
Search and extraction details

primary comparison tables located

Searches

  • Understanding the natural language of DNA using encoder–decoder foundation models with byte-level precision 10.1093/bioadv/vbae117

Evidence locations

  • Table 1.; XML table vbae117-T1
  • Table 2.; XML table vbae117-T2

Strengths and limitations

Strengths and considerations

No source-reviewed explanatory claims are recorded here yet.

Limitations and conditions

Profile review details

Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.

Stable record: reported-task-132da895d4c381

Evidence table

Inspect claims, sources and review details

Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.

One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.

24 evidence rows matching the loaded filters

Claims, original sources and review scope · Release 2026-09-17-d277315f7d76
Property and statementOriginal source and locationReview and provenance
Diagram caption

Conceptual summary of the cited evaluation; exact task configuration and source version remain part of the protocol.

Individual claims
Understanding the natural language of DNA using encoder–decoder foundation models with byte-level precision

Original source ↗

Results §1.2.1; Methods §§2.5.1–2.5.2; cached text lines 18–19, 70–74; matching task comparison table/ablation captions; Supplement: Nucleotide Transformer benchmark dataset description and Table 1

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: version of record
Retrieved: 2026-09-16T10:33:35.378Z

source checked

automated source review · 2026-09-16

Audit details

Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.

Field: attributes.profile.diagram.caption

Source artifact SHA-256: e95d4be70d32e61af5a92eda8ea66f25a2cc629e3f83e7b5241b13cde8bdb83b

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Diagram caption

Conceptual summary of the cited evaluation; exact task configuration and source version remain part of the protocol.

Individual claims
enbed-2024__vbae117_supplementary_data.pdf

Original source ↗

Results §1.2.1; Methods §§2.5.1–2.5.2; cached text lines 18–19, 70–74; matching task comparison table/ablation captions; Supplement: Nucleotide Transformer benchmark dataset description and Table 1

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: Retrieved 2026-09-16; sha256:a169bbebba2298e9c98c33c28053c1a0b42d0c3e88e5a8795934725c4408cc32
Retrieved: 2026-09-16T21:05:56.086158+00:00

source checked

automated source review · 2026-09-16

Audit details

Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.

Field: attributes.profile.diagram.caption

Source artifact SHA-256: a169bbebba2298e9c98c33c28053c1a0b42d0c3e88e5a8795934725c4408cc32

Hash scope: Hash scope not separately documented; inspect source record

Archive member: vbae117_supplementary_data.pdf

Inspected artifact

Diagram steps

["Input: Genomic sequence.","Evaluation: For the Nucleotide Transformer enhancer tasks, Supplementary Table 1 lists 14,968 training sequences and 400 test sequences of length 200. This set combines the original strong/weak/non-enhancer collection with 6,000 synthetic enhancers and 6,000 synthetic non-enhancers. Those counts do not describe the separate Genomic Benchmarks enhancer datasets.","Readout: Classification accuracy is described for the benchmark comparison."]

Individual claims
Understanding the natural language of DNA using encoder–decoder foundation models with byte-level precision

Original source ↗

Results §1.2.1; Methods §§2.5.1–2.5.2; cached text lines 18–19, 70–74; matching task comparison table/ablation captions; Supplement: Nucleotide Transformer benchmark dataset description and Table 1

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: version of record
Retrieved: 2026-09-16T10:33:35.378Z

source checked

automated source review · 2026-09-16

Audit details

Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.

Field: attributes.profile.diagram.steps

Source artifact SHA-256: e95d4be70d32e61af5a92eda8ea66f25a2cc629e3f83e7b5241b13cde8bdb83b

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Diagram steps

["Input: Genomic sequence.","Evaluation: For the Nucleotide Transformer enhancer tasks, Supplementary Table 1 lists 14,968 training sequences and 400 test sequences of length 200. This set combines the original strong/weak/non-enhancer collection with 6,000 synthetic enhancers and 6,000 synthetic non-enhancers. Those counts do not describe the separate Genomic Benchmarks enhancer datasets.","Readout: Classification accuracy is described for the benchmark comparison."]

Individual claims
enbed-2024__vbae117_supplementary_data.pdf

Original source ↗

Results §1.2.1; Methods §§2.5.1–2.5.2; cached text lines 18–19, 70–74; matching task comparison table/ablation captions; Supplement: Nucleotide Transformer benchmark dataset description and Table 1

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: Retrieved 2026-09-16; sha256:a169bbebba2298e9c98c33c28053c1a0b42d0c3e88e5a8795934725c4408cc32
Retrieved: 2026-09-16T21:05:56.086158+00:00

source checked

automated source review · 2026-09-16

Audit details

Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.

Field: attributes.profile.diagram.steps

Source artifact SHA-256: a169bbebba2298e9c98c33c28053c1a0b42d0c3e88e5a8795934725c4408cc32

Hash scope: Hash scope not separately documented; inspect source record

Archive member: vbae117_supplementary_data.pdf

Inspected artifact

Diagram title

Computational evaluation flow

Individual claims
Understanding the natural language of DNA using encoder–decoder foundation models with byte-level precision

Original source ↗

Results §1.2.1; Methods §§2.5.1–2.5.2; cached text lines 18–19, 70–74; matching task comparison table/ablation captions; Supplement: Nucleotide Transformer benchmark dataset description and Table 1

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: version of record
Retrieved: 2026-09-16T10:33:35.378Z

source checked

automated source review · 2026-09-16

Audit details

Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.

Field: attributes.profile.diagram.title

Source artifact SHA-256: e95d4be70d32e61af5a92eda8ea66f25a2cc629e3f83e7b5241b13cde8bdb83b

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Diagram title

Computational evaluation flow

Individual claims
enbed-2024__vbae117_supplementary_data.pdf

Original source ↗

Results §1.2.1; Methods §§2.5.1–2.5.2; cached text lines 18–19, 70–74; matching task comparison table/ablation captions; Supplement: Nucleotide Transformer benchmark dataset description and Table 1

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: Retrieved 2026-09-16; sha256:a169bbebba2298e9c98c33c28053c1a0b42d0c3e88e5a8795934725c4408cc32
Retrieved: 2026-09-16T21:05:56.086158+00:00

source checked

automated source review · 2026-09-16

Audit details

Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.

Field: attributes.profile.diagram.title

Source artifact SHA-256: a169bbebba2298e9c98c33c28053c1a0b42d0c3e88e5a8795934725c4408cc32

Hash scope: Hash scope not separately documented; inspect source record

Archive member: vbae117_supplementary_data.pdf

Inspected artifact

Datasets

Genomic Benchmarks includes human enhancer datasets; Nucleotide Transformer tasks include enhancer/non-enhancer and enhancer-strength labels.

Individual claims
Understanding the natural language of DNA using encoder–decoder foundation models with byte-level precision

Original source ↗

Results §1.2.1; Methods §§2.5.1–2.5.2; cached text lines 18–19, 70–74; matching task comparison table/ablation captions

Version: version of record
Retrieved: 2026-09-16T10:33:35.378Z

source checked

automated source review · 2026-09-16

Audit details

Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.

Field: attributes.profile.facts.0.value

Source artifact SHA-256: e95d4be70d32e61af5a92eda8ea66f25a2cc629e3f83e7b5241b13cde8bdb83b

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Splits

For the Nucleotide Transformer enhancer tasks, Supplementary Table 1 lists 14,968 training sequences and 400 test sequences of length 200. This set combines the original strong/weak/non-enhancer collection with 6,000 synthetic enhancers and 6,000 synthetic non-enhancers. Those counts do not describe the separate Genomic Benchmarks enhancer datasets.

Individual claims
Understanding the natural language of DNA using encoder–decoder foundation models with byte-level precision

Original source ↗

Supplement: Nucleotide Transformer benchmark dataset description and Table 1

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: version of record
Retrieved: 2026-09-16T10:33:35.378Z

source checked

automated source review · 2026-09-16

Audit details

Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.

Field: attributes.profile.facts.1.value

Source artifact SHA-256: e95d4be70d32e61af5a92eda8ea66f25a2cc629e3f83e7b5241b13cde8bdb83b

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Splits

For the Nucleotide Transformer enhancer tasks, Supplementary Table 1 lists 14,968 training sequences and 400 test sequences of length 200. This set combines the original strong/weak/non-enhancer collection with 6,000 synthetic enhancers and 6,000 synthetic non-enhancers. Those counts do not describe the separate Genomic Benchmarks enhancer datasets.

Individual claims
enbed-2024__vbae117_supplementary_data.pdf

Original source ↗

Supplement: Nucleotide Transformer benchmark dataset description and Table 1

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: Retrieved 2026-09-16; sha256:a169bbebba2298e9c98c33c28053c1a0b42d0c3e88e5a8795934725c4408cc32
Retrieved: 2026-09-16T21:05:56.086158+00:00

source checked

automated source review · 2026-09-16

Audit details

Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.

Field: attributes.profile.facts.1.value

Source artifact SHA-256: a169bbebba2298e9c98c33c28053c1a0b42d0c3e88e5a8795934725c4408cc32

Hash scope: Hash scope not separately documented; inspect source record

Archive member: vbae117_supplementary_data.pdf

Inspected artifact

Adaptation

Task-level enhancer classification; exact task configuration must be selected from the evaluated benchmark collection.

Individual claims
Understanding the natural language of DNA using encoder–decoder foundation models with byte-level precision

Original source ↗

Results §1.2.1; Methods §§2.5.1–2.5.2; cached text lines 18–19, 70–74; matching task comparison table/ablation captions

Version: version of record
Retrieved: 2026-09-16T10:33:35.378Z

source checked

automated source review · 2026-09-16

Audit details

Targeted full-paper and supplement review of the outstanding task fields, with original dataset metadata checked where accessible. Source-scoped omissions are explicit; no independent benchmark reproduction or numerical-result change.

Field: attributes.profile.facts.10.value

Source artifact SHA-256: e95d4be70d32e61af5a92eda8ea66f25a2cc629e3f83e7b5241b13cde8bdb83b

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Sources and history

Release 2026-09-17-d277315f7d76 · Record review: needs review

3 source records and release historyDownload this release
Technical metadata and extraction receipts

Stable ID: reported-task-132da895d4c381

areas
dna-genomes
tasks
Enhancer classification
entity level
task
version
Not reported
task
Enhancer classification
scope note
Paper-specific evaluation task; protocol completeness requires further extraction.
benchmark research
review date: 2026-09-17; status: primary_comparison_tables_located; primary sources: evidence-expansion-enbed-2024-e95d4be7; inspected locators: Table 1.; XML table vbae117-T1; Table 2.; XML table vbae117-T2; searched queries: Understanding the natural language of DNA using encoder–decoder foundation models with byte-level precision 10.1093/bioadv/vbae117; gaps: complete numerical transcription and independent cell review: Full primary artifact and table inventory preserved; no new numeric row is published from this audit alone.; exact checkpoint hashes and per-method scored denominators: Table labels alone do not establish these fields; do not infer checkpoint or scored count from model name or dataset size.; claim scope: Dated primary-source discovery and protocol/table screening. Source checking does not mean experimental reproduction. Only separately extracted and independently reviewed numeric batches are publishable.
historical missing metadata
protocol version: not_reported_in_legacy_extract; split: not_reported_in_legacy_extract
metadata review scope
historical_missing_metadata preserves the original discovery state. Current descriptive evidence and missingness are recorded in profile.facts; numerical-result review is separate.
legacy kinds
benchmark
entity classification
review date: 2026-09-17; rationale: This source-scoped record identifies the biological prediction task and holds its paper context. Preserve the existing task identity; exact split, model adaptation and scoring remain in linked evaluations or separate protocol records.; source ids: enbed-2024; source locator: Results §1.2.1; Methods §§2.5.1–2.5.2; cached text lines 18–19, 70–74; matching task comparison table/ablation captions; ambiguities: A paper- or suite-specific task may constrain some inputs or metrics; that alone does not make it interchangeable with a complete versioned protocol. No protocol equivalence is inferred.; Some legacy profile Entity type facts use the generic phrase computational evaluation protocol. That boilerplate is not sufficient to establish a single fixed protocol identity or to merge this task with another protocol record.
Related records

Suggest a correction