Model type
Transformer representation pipeline; this record is the paper-specific evaluated configuration.
ENBED is a byte-level encoder–decoder transformer for genomic sequence representation and sequence-to-sequence tasks.
Conceptual input–method–output guide. Check the procedure text and linked evaluation for fitted components, additional inputs and exact settings.
Transformer representation pipeline; this record is the paper-specific evaluated configuration.
DNA sequences at single-byte nucleotide resolution
Task-specific classifications or generated DNA sequences
limited source coverage · Automated source review, 2026-09-16. All specifications and missing details
Release 2026-09-17-d277315f7d76 · 1 evaluation · 1 metric row. Different protocols are not a single leaderboard.
| Metric and finding | Coverage and uncertainty | Evidence |
|---|---|---|
| ENBED (GRCh38): Enhancer classification Configuration: ENBED (GRCh38)Task: Enhancer classificationDataset: Genomic Benchmarks Mouse Enhancers ENBED trained on GRCh38; reported Genomic Benchmarks classification accuracy. Author-reported evaluation · Evaluation metadata: needs review | ||
| 81.1 Accuracy Unit: % · Direction: unknown | Uncertainty: not reported in legacy extract Scored: Not reported · Eligible: Not reported | source checkedUnderstanding the natural language of DNA using encoder–decoder foundation models with byte-level precision · Table 2, Mouse Enhancers row, ENBED (GRCh38) column Source checking is not independent reproduction. |
Byte-level nucleotide inputs feed encoder and decoder transformer blocks with a subquadratic attention implementation. Masked-language pretraining precedes task-specific adaptation.
The linked evaluation record identifies ENBED (GRCh38): Enhancer classification. Its dataset, split, adaptation and evidence origin remain attached to the reported results.
Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.
Stable record: reported-model-bdb1db16d3389dExplanatory profile: limited source coverage · Automated source review, 2026-09-16. Review applies to the cited claims; unresolved fields are listed below. Numerical results retain their own review status.
| Property | Description and evidence |
|---|---|
| Model type | Transformer representation pipeline; this record is the paper-specific evaluated configuration.SourcesUnderstanding the natural language of DNA using encoder–decoder foundation models with byte-level precision · 1 Introduction/1.2 Our contributions (paragraph 1); 4 Ablation studies/4.1 Encoder–decoder architecture (paragraph 2) |
| Architecture / procedure | Byte-level nucleotide inputs feed encoder and decoder transformer blocks with a subquadratic attention implementation. Masked-language pretraining precedes task-specific adaptation.SourcesUnderstanding the natural language of DNA using encoder–decoder foundation models with byte-level precision · 1 Introduction/1.2 Our contributions (paragraph 1); 4 Ablation studies/4.1 Encoder–decoder architecture (paragraph 2) |
| Biological inputs | DNA sequences at single-byte nucleotide resolutionSourcesUnderstanding the natural language of DNA using encoder–decoder foundation models with byte-level precision · 1 Introduction/1.2 Our contributions (paragraph 2); 5 Discussion (paragraph 1) |
| Outputs | Task-specific classifications or generated DNA sequencesSourcesUnderstanding the natural language of DNA using encoder–decoder foundation models with byte-level precision · 2 Methods/2.4 Applications of foundation models using transfer learning/2.4.2 Fine-tuning for downstream tasks (paragraph 1); 2 Methods/2.5 Application domains/2.5.1 Genomic benchmarks (paragraph 1) |
| Parameters | 1.2 billion trainable parameters in the full encoder–decoder model.SourcesUnderstanding the natural language of DNA using encoder–decoder foundation models with byte-level precision · 4 Ablation studies (paragraph 1); 4 Ablation studies/4.1 Encoder–decoder architecture (paragraph 1) |
| Known versions / configuration | ENBED (GRCh38) is the comparison-table label; that label does not specify an immutable weight revision. · Not reported in inspected sourcesSourcesUnderstanding the natural language of DNA using encoder–decoder foundation models with byte-level precision · Model identification in the comparison table and corresponding Methods; immutable checkpoint revision is not supplied by the table label. |
| Training data / fitting | Reference-genome sequences; the GRCh38-labelled row is a distinct human-reference configuration.SourcesUnderstanding the natural language of DNA using encoder–decoder foundation models with byte-level precision · 1 Introduction/1.2 Our contributions/1.2.1 Evaluation of performance on genomic benchmark datasets (paragraph 1); 3 Results/3.1 ENBED outperforms state-of-the-art models on GB datasets (paragraph 2) |
| Context limits | 16,384 input/output tokens using local sliding-window plus global attention. The 512-token value in Methods describes the dense-attention hardware baseline, not ENBED’s final context.SourcesUnderstanding the natural language of DNA using encoder–decoder foundation models with byte-level precision · 2 Methods/2.3 Attention (paragraph 2); 2 Methods/2.3 Attention/2.3.1 Sliding-window attention (paragraph 1) |
| Access | The authors provide implementation code at https://github.itap.purdue.edu/Clan-labs/ENBED and model weights through https://huggingface.co/malusare. A table-specific checkpoint hash is not supplied by these account-level links.SourcesUnderstanding the natural language of DNA using encoder–decoder foundation models with byte-level precision · Data availability |
| Code licence | No explicit code licence was established from the paper’s availability statement and inspected repository-root documentation. · Not reported in inspected sourcesSourcesUnderstanding the natural language of DNA using encoder–decoder foundation models with byte-level precision · Abstract (paragraph 1); 2 Methods/2.4 Applications of foundation models using transfer learning/2.4.1 Building the foundation model (paragraph 1) |
| Weights licence | The inspected model-access documentation does not explicitly identify terms for this exact evaluated checkpoint or fitted head; repository code terms are shown separately. · Not reported in inspected sourcesSourcesUnderstanding the natural language of DNA using encoder–decoder foundation models with byte-level precision · Abstract (paragraph 1); 2 Methods/2.1 Encoder–decoder model architecture (paragraph 1) |
Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.
One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.
19 evidence rows matching the loaded filters
| Property and statement | Original source and location | Review and provenance |
|---|---|---|
| Diagram caption Conceptual input–method–output guide. Check the procedure text and linked evaluation for fitted components, additional inputs and exact settings. Individual claims | Understanding the natural language of DNA using encoder–decoder foundation models with byte-level precision 1 Introduction/1.2 Our contributions (paragraph 1); 4 Ablation studies/4.1 Encoder–decoder architecture (paragraph 2) Version: version of record | source checked automated source review · 2026-09-16 Audit detailsPrimary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Diagram steps ["DNA sequences at single-byte nucleotide resolution","ENBED (GRCh38)","Task-specific classifications or generated DNA sequences"] Individual claims | Understanding the natural language of DNA using encoder–decoder foundation models with byte-level precision 1 Introduction/1.2 Our contributions (paragraph 1); 4 Ablation studies/4.1 Encoder–decoder architecture (paragraph 2) Version: version of record | source checked automated source review · 2026-09-16 Audit detailsPrimary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Diagram title Evaluated procedure (conceptual) Individual claims | Understanding the natural language of DNA using encoder–decoder foundation models with byte-level precision 1 Introduction/1.2 Our contributions (paragraph 1); 4 Ablation studies/4.1 Encoder–decoder architecture (paragraph 2) Version: version of record | source checked automated source review · 2026-09-16 Audit detailsPrimary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Model type Transformer representation pipeline; this record is the paper-specific evaluated configuration. Individual claims | Understanding the natural language of DNA using encoder–decoder foundation models with byte-level precision 1 Introduction/1.2 Our contributions (paragraph 1); 4 Ablation studies/4.1 Encoder–decoder architecture (paragraph 2) Version: version of record | source checked automated source review · 2026-09-16 Audit detailsPrimary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Architecture / procedure Byte-level nucleotide inputs feed encoder and decoder transformer blocks with a subquadratic attention implementation. Masked-language pretraining precedes task-specific adaptation. Individual claims | Understanding the natural language of DNA using encoder–decoder foundation models with byte-level precision 1 Introduction/1.2 Our contributions (paragraph 1); 4 Ablation studies/4.1 Encoder–decoder architecture (paragraph 2) Version: version of record | source checked automated source review · 2026-09-16 Audit detailsPrimary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Weights licence The inspected model-access documentation does not explicitly identify terms for this exact evaluated checkpoint or fitted head; repository code terms are shown separately. Individual claims | Understanding the natural language of DNA using encoder–decoder foundation models with byte-level precision Abstract (paragraph 1); 2 Methods/2.1 Encoder–decoder model architecture (paragraph 1) Version: version of record | unreported automated source review · 2026-09-16 Audit detailsPrimary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Biological inputs DNA sequences at single-byte nucleotide resolution Individual claims | Understanding the natural language of DNA using encoder–decoder foundation models with byte-level precision 1 Introduction/1.2 Our contributions (paragraph 2); 5 Discussion (paragraph 1) Version: version of record | source checked automated source review · 2026-09-16 Audit detailsPrimary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Outputs Task-specific classifications or generated DNA sequences Individual claims | Understanding the natural language of DNA using encoder–decoder foundation models with byte-level precision 2 Methods/2.4 Applications of foundation models using transfer learning/2.4.2 Fine-tuning for downstream tasks (paragraph 1); 2 Methods/2.5 Application domains/2.5.1 Genomic benchmarks (paragraph 1) Version: version of record | source checked automated source review · 2026-09-16 Audit detailsPrimary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Parameters 1.2 billion trainable parameters in the full encoder–decoder model. Individual claims | Understanding the natural language of DNA using encoder–decoder foundation models with byte-level precision 4 Ablation studies (paragraph 1); 4 Ablation studies/4.1 Encoder–decoder architecture (paragraph 1) Version: version of record | source checked automated source review · 2026-09-16 Audit detailsPrimary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Known versions / configuration ENBED (GRCh38) is the comparison-table label; that label does not specify an immutable weight revision. Individual claims | Understanding the natural language of DNA using encoder–decoder foundation models with byte-level precision Model identification in the comparison table and corresponding Methods; immutable checkpoint revision is not supplied by the table label. Version: version of record | unreported automated source review · 2026-09-16 Audit detailsPrimary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
Release 2026-09-17-d277315f7d76 · Record review: needs review
Stable ID: reported-model-bdb1db16d3389d