Strengths and considerations
No source-reviewed explanatory claims are recorded here yet.
Regulatory-sequence classification compares genomic models and tokenizers across established benchmark collections.
Explanatory profile: limited source coverage · Automated source review, 2026-09-16. Review applies to the cited claims; unresolved fields are listed below. Numerical results retain their own review status.
| Property | Description and evidence |
|---|---|
| Datasets | Nucleotide Transformer tasks, Genomic Benchmarks and GUE.SourcesThe impact of tokenizer selection in genomic language models · Methods §2.2 Benchmarks; Results §3.1; cached text lines 20–21, 43–44 |
| Splits | The experiment reuses three published fine-tuning benchmark collections and repeats every task at least ten times. The broad catalogue label spans multiple task partitions; no single regulatory-element split is defined by that label.SourcesThe impact of tokenizer selection in genomic language models · §2.2 Benchmarks |
| Metrics | Mean accuracy and Matthews correlation coefficient are reported per fine-tuning task; multiclass tasks use macro averaging. Results remain task-specific across the three benchmark collections.SourcesThe impact of tokenizer selection in genomic language models · §2.2–2.3 Benchmarks and Metrics |
| Baselines | Attention-based and state-space genomic language models with different tokenizer choices.SourcesThe impact of tokenizer selection in genomic language models · Methods §2.2 Benchmarks; Results §3.1; cached text lines 20–21, 43–44 |
| Leakage controls | The benchmark reuses GB, NTTv2 and GUE task datasets and compares models pretrained on different corpora, including the human reference genome. Its main Methods do not report a unified overlap-removal audit between all task examples and every model’s pretraining corpus. · Not reported in inspected sourcesSourcesThe impact of tokenizer selection in genomic language models · §2 Materials and methods, model/pretraining and task-dataset descriptions; Table 1 |
| Uncertainty | Each fine-tuning task is replicated at least ten times; hyperparameter search differs by model family.SourcesThe impact of tokenizer selection in genomic language models · Methods §2.2 Benchmarks; Results §3.1; cached text lines 20–21, 43–44 |
| Entity type | Paper-specific computational evaluation protocol.SourcesThe impact of tokenizer selection in genomic language models · Methods §2.2 Benchmarks; Results §3.1; cached text lines 20–21, 43–44 |
| Organisms | Dataset-specific organisms in Nucleotide Transformer, Genomic Benchmarks and GUE.SourcesThe impact of tokenizer selection in genomic language models · Methods §2.2 Benchmarks; Results §3.1; cached text lines 20–21, 43–44 |
| Assays | Regulatory and other genomic classification labels.SourcesThe impact of tokenizer selection in genomic language models · Methods §2.2 Benchmarks; Results §3.1; cached text lines 20–21, 43–44 |
| Allowed inputs | Tokenized DNA sequence.SourcesThe impact of tokenizer selection in genomic language models · Methods §2.2 Benchmarks; Results §3.1; cached text lines 20–21, 43–44 |
| Adaptation | Repeated task fine-tuning compares tokenizer/model configurations.SourcesThe impact of tokenizer selection in genomic language models · Methods §2.2 Benchmarks; Results §3.1; cached text lines 20–21, 43–44 |
Conceptual summary of the cited evaluation; exact task configuration and source version remain part of the protocol.
Nucleotide Transformer tasks, Genomic Benchmarks and GUE. Mean accuracy and Matthews correlation coefficient are reported per fine-tuning task; multiclass tasks use macro averaging. Results remain task-specific across the three benchmark collections. Attention-based and state-space genomic language models with different tokenizer choices.
Each evaluation records what was tested and under which conditions.
Release 2026-09-17-d277315f7d76 · 1 evaluation · 1 metric row. Different protocols are not a single leaderboard.
| Metric and finding | Coverage and uncertainty | Evidence |
|---|---|---|
| Caduceus (character tokens): regulatory sequence classification Configuration: Caduceus (character tokens)Task: regulatory sequence classificationDataset: genomic benchmark categories task-category MCC across benchmark datasets Independent external evaluation · Evaluation metadata: needs review | ||
| 0.778 MCC Unit: unitless · Direction: unknown | Uncertainty: not reported in legacy extract Scored: Not reported · Eligible: Not reported | source checkedThe impact of tokenizer selection in genomic language models · Table 2, Regulatory row, Caduceus (char) MCC column Source checking is not independent reproduction. |
Last literature check: 2026-09-17. Primary-paper discovery and source inspection. Source-checked results are not independently reproduced experiments.
| Paper or primary resource | Version | Reference |
|---|---|---|
| The impact of tokenizer selection in genomic language models | journal full text in PMC | Read source DOI: 10.1093/bioinformatics/btaf456 |
primary comparison table screened
No source-reviewed explanatory claims are recorded here yet.
Relevant full-paper computational evaluation sections, tables/captions and cited supplementary task passages were reviewed. Reporting omissions are scoped to the inspected sources. Original numerical results are unchanged.
Stable record: reported-task-cd127e56fb1f04Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.
One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.
17 evidence rows matching the loaded filters
| Property and statement | Original source and location | Review and provenance |
|---|---|---|
| Diagram caption Conceptual summary of the cited evaluation; exact task configuration and source version remain part of the protocol. Individual claims | The impact of tokenizer selection in genomic language models Methods §2.2 Benchmarks; Results §3.1; cached text lines 20–21, 43–44; Methods §2.2 Benchmarks; Results §3.1; cached text lines 20–21, 43–44; §2.2–2.3 Benchmarks and Metrics Version: journal full text in PMC | source checked automated source review · 2026-09-16 Audit detailsRelevant full-paper computational evaluation sections, tables/captions and cited supplementary task passages were reviewed. Reporting omissions are scoped to the inspected sources. Original numerical results are unchanged. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Diagram steps ["Allowed inputs: Tokenized DNA sequence.","Datasets: Nucleotide Transformer tasks, Genomic Benchmarks and GUE.","Metrics: Mean accuracy and Matthews correlation coefficient are reported per fine-tuning task; multiclass tasks use macro averaging. Results remain task-specific across the three benchmark collections."] Individual claims | The impact of tokenizer selection in genomic language models Methods §2.2 Benchmarks; Results §3.1; cached text lines 20–21, 43–44; Methods §2.2 Benchmarks; Results §3.1; cached text lines 20–21, 43–44; §2.2–2.3 Benchmarks and Metrics Version: journal full text in PMC | source checked automated source review · 2026-09-16 Audit detailsRelevant full-paper computational evaluation sections, tables/captions and cited supplementary task passages were reviewed. Reporting omissions are scoped to the inspected sources. Original numerical results are unchanged. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Diagram title Computational evaluation flow Individual claims | The impact of tokenizer selection in genomic language models Methods §2.2 Benchmarks; Results §3.1; cached text lines 20–21, 43–44; Methods §2.2 Benchmarks; Results §3.1; cached text lines 20–21, 43–44; §2.2–2.3 Benchmarks and Metrics Version: journal full text in PMC | source checked automated source review · 2026-09-16 Audit detailsRelevant full-paper computational evaluation sections, tables/captions and cited supplementary task passages were reviewed. Reporting omissions are scoped to the inspected sources. Original numerical results are unchanged. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Datasets Nucleotide Transformer tasks, Genomic Benchmarks and GUE. Individual claims | The impact of tokenizer selection in genomic language models Methods §2.2 Benchmarks; Results §3.1; cached text lines 20–21, 43–44 Version: journal full text in PMC | source checked automated source review · 2026-09-16 Audit detailsRelevant full-paper computational evaluation sections, tables/captions and cited supplementary task passages were reviewed. Reporting omissions are scoped to the inspected sources. Original numerical results are unchanged. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Splits The experiment reuses three published fine-tuning benchmark collections and repeats every task at least ten times. The broad catalogue label spans multiple task partitions; no single regulatory-element split is defined by that label. Individual claims | The impact of tokenizer selection in genomic language models §2.2 Benchmarks Version: journal full text in PMC | source checked automated source review · 2026-09-16 Audit detailsRelevant full-paper computational evaluation sections, tables/captions and cited supplementary task passages were reviewed. Reporting omissions are scoped to the inspected sources. Original numerical results are unchanged. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Adaptation Repeated task fine-tuning compares tokenizer/model configurations. Individual claims | The impact of tokenizer selection in genomic language models Methods §2.2 Benchmarks; Results §3.1; cached text lines 20–21, 43–44 Version: journal full text in PMC | source checked automated source review · 2026-09-16 Audit detailsRelevant full-paper computational evaluation sections, tables/captions and cited supplementary task passages were reviewed. Reporting omissions are scoped to the inspected sources. Original numerical results are unchanged. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Metrics Mean accuracy and Matthews correlation coefficient are reported per fine-tuning task; multiclass tasks use macro averaging. Results remain task-specific across the three benchmark collections. Individual claims | The impact of tokenizer selection in genomic language models §2.2–2.3 Benchmarks and Metrics Version: journal full text in PMC | source checked automated source review · 2026-09-16 Audit detailsRelevant full-paper computational evaluation sections, tables/captions and cited supplementary task passages were reviewed. Reporting omissions are scoped to the inspected sources. Original numerical results are unchanged. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Baselines Attention-based and state-space genomic language models with different tokenizer choices. Individual claims | The impact of tokenizer selection in genomic language models Methods §2.2 Benchmarks; Results §3.1; cached text lines 20–21, 43–44 Version: journal full text in PMC | source checked automated source review · 2026-09-16 Audit detailsRelevant full-paper computational evaluation sections, tables/captions and cited supplementary task passages were reviewed. Reporting omissions are scoped to the inspected sources. Original numerical results are unchanged. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Leakage controls The benchmark reuses GB, NTTv2 and GUE task datasets and compares models pretrained on different corpora, including the human reference genome. Its main Methods do not report a unified overlap-removal audit between all task examples and every model’s pretraining corpus. Individual claims | The impact of tokenizer selection in genomic language models §2 Materials and methods, model/pretraining and task-dataset descriptions; Table 1 Version: journal full text in PMC | unreported automated source review · 2026-09-16 Audit detailsRelevant full-paper computational evaluation sections, tables/captions and cited supplementary task passages were reviewed. Reporting omissions are scoped to the inspected sources. Original numerical results are unchanged. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Uncertainty Each fine-tuning task is replicated at least ten times; hyperparameter search differs by model family. Individual claims | The impact of tokenizer selection in genomic language models Methods §2.2 Benchmarks; Results §3.1; cached text lines 20–21, 43–44 Version: journal full text in PMC | source checked automated source review · 2026-09-16 Audit detailsRelevant full-paper computational evaluation sections, tables/captions and cited supplementary task passages were reviewed. Reporting omissions are scoped to the inspected sources. Original numerical results are unchanged. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
Release 2026-09-17-d277315f7d76 · Record review: needs review
Stable ID: reported-task-cd127e56fb1f04