Model type
DNA state-space model; this record is the paper-specific evaluated configuration.
This Caduceus configuration is a character-token baseline in a study of genomic tokenisation.
Conceptual input–method–output guide. Check the procedure text and linked evaluation for fitted components, additional inputs and exact settings.
DNA state-space model; this record is the paper-specific evaluated configuration.
DNA sequences represented by individual nucleotide characters
Task-specific genomic classification predictions
limited source coverage · Automated source review, 2026-09-16. All specifications and missing details
Release 2026-09-17-d277315f7d76 · 1 evaluation · 1 metric row. Different protocols are not a single leaderboard.
| Metric and finding | Coverage and uncertainty | Evidence |
|---|---|---|
| Caduceus (character tokens): regulatory sequence classification Configuration: Caduceus (character tokens)Task: regulatory sequence classificationDataset: genomic benchmark categories task-category MCC across benchmark datasets Independent external evaluation · Evaluation metadata: needs review | ||
| 0.778 MCC Unit: unitless · Direction: unknown | Uncertainty: not reported in legacy extract Scored: Not reported · Eligible: Not reported | source checkedThe impact of tokenizer selection in genomic language models · Table 2, Regulatory row, Caduceus (char) MCC column Source checking is not independent reproduction. |
A state-space genomic sequence model processes individual nucleotide tokens. The study compares character, k-mer and byte-pair schemes across downstream classifiers.
Caduceus exposes distinct Ph and PS configurations. The documented Ph-131k checkpoint uses 16 layers, width 256 and reverse-complement data augmentation; PS implements reverse-complement equivariance. These training choices are not interchangeable.
The linked evaluation record identifies Caduceus (character tokens): regulatory sequence classification. Its dataset, split, adaptation and evidence origin remain attached to the reported results.
Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.
Stable record: reported-model-d5bc536ca6f0d3Explanatory profile: limited source coverage · Automated source review, 2026-09-16. Review applies to the cited claims; unresolved fields are listed below. Numerical results retain their own review status.
| Property | Description and evidence |
|---|---|
| Model type | DNA state-space model; this record is the paper-specific evaluated configuration.Sourceskuleshov-group/caduceus README.md · README.md model description |
| Architecture / procedure | A state-space genomic sequence model processes individual nucleotide tokens. The study compares character, k-mer and byte-pair schemes across downstream classifiers.SourcesThe impact of tokenizer selection in genomic language models · 1 Introduction (paragraph 9); Abstract (paragraph 1) |
| Biological inputs | DNA sequences represented by individual nucleotide charactersSourcesThe impact of tokenizer selection in genomic language models · 1 Introduction (paragraph 6); 1 Introduction (paragraph 3) |
| Outputs | Task-specific genomic classification predictionsSourcesThe impact of tokenizer selection in genomic language models · 2 Materials and methods/2.6 Genomic tasks (paragraph 4); 2 Materials and methods/2.6 Genomic tasks (paragraph 6) |
| Parameters | 3.9-million-parameter variant as reported for this rowSourcesThe impact of tokenizer selection in genomic language models · 2 Materials and methods/2.3 Metrics (paragraph 1); 2 Materials and methods/2.6 Genomic tasks (paragraph 6) |
| Known versions / configuration | 3.9M parameter variantSourcesThe impact of tokenizer selection in genomic language models · 2 Materials and methods/2.6 Genomic tasks (paragraph 6); 5 Conclusions (paragraph 1) |
| Training data / fitting | The tokenisation study’s model pretraining and 44 classification fine-tuning tasks; each task retains its own adaptation and split.SourcesThe impact of tokenizer selection in genomic language models · 2 Materials and methods/2.3 Metrics (paragraph 1); 4 Discussion (paragraph 5) |
| Context limits | The paper uses a 4,000-nucleotide context for the compared genomic tokenisation models.SourcesThe impact of tokenizer selection in genomic language models · 1 Introduction (paragraph 5); 1 Introduction (paragraph 9) |
| Access | Official upstream implementation and usage documentation: https://github.com/kuleshov-group/caduceus/blob/0060a6d8079b6a040fc55d505e15972a327b70a6/README.md. This pinned documentation revision is not automatically the evaluated weight revision.Sourceskuleshov-group/caduceus README.md · README.md; installation, model download and usage instructions |
| Code licence | Apache 2.0 (upstream repository code at the cited revision; this does not establish every dependency or historical checkpoint licence).Sourceskuleshov-group/caduceus LICENSE · LICENSE; complete licence text |
| Weights licence | The inspected model-access documentation does not explicitly identify terms for this exact evaluated checkpoint or fitted head; repository code terms are shown separately. · Not reported in inspected sourcesSourceskuleshov-group/caduceus README.md · README.md; checkpoint/access documentation and licence scope |
Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.
One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.
20 evidence rows matching the loaded filters
| Property and statement | Original source and location | Review and provenance |
|---|---|---|
| Diagram caption Conceptual input–method–output guide. Check the procedure text and linked evaluation for fitted components, additional inputs and exact settings. Individual claims | The impact of tokenizer selection in genomic language models 1 Introduction (paragraph 9); Abstract (paragraph 1) Version: journal full text in PMC | source checked automated source review · 2026-09-16 Audit detailsPrimary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Diagram steps ["DNA sequences represented by individual nucleotide characters","Caduceus (character tokens)","Task-specific genomic classification predictions"] Individual claims | The impact of tokenizer selection in genomic language models 1 Introduction (paragraph 9); Abstract (paragraph 1) Version: journal full text in PMC | source checked automated source review · 2026-09-16 Audit detailsPrimary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Diagram title Evaluated procedure (conceptual) Individual claims | The impact of tokenizer selection in genomic language models 1 Introduction (paragraph 9); Abstract (paragraph 1) Version: journal full text in PMC | source checked automated source review · 2026-09-16 Audit detailsPrimary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Model type DNA state-space model; this record is the paper-specific evaluated configuration. Individual claims | kuleshov-group/caduceus README.md README.md model description Version: 0060a6d8079b6a040fc55d505e15972a327b70a6 | source checked automated source review · 2026-09-16 Audit detailsPrimary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Architecture / procedure A state-space genomic sequence model processes individual nucleotide tokens. The study compares character, k-mer and byte-pair schemes across downstream classifiers. Individual claims | The impact of tokenizer selection in genomic language models 1 Introduction (paragraph 9); Abstract (paragraph 1) Version: journal full text in PMC | source checked automated source review · 2026-09-16 Audit detailsPrimary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Weights licence The inspected model-access documentation does not explicitly identify terms for this exact evaluated checkpoint or fitted head; repository code terms are shown separately. Individual claims | kuleshov-group/caduceus README.md README.md; checkpoint/access documentation and licence scope Version: 0060a6d8079b6a040fc55d505e15972a327b70a6 | unreported automated source review · 2026-09-16 Audit detailsPrimary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Biological inputs DNA sequences represented by individual nucleotide characters Individual claims | The impact of tokenizer selection in genomic language models 1 Introduction (paragraph 6); 1 Introduction (paragraph 3) Version: journal full text in PMC | source checked automated source review · 2026-09-16 Audit detailsPrimary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Outputs Task-specific genomic classification predictions Individual claims | The impact of tokenizer selection in genomic language models 2 Materials and methods/2.6 Genomic tasks (paragraph 4); 2 Materials and methods/2.6 Genomic tasks (paragraph 6) Version: journal full text in PMC | source checked automated source review · 2026-09-16 Audit detailsPrimary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Parameters 3.9-million-parameter variant as reported for this row Individual claims | The impact of tokenizer selection in genomic language models 2 Materials and methods/2.3 Metrics (paragraph 1); 2 Materials and methods/2.6 Genomic tasks (paragraph 6) Version: journal full text in PMC | source checked automated source review · 2026-09-16 Audit detailsPrimary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Known versions / configuration 3.9M parameter variant Individual claims | The impact of tokenizer selection in genomic language models 2 Materials and methods/2.6 Genomic tasks (paragraph 6); 5 Conclusions (paragraph 1) Version: journal full text in PMC | source checked automated source review · 2026-09-16 Audit detailsPrimary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
Release 2026-09-17-d277315f7d76 · Record review: needs review
Stable ID: reported-model-d5bc536ca6f0d3