rewire.it
Configuration

Caduceus (character tokens)

This Caduceus configuration is a character-token baseline in a study of genomic tokenisation.

SourcesThe impact of tokenizer selection in genomic language models · 1 Introduction (paragraph 9); 4 Discussion (paragraph 4)

1 evaluation · 1 metric row

How it worksEvaluated procedure (conceptual)
Evaluated procedure (conceptual)1. DNA sequences represented by individual nucleotide characters. Then: 2. Caduceus (character tokens). Then: 3. Task-specific genomic classification predictionsEvaluated procedure (conceptual)1. DNA sequences represented by individual nucleotide characters. Then: 2. Caduceus (character tokens). Then: 3. Task-specific genomic classification predictionsEvaluated procedure (conceptual)1. DNA sequences represented by individual nucleotide characters. Then: 2. Caduceus (character tokens). Then: 3. Task-specific genomic classification predictions

Conceptual input–method–output guide. Check the procedure text and linked evaluation for fitted components, additional inputs and exact settings.

SourcesThe impact of tokenizer selection in genomic language models · 1 Introduction (paragraph 9); Abstract (paragraph 1)

At a glance

limited source coverage · Automated source review, 2026-09-16. All specifications and missing details

Evaluations and results

Release 2026-09-17-d277315f7d76 · 1 evaluation · 1 metric row. Different protocols are not a single leaderboard.

Results grouped by the exact reported evaluation
Metric and findingCoverage and uncertaintyEvidence
Caduceus (character tokens): regulatory sequence classification

task-category MCC across benchmark datasets

Independent external evaluation · Evaluation metadata: needs review

0.778 MCC

Unit: unitless · Direction: unknown

Uncertainty: not reported in legacy extract

Scored: Not reported · Eligible: Not reported

source checkedThe impact of tokenizer selection in genomic language models · Table 2, Regulatory row, Caduceus (char) MCC column

Source checking is not independent reproduction.

How it works

How the evaluated method works

A state-space genomic sequence model processes individual nucleotide tokens. The study compares character, k-mer and byte-pair schemes across downstream classifiers.

SourcesThe impact of tokenizer selection in genomic language models · 1 Introduction (paragraph 9); Abstract (paragraph 1)
Underlying method and version boundaries

Caduceus exposes distinct Ph and PS configurations. The documented Ph-131k checkpoint uses 16 layers, width 256 and reverse-complement data augmentation; PS implements reverse-complement equivariance. These training choices are not interchangeable.

Sourceskuleshov-group/caduceus README.md · README.md; introduction, model description, pretrained-model and usage sections at pinned revision
What was evaluated

The linked evaluation record identifies Caduceus (character tokens): regulatory sequence classification. Its dataset, split, adaptation and evidence origin remain attached to the reported results.

SourcesThe impact of tokenizer selection in genomic language models · The named evaluation’s methods and comparison table; exact preserved evaluation IDs: evaluation-b2-genomic-tokenizer-selection-2025

Strengths and limitations

Limitations and conditions

  • Tokenisation benefits are task-dependent; the paper reports different behaviour for SARS-CoV-2 classification and does not establish a universal best tokenizer.
    SourcesThe impact of tokenizer selection in genomic language models · 4 Discussion (paragraph 5); 3 Results/3.1 Tokenization differentially impacts performance of gLMs on specific benchmark tasks (paragraph 1)
Profile review details

Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.

Stable record: reported-model-d5bc536ca6f0d3

Specifications

Inputs, training, access and other details

Explanatory profile: limited source coverage · Automated source review, 2026-09-16. Review applies to the cited claims; unresolved fields are listed below. Numerical results retain their own review status.

Inputs, outputs and configuration
PropertyDescription and evidence
Model typeDNA state-space model; this record is the paper-specific evaluated configuration.
Sourceskuleshov-group/caduceus README.md · README.md model description
Architecture / procedureA state-space genomic sequence model processes individual nucleotide tokens. The study compares character, k-mer and byte-pair schemes across downstream classifiers.
SourcesThe impact of tokenizer selection in genomic language models · 1 Introduction (paragraph 9); Abstract (paragraph 1)
Biological inputsDNA sequences represented by individual nucleotide characters
SourcesThe impact of tokenizer selection in genomic language models · 1 Introduction (paragraph 6); 1 Introduction (paragraph 3)
OutputsTask-specific genomic classification predictions
SourcesThe impact of tokenizer selection in genomic language models · 2 Materials and methods/2.6 Genomic tasks (paragraph 4); 2 Materials and methods/2.6 Genomic tasks (paragraph 6)
Parameters3.9-million-parameter variant as reported for this row
SourcesThe impact of tokenizer selection in genomic language models · 2 Materials and methods/2.3 Metrics (paragraph 1); 2 Materials and methods/2.6 Genomic tasks (paragraph 6)
Known versions / configuration3.9M parameter variant
SourcesThe impact of tokenizer selection in genomic language models · 2 Materials and methods/2.6 Genomic tasks (paragraph 6); 5 Conclusions (paragraph 1)
Training data / fittingThe tokenisation study’s model pretraining and 44 classification fine-tuning tasks; each task retains its own adaptation and split.
SourcesThe impact of tokenizer selection in genomic language models · 2 Materials and methods/2.3 Metrics (paragraph 1); 4 Discussion (paragraph 5)
Context limitsThe paper uses a 4,000-nucleotide context for the compared genomic tokenisation models.
SourcesThe impact of tokenizer selection in genomic language models · 1 Introduction (paragraph 5); 1 Introduction (paragraph 9)
AccessOfficial upstream implementation and usage documentation: https://github.com/kuleshov-group/caduceus/blob/0060a6d8079b6a040fc55d505e15972a327b70a6/README.md. This pinned documentation revision is not automatically the evaluated weight revision.
Sourceskuleshov-group/caduceus README.md · README.md; installation, model download and usage instructions
Code licenceApache 2.0 (upstream repository code at the cited revision; this does not establish every dependency or historical checkpoint licence).
Sourceskuleshov-group/caduceus LICENSE · LICENSE; complete licence text
Weights licenceThe inspected model-access documentation does not explicitly identify terms for this exact evaluated checkpoint or fitted head; repository code terms are shown separately. · Not reported in inspected sources
Sourceskuleshov-group/caduceus README.md · README.md; checkpoint/access documentation and licence scope

Evidence table

Inspect claims, sources and review details

Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.

One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.

20 evidence rows matching the loaded filters

Claims, original sources and review scope · Release 2026-09-17-d277315f7d76
Property and statementOriginal source and locationReview and provenance
Diagram caption

Conceptual input–method–output guide. Check the procedure text and linked evaluation for fitted components, additional inputs and exact settings.

Individual claims
The impact of tokenizer selection in genomic language models

Original source ↗

1 Introduction (paragraph 9); Abstract (paragraph 1)

Version: journal full text in PMC
Retrieved: 2026-09-16T10:38:57.558210+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.

Field: attributes.profile.diagram.caption

Source artifact SHA-256: 0a01c36fdd63f3f6db509777e61c3f87e8a298c810f8aef7974915aaa0655342

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Diagram steps

["DNA sequences represented by individual nucleotide characters","Caduceus (character tokens)","Task-specific genomic classification predictions"]

Individual claims
The impact of tokenizer selection in genomic language models

Original source ↗

1 Introduction (paragraph 9); Abstract (paragraph 1)

Version: journal full text in PMC
Retrieved: 2026-09-16T10:38:57.558210+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.

Field: attributes.profile.diagram.steps

Source artifact SHA-256: 0a01c36fdd63f3f6db509777e61c3f87e8a298c810f8aef7974915aaa0655342

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Diagram title

Evaluated procedure (conceptual)

Individual claims
The impact of tokenizer selection in genomic language models

Original source ↗

1 Introduction (paragraph 9); Abstract (paragraph 1)

Version: journal full text in PMC
Retrieved: 2026-09-16T10:38:57.558210+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.

Field: attributes.profile.diagram.title

Source artifact SHA-256: 0a01c36fdd63f3f6db509777e61c3f87e8a298c810f8aef7974915aaa0655342

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Model type

DNA state-space model; this record is the paper-specific evaluated configuration.

Individual claims
kuleshov-group/caduceus README.md

Original source ↗

README.md model description

Version: 0060a6d8079b6a040fc55d505e15972a327b70a6
Retrieved: 2026-09-16T20:00:02.715892+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.

Field: attributes.profile.facts.0.value

Source artifact SHA-256: e508e5199d0cfb9c36dbd503cdccf50d734f419fa8cf1f890dec0db7e741ad42

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Architecture / procedure

A state-space genomic sequence model processes individual nucleotide tokens. The study compares character, k-mer and byte-pair schemes across downstream classifiers.

Individual claims
The impact of tokenizer selection in genomic language models

Original source ↗

1 Introduction (paragraph 9); Abstract (paragraph 1)

Version: journal full text in PMC
Retrieved: 2026-09-16T10:38:57.558210+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.

Field: attributes.profile.facts.1.value

Source artifact SHA-256: 0a01c36fdd63f3f6db509777e61c3f87e8a298c810f8aef7974915aaa0655342

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Weights licence

The inspected model-access documentation does not explicitly identify terms for this exact evaluated checkpoint or fitted head; repository code terms are shown separately.

Individual claims
kuleshov-group/caduceus README.md

Original source ↗

README.md; checkpoint/access documentation and licence scope

Version: 0060a6d8079b6a040fc55d505e15972a327b70a6
Retrieved: 2026-09-16T20:00:02.715892+00:00

unreported

automated source review · 2026-09-16

Audit details

Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.

Field: attributes.profile.facts.10.value

Source artifact SHA-256: e508e5199d0cfb9c36dbd503cdccf50d734f419fa8cf1f890dec0db7e741ad42

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Biological inputs

DNA sequences represented by individual nucleotide characters

Individual claims
The impact of tokenizer selection in genomic language models

Original source ↗

1 Introduction (paragraph 6); 1 Introduction (paragraph 3)

Version: journal full text in PMC
Retrieved: 2026-09-16T10:38:57.558210+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.

Field: attributes.profile.facts.2.value

Source artifact SHA-256: 0a01c36fdd63f3f6db509777e61c3f87e8a298c810f8aef7974915aaa0655342

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Outputs

Task-specific genomic classification predictions

Individual claims
The impact of tokenizer selection in genomic language models

Original source ↗

2 Materials and methods/2.6 Genomic tasks (paragraph 4); 2 Materials and methods/2.6 Genomic tasks (paragraph 6)

Version: journal full text in PMC
Retrieved: 2026-09-16T10:38:57.558210+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.

Field: attributes.profile.facts.3.value

Source artifact SHA-256: 0a01c36fdd63f3f6db509777e61c3f87e8a298c810f8aef7974915aaa0655342

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Parameters

3.9-million-parameter variant as reported for this row

Individual claims
The impact of tokenizer selection in genomic language models

Original source ↗

2 Materials and methods/2.3 Metrics (paragraph 1); 2 Materials and methods/2.6 Genomic tasks (paragraph 6)

Version: journal full text in PMC
Retrieved: 2026-09-16T10:38:57.558210+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.

Field: attributes.profile.facts.4.value

Source artifact SHA-256: 0a01c36fdd63f3f6db509777e61c3f87e8a298c810f8aef7974915aaa0655342

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Known versions / configuration

3.9M parameter variant

Individual claims
The impact of tokenizer selection in genomic language models

Original source ↗

2 Materials and methods/2.6 Genomic tasks (paragraph 6); 5 Conclusions (paragraph 1)

Version: journal full text in PMC
Retrieved: 2026-09-16T10:38:57.558210+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.

Field: attributes.profile.facts.5.value

Source artifact SHA-256: 0a01c36fdd63f3f6db509777e61c3f87e8a298c810f8aef7974915aaa0655342

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Sources and history

Release 2026-09-17-d277315f7d76 · Record review: needs review

3 source records and release historyDownload this release
Technical metadata and extraction receipts

Stable ID: reported-model-d5bc536ca6f0d3

areas
dna-genomes
entity level
method
version
3.9M parameter variant
reported name
Caduceus (character tokens)
historical missing metadata
checkpoint revision: not_reported_in_legacy_extract; training data: not_reported_in_legacy_extract; licence: not_reported_in_legacy_extract
metadata review scope
historical_missing_metadata preserves the original discovery state. Current descriptive evidence and missingness are recorded in profile.facts; numerical-result review is separate.
legacy kinds
model
entity classification
review date: 2026-09-17; rationale: This source-scoped entry preserves the method/configuration actually named in an evaluation. It is neither a global family identity nor proof of an immutable checkpoint; the linked evaluation retains adaptation, fitting and scoring details.; source ids: genomic-tokenizer-selection-2025; evidence-reported-base-caduceus-readme-md; source locator: 1 Introduction (paragraph 9); Abstract (paragraph 1) | README.md model description | 1 Introduction (paragraph 9); 4 Discussion (paragraph 4); ambiguities: Configuration means the source-labelled evaluated identity. It does not establish missing checkpoint hashes, default settings or equivalence to same-named records in other papers.
Related records

Suggest a correction