rewire.it
benchmark · task

regulatory sequence classification

This paper-specific evaluation tests regulatory sequence classification using genomic benchmark categories.

1 evaluations · 1 metric rows

At a glance

Explanatory profile: limited source coverage · Automated source review, 2026-09-16. This does not change the review status of its results.

Data, procedure and scoring
PropertyDescription and evidence
Record typePaper-specific task; protocol incompletely extractedThe impact of tokenizer selection in genomic language models · Table 2, Regulatory row, Caduceus (char) MCC column
Inputsgenomic benchmark categoriesThe impact of tokenizer selection in genomic language models · Table 2, Regulatory row, Caduceus (char) MCC column
AssessmentMCCThe impact of tokenizer selection in genomic language models · Table 2, Regulatory row, Caduceus (char) MCC column
Recorded split or evaluation settingpaper benchmark summaryThe impact of tokenizer selection in genomic language models · Table 2, Regulatory row, Caduceus (char) MCC column
DatasetsNot extracted or verified for this record.
OrganismsNot extracted or verified for this record.
AssaysNot extracted or verified for this record.
AdaptationNot extracted or verified for this record.
BaselinesNot extracted or verified for this record.

How it works

Reported evaluation outline

Outline of the existing paper extraction. Split membership, fitting details and scorer implementation remain incompletely reviewed.

Reported evaluation outlinegenomic benchmark categories. Then: Recorded fitting or scoring procedure. Then: Assess MCCgenomic benchmark categoriesRecorded fitting or scoringprocedureAssess MCC
Read the diagram as text
  1. genomic benchmark categories
  2. Recorded fitting or scoring procedure
  3. Assess MCC
The impact of tokenizer selection in genomic language models · Table 2, Regulatory row, Caduceus (char) MCC column

Evaluation context

The existing paper extraction describes: task-category MCC across benchmark datasets. This description is retained with the exact evaluation records; it is not a new protocol reconstruction.

The impact of tokenizer selection in genomic language models · Table 2, Regulatory row, Caduceus (char) MCC column

Tested models and results

Release 2026-09-16-d74d282221a9 · 1 evaluation · 1 metric row. Different protocols are not a single leaderboard.

Results grouped by the exact reported evaluation
Metric and findingCoverage and uncertaintyEvidence
Caduceus (character tokens): regulatory sequence classification

task-category MCC across benchmark datasets

Independent external evaluation · Evaluation metadata: needs review

0.778 MCC

Unit: unitless · Direction: unknown

Uncertainty: not reported in legacy extract

Scored: Not reported · Eligible: Not reported

source checkedThe impact of tokenizer selection in genomic language models · Table 2, Regulatory row, Caduceus (char) MCC column

Source checking is not independent reproduction.

Strengths and limitations

Profile review details

Catalogue extraction inspected; protocol claims remain limited to the cited evidence. Missing details are not presumed absent from the original paper.

Stable record: reported-task-cd127e56fb1f04

Sources and history

Release 2026-09-16-d74d282221a9 · Record review: needs review

Download this release
Technical metadata and extraction receipts

Stable ID: reported-task-cd127e56fb1f04

areas
dna-genomes
tasks
regulatory sequence classification
entity level
task
version
Not reported
task
regulatory sequence classification
scope note
Paper-specific evaluation task; protocol completeness requires further extraction.
missing metadata
protocol version: not_reported_in_legacy_extract
Related records

Suggest a correction