rewire.itbenchmarks

← Back to use cases

RNA and transcriptomes · Use case

Set baselines for UTR translation experiments

Before training a more complex model of reporter translation, which sequence-only controls should be measured on the same held-out data?

Your decision and inputs

Establish training-mean and sequence-composition controls on a fixed held-out reporter dataset before deciding whether a more complex model adds useful prediction. The current evidence contains these two procedural controls and no pretrained model.

Who this is for
Computational researchers building models of reporter translation; Experimental researchers evaluating a designed 5′UTR library before selecting a modelling strategy
Context
Research
Inputs
  • Complete processed sequences and measured mean ribosome load from a defined reporter library
  • A fixed training, validation and test assignment, with test labels withheld from fitting and model selection
Expected output
Two exact baseline configurations, their matched held-out evaluations and a repeat recipe; no foundation-model ranking or prediction of therapeutic performance.
Biological setting
The current comparison covers only the Sample designed subset in mRNABench and its measured mean ribosome load. Both controls use the full processed source sequence and the same seed-2541 split: 70,011 training, 15,003 validation and 15,003 test records. Both score all 15,003 held-out test records.

Outside this use case

  • Benchmark-wide mRNABench performance or comparisons with the paper’s aggregated model scores
  • Claims about novel combinations of sequence motifs, homology-separated generalisation, other reporter systems or RNA chemistries
  • Therapeutic potency, in-vivo protein output and clinical decisions

What this establishes for clinical research

This is research evidence for designing a baseline comparison. Reporter mean ribosome load does not establish therapeutic efficacy or clinical suitability.

Which evaluations inform this question?

Evidence is grouped by its protocol. Relevance refers to the stated endpoint and context; it is separate from clinical validation and from the review method.

mRNABench Sample designed MRL

Current source-reviewed mapping

Direct evidence for the stated endpoint

The matched sequence-composition and training-mean controls provide an auditable starting point for deciding whether a more complex method adds predictive value on this defined endpoint. They do not establish transfer to a different library or supply a pretrained-model comparison.

Assessed endpoint
Predicting target_mrl_designed on the complete 15,003-record canonical test split of the Sample designed reporter dataset.
Evaluation protocol
mRNABench Sample designed MRL
Computational task
A reviewed task relationship is not recorded for this protocol.
Input and population constraints
  • Both controls use the same full processed sequences, seed-2541 split, training labels and complete held-out test population. No validation or test labels select hyperparameters.
  • Composition features are log1p length, A/C/G/T fractions and unknown fraction, fitted with train-only RidgeCV over the recorded five alpha values and no normalization. The constant control is the training-target arithmetic mean.
  • MSE is the matched primary endpoint. Constant-control Pearson and Spearman are unavailable; do not replace missing values with zeros.

Limits on interpretation

  • One target, one split and one execution per method; no uncertainty or seed variability. No homology-separated or compositional-generalisation claim.
  • This training-only protocol differs from upstream probing and is not an aggregate mRNABench score or a reproduction of a published model result.
  • The original run’s raw inputs and saved predictions are not publicly archived. The recipe obtains source data and runs the controls again, including separate protein controls; it does not rescore prior predictions. Source-data reuse terms and cross-platform reproducibility remain unresolved.
  • Research baseline evidence only, with no new execution, human domain review or clinical validation.

Automated source review · 2026-09-28 · Codex research curation

Primary-source curation and separate automated cross-review checked exact evidence identities, comparator coverage, endpoint relevance and transfer limits. No new execution, human scientific review, independent replication or clinical validation.

Evaluated configurations

Each configuration below belongs to this protocol. Inspect its inputs, population and scoring conditions before comparing it with another evaluation.

Sequence composition + RidgeCV (mRNABench Sample designed MRL)

Rewire evaluation · Source checked

Predict target_mrl_designed from sequence on the complete 15,003-record canonical test split. Train-only RidgeCV differs from upstream default validation evaluation; this is not an aggregate mRNABench score.

Population and split
15003/15003 · canonical test
Inputs and adaptation
Complete source sequence only; no assay labels during prediction · train-only fitting
Evaluation budget
one local CPU evaluation; no hyperparameter search outside training
Runtime and memory

Inference and fit section: 0.955 s.

Metric evaluation section: 1.18 s.

Device: Not reported. Batch size: 32.

These timings describe the recorded sections of this run, not total runtime or a general hardware benchmark. Peak memory is not reported.

Recorded results for this configuration
MetricValueCoverageUncertainty and source
mse1.91

squared mean ribosome load · lower

15003/15003

Not reported

Result provenance
pearson0.437

dimensionless · higher

15003/15003

Not reported

Result provenance
spearman0.495

dimensionless · higher

15003/15003

Not reported

Result provenance

Uncertainty: Not estimated; one complete selected evaluation.

Evaluation methods, evidence and reproduction

Recipe: generate and evaluate predictions

This committed script regenerates the selected local evaluation with the recorded inputs and configuration. Sequence-control instructions execute all four controls; select this evaluation from their outputs. Raw prior predictions are private, so public evidence alone cannot rescore them.

Training-mean control (mRNABench Sample designed MRL)

Rewire evaluation · Source checked

Predict target_mrl_designed from sequence on the complete 15,003-record canonical test split. Train-only RidgeCV differs from upstream default validation evaluation; this is not an aggregate mRNABench score.

Population and split
15003/15003 · canonical test
Inputs and adaptation
Complete source sequence only; no assay labels during prediction · train-only fitting
Evaluation budget
one local CPU evaluation; no hyperparameter search outside training
Runtime and memory

Inference and fit section: 0.0676 s.

Metric evaluation section: 1.18 s.

Device: Not reported. Batch size: 32.

These timings describe the recorded sections of this run, not total runtime or a general hardware benchmark. Peak memory is not reported.

Recorded results for this configuration
MetricValueCoverageUncertainty and source
mse2.37

squared mean ribosome load · lower

15003/15003

Not reported

Result provenance
pearsonundefined

dimensionless · higher

15003/15003

Not reported

Result provenance
spearmanundefined

dimensionless · higher

15003/15003

Not reported

Result provenance

Uncertainty: Not estimated; one complete selected evaluation.

Evaluation methods, evidence and reproduction

Recipe: generate and evaluate predictions

This committed script regenerates the selected local evaluation with the recorded inputs and configuration. Sequence-control instructions execute all four controls; select this evaluation from their outputs. Raw prior predictions are private, so public evidence alone cannot rescore them.

Open the protocol's results and comparison checks →

Original execution documentation ↗

Mapping sources and review metadata

Mapping use-case-mapping-utr-translation-baselines-designed-v1 · revision 1

Initial applicability review: primary sources and separate automated cross-review support this exact protocol, its complete selected comparator group and the stated limits.

Reviewed evidence fingerprint 1034171098b5ead5f29764df1c6efaccee4aa4c24f235945e36b8536ec30caf5

What evidence is still missing?

  • Only two procedural controls have been evaluated here; no pretrained model or more complex sequence model has a matched result in this protocol.
  • There is one split and one execution per configuration, with no uncertainty interval or seed-variability estimate. The split does not establish homology separation or compositional generalisation.
  • RidgeCV uses training labels only and leaves the validation split unused. This differs from upstream probing; the selected test MSE must not be compared as if it were the paper’s aggregated Pearson score or default validation result.
  • The training-mean control has undefined Pearson and Spearman correlations because its predictions are constant. Unavailable correlations are not zero performance scores.
  • The original run’s raw inputs and saved predictions are not publicly archived. The public reports record hashes and a recipe for obtaining source data and running the controls again; prior predictions cannot be rescored from those reports. Dataset reuse terms are unreported.
  • The recorded repeat recipe executes all four local sequence controls, including a separate protein dataset. Portable command examples have not been rerun verbatim during this review. Timings cover the reported calculation sections, not complete setup or cross-machine performance.

Contribute evidence or propose a correction

Sources and review

Automated source review · 2026-09-28 · Codex research curation

Primary-source curation and separate automated cross-review checked exact evidence identities, comparator coverage, endpoint relevance and transfer limits. No new execution, human scientific review, independent replication or clinical validation.

Release provenance and downloads

Release 2026-09-28-c7b5ac6d34f2

Use-case input digest a0dd27a5f430ec387d309fd6e8615083250873ce4cff5e882c6fb8458ef80e95

Download questions, collection plans and review metadata (JSON) · Verify release checksums

Question use-case-utr-translation-baselines. Any numerical results on this page come from this release's existing evaluation records.