rewire.itbenchmarks

← Back to use cases

DNA and genomes · Use case

Prioritise variants for splicing experiments

Which evaluated configurations can inform selection of human SNVs for follow-up splicing experiments?

Your decision and inputs

Inspect the matched MFASS configurations and their limitations before selecting a method and designing validation in the intended experimental setting.

Who this is for
Experimental researchers selecting variants for functional follow-up; Computational researchers comparing splice-effect configurations; Clinical researchers investigating the limits of assay evidence
Context
Research; Clinical research
Inputs
  • Human single nucleotide variants with alleles, genome assembly and gene/transcript context
  • The intended experimental endpoint and the number of variants that can be followed up
Expected output
A sourced set of exact evaluated configurations and remaining validation needs; no patient-level variant classification or recommendation.
Biological setting
MFASS measures exon recognition in an artificial minigene reporter. The matched study evaluates genomic-context SpliceAI and Pangolin configurations against this functional endpoint on a fixed held-out population.

Outside this use case

  • Indels and variant types outside the reviewed human SNV protocol
  • Patient-RNA effects, disease pathogenicity, diagnostic yield and treatment decisions
  • Pooling the matched annotation study with historical MFASS runs using other annotations or scoring populations

What this establishes for clinical research

Clinical applicability is not established. Reporter-assay ranking does not demonstrate patient-RNA performance, pathogenicity classification or clinical yield; validation in the intended population and workflow is still needed.

Which evaluations inform this question?

Evidence is grouped by its protocol. Relevance refers to the stated endpoint and context; it is separate from clinical validation and from the review method.

MFASS: matched GENCODE 44 canonical annotation

Current source-reviewed mapping

Proxy evidence: transfer to this question is limited

The recorded endpoint informs assay-oriented prioritisation under its declared conditions. Selecting variants for a different follow-up experiment requires transfer validation; the endpoint is not patient RNA or clinical pathogenicity.

Assessed endpoint
Ranking held-out MFASS SNVs by reporter-assay splice disruption under matched canonical annotation.
Evaluation protocol
MFASS: matched GENCODE 44 canonical annotation
Computational task
MFASS splice-variant prioritisation (Not yet reviewed)
Input and population constraints
  • Identical scored population: 8,297 of 8,324 held-out variants, 314 scored positives and 460 groups in all four configurations.
  • Shared GENCODE 44 canonical transcript selection and FASTA; 50-base distance; distinct SpliceAI and Pangolin masking settings.
  • Keep historical annotation conditions and scoring populations in their existing separate comparison groups.

Limits on interpretation

  • 23 assembly-orientation mismatches and four canonical-transcript-span exclusions remain unscored. The latter are protocol exclusions, not established faulty variants; full-population performance is unknown.
  • No top-100 precision difference is established and masked Pangolin is tie-sensitive. Paired contrast intervals are not individual-condition intervals.
  • Exploratory source review only; no independent human review, replication or clinical validation. Assembly-issue author confirmation is not established.

Automated source review · 2026-09-25 · Codex research curation

Reviewed exact configuration and population scope against pinned source bytes. Applicability remains proxy evidence; no human domain review or clinical validation.

Evaluated configurations

Each configuration below belongs to this protocol. Inspect its inputs, population and scoring conditions before comparing it with another evaluation.

SpliceAI 1.3.1 · mask 0 (S0)

Rewire evaluation · Source checked

8,297 of 8,324 held-out variants scored in every configuration (314 positives, 460 exon groups). The same 27 rows are excluded: 23 assembly-orientation mismatches and four outside the selected canonical transcript spans. Missing scores are not zero or negative predictions.

Population and split
8297/8324 · split-v2 test
Inputs and adaptation
GRCh38 genomic context; matched GENCODE44 canonical annotation · zero-shot pretrained specialists
Evaluation budget
one frozen execution per condition
Runtime and memory

Runtime and memory measurements are not reported in this evaluation. A study budget is not a runtime or memory measurement.

Recorded results for this configuration
MetricValueCoverageUncertainty and source
AUROC0.804

dimensionless · higher

8297/8324

Not reported

Result provenance
Average precision0.295

dimensionless · higher

8297/8324

Not reported

Result provenance
Precision at 1000.63

dimensionless · higher

8297/8324

Not reported

Result provenance
Recall at 1000.201

dimensionless · higher

8297/8324

Not reported

Result provenance

Uncertainty: Point estimate; paired contrast intervals are reported separately in the source and must not be used as intervals for this individual condition.

Evaluation methods, evidence and reproduction

No execution recipe has been verified for this exact configuration and evaluation. Inspect its methods and original run documentation before attempting reproduction.

SpliceAI 1.3.1 · mask 1 (S1)

Rewire evaluation · Source checked

8,297 of 8,324 held-out variants scored in every configuration (314 positives, 460 exon groups). The same 27 rows are excluded: 23 assembly-orientation mismatches and four outside the selected canonical transcript spans. Missing scores are not zero or negative predictions.

Population and split
8297/8324 · split-v2 test
Inputs and adaptation
GRCh38 genomic context; matched GENCODE44 canonical annotation · zero-shot pretrained specialists
Evaluation budget
one frozen execution per condition
Runtime and memory

Runtime and memory measurements are not reported in this evaluation. A study budget is not a runtime or memory measurement.

Recorded results for this configuration
MetricValueCoverageUncertainty and source
AUROC0.815

dimensionless · higher

8297/8324

Not reported

Result provenance
Average precision0.313

dimensionless · higher

8297/8324

Not reported

Result provenance
Precision at 1000.65

dimensionless · higher

8297/8324

Not reported

Result provenance
Recall at 1000.207

dimensionless · higher

8297/8324

Not reported

Result provenance

Uncertainty: Point estimate; paired contrast intervals are reported separately in the source and must not be used as intervals for this individual condition.

Evaluation methods, evidence and reproduction

No execution recipe has been verified for this exact configuration and evaluation. Inspect its methods and original run documentation before attempting reproduction.

Pangolin 1.0.2 + per-gene masking patch · mask False (P0)

Rewire evaluation · Source checked

8,297 of 8,324 held-out variants scored in every configuration (314 positives, 460 exon groups). The same 27 rows are excluded: 23 assembly-orientation mismatches and four outside the selected canonical transcript spans. Missing scores are not zero or negative predictions.

Population and split
8297/8324 · split-v2 test
Inputs and adaptation
GRCh38 genomic context; matched GENCODE44 canonical annotation · zero-shot pretrained specialists
Evaluation budget
one frozen execution per condition
Runtime and memory

Runtime and memory measurements are not reported in this evaluation. A study budget is not a runtime or memory measurement.

Recorded results for this configuration
MetricValueCoverageUncertainty and source
AUROC0.876

dimensionless · higher

8297/8324

Not reported

Result provenance
Average precision0.389

dimensionless · higher

8297/8324

Not reported

Result provenance
Precision at 1000.65

dimensionless · higher

8297/8324

Not reported

Result provenance
Recall at 1000.207

dimensionless · higher

8297/8324

Not reported

Result provenance

Uncertainty: Point estimate; paired contrast intervals are reported separately in the source and must not be used as intervals for this individual condition.

Evaluation methods, evidence and reproduction

No execution recipe has been verified for this exact configuration and evaluation. Inspect its methods and original run documentation before attempting reproduction.

Pangolin 1.0.2 + per-gene masking patch · mask True (P1)

Rewire evaluation · Source checked

8,297 of 8,324 held-out variants scored in every configuration (314 positives, 460 exon groups). The same 27 rows are excluded: 23 assembly-orientation mismatches and four outside the selected canonical transcript spans. Missing scores are not zero or negative predictions.

Population and split
8297/8324 · split-v2 test
Inputs and adaptation
GRCh38 genomic context; matched GENCODE44 canonical annotation · zero-shot pretrained specialists
Evaluation budget
one frozen execution per condition
Runtime and memory

Runtime and memory measurements are not reported in this evaluation. A study budget is not a runtime or memory measurement.

Recorded results for this configuration
MetricValueCoverageUncertainty and source
AUROC0.873

dimensionless · higher

8297/8324

Not reported

Result provenance
Average precision0.411

dimensionless · higher

8297/8324

Not reported

Result provenance
Precision at 1000.66

dimensionless · higher

8297/8324

Not reported

Result provenance
Recall at 1000.21

dimensionless · higher

8297/8324

Not reported

Result provenance

Uncertainty: Point estimate; paired contrast intervals are reported separately in the source and must not be used as intervals for this individual condition.

Evaluation methods, evidence and reproduction

No execution recipe has been verified for this exact configuration and evaluation. Inspect its methods and original run documentation before attempting reproduction.

Open the protocol's results and comparison checks →

Original execution documentation ↗

Mapping sources and review metadata

Mapping use-case-mapping-splicing-mfass-matched-v1 · revision 1

Initial bounded applicability review of the matched MFASS study.

Reviewed evidence fingerprint c4cfbe66c30d59308fc8d2226f42720a09a6e54d2f407c7141b08e01e39c0a47

What evidence is still missing?

  • All four conditions scored 8,297 of 8,324 held-out variants. The same 27 exclusions comprise 23 hg19-to-hg38 assembly-orientation mismatches and four canonical-transcript-span exclusions; the latter are not established faulty variants. Missing predictions are not negative predictions.
  • No top-100 precision difference is established; Pangolin's masked top-100 result is sensitive to the registered tie order. Precision at 100 does not transfer automatically to another follow-up capacity or prevalence.
  • Individual-condition uncertainty intervals are not recorded. Paired-contrast intervals concern differences between conditions and must not be shown as each condition's uncertainty.
  • The study is exploratory: prior outcomes were inspected and nine contrast intervals are unadjusted. Matching annotation does not isolate model architecture.
  • Author confirmation of the assembly-orientation finding is not established. Human scientific review and independent replication remain outstanding.
  • These mappings do not change the MFASS task's discovered status. A source-reviewed applicability mapping is separate from reviewing the task record.

Contribute evidence or propose a correction

Sources and review

Automated source review · 2026-09-25 · Codex research curation

Bounded review of pinned reports, protocol documentation and intake narrative. No new model execution, human domain review, independent replication or clinical validation.

Release provenance and downloads

Release 2026-09-25-8af07e960e5f

Use-case input digest d0c76e33f58fe8d4845cdba114d6bb6bea1920fc68c5a5e43922801c7e7ca27b

Download the mappings and review metadata (JSON) · Verify release checksums

Question use-case-splicing-follow-up. Numerical values above come from this release's existing evaluation records.