rewire.itbenchmarks

← Back to use cases

Proteins and complexes · Use case

Assess methods for protein stability experiments

What evidence supports ranking protein substitutions by folding stability, and what must be validated before choosing a method?

Your decision and inputs

Inspect the available AMFR assay results as a narrow example, then identify the missing singles-only comparison and validation required for the target protein and endpoint.

Who this is for
Experimental researchers selecting variants for protein stability experiments; Computational researchers evaluating sequence-based variant rankings; Translational researchers checking whether stability evidence answers a disease question
Context
Research; Clinical research
Inputs
  • Wild-type protein sequence and amino acid substitutions
  • A defined construct and experimental stability endpoint; alignment-derived methods additionally require a traceable alignment and model
Expected output
Exact existing AMFR configurations, their assay scope and unresolved comparisons; no universal model ranking or pathogenicity prediction.
Biological setting
The current evidence is limited to the 47-residue AMFR_HUMAN_Tsuboyama_2023_4G3O construct in ProteinGym v1.3. Its cDNA-display proteolysis assay infers folding stability. Completed evaluations cover 2,972 variants: 820 single and 2,152 double substitutions.

Outside this use case

  • Whole-protein function, cellular activity, organismal fitness and clinical pathogenicity
  • Generalisation to other proteins, other ESM-2 checkpoints or the full ProteinGym track
  • Treating the existing mixed cohort as a completed matched single-substitution comparison

What this establishes for clinical research

Clinical applicability is not established. Stability of this short experimental construct is not evidence of clinical pathogenicity or suitability for diagnosis or treatment.

Which evaluations inform this question?

Evidence is grouped by its protocol. Relevance refers to the stated endpoint and context; it is separate from clinical validation and from the review method.

ProteinGym v1.3 AMFR substitution assay

Current source-reviewed mapping

Proxy evidence: transfer to this question is limited

This completed short-construct assay is a narrow example for assessing stability-ranking evidence. Its mixed single/double cohort does not directly answer the planned single-substitution comparison or establish transfer to another protein.

Assessed endpoint
Ranking the mixed AMFR assay cohort by proteolysis-inferred folding stability using ESM-2 8M masked marginals.
Evaluation protocol
ProteinGym v1.3 AMFR substitution assay
Computational task
A reviewed task relationship is not recorded for this protocol.
Input and population constraints
  • One 47-residue AMFR construct and all 2,972 variants, comprising 820 singles and 2,152 doubles; no train/test split.
  • Exact esm2_t6_8M_UR50D checkpoint, wild-type-context masked marginals summed over substituted sites, CPU, one thread; no assay labels, alignment or structure as scorer inputs.
  • Protocol-only mapping: a direct reviewed task-membership relationship is not recorded.

Limits on interpretation

  • No interval or seed-variability estimate; training overlap remains unknown and assay bytes have not been independently authenticated against an upstream checksum.
  • Recorded inference timing excludes loading, preparation and metrics; peak memory is unreported.
  • The site-independent and full EVmutation methods in the planned singles study have no completed comparison here.
  • No general whole-protein, ProteinGym-wide or clinical inference; automated source review is not human scientific validation.

Automated source review · 2026-09-25 · Codex research curation

Reviewed existing measured evidence separately from the planned singles comparison. No new experiment, human domain review or clinical validation.

Evaluated configurations

Each configuration below belongs to this protocol. Inspect its inputs, population and scoring conditions before comparing it with another evaluation.

ESM-2 8M masked-marginal scoring (ProteinGym v1.3 AMFR substitution assay)

Rewire evaluation · Source checked

Score all 2,972 variants in AMFR_HUMAN_Tsuboyama_2023_4G3O with ESM-2 8M masked marginals. This single-assay result is not a full ProteinGym track or suite score.

Population and split
2972/2972 · Full assay; no train/test split
Inputs and adaptation
Wild-type protein sequence and amino-acid substitutions; no MSA, structure or assay labels · zero-shot masked marginals
Evaluation budget
one local CPU evaluation; no hyperparameter search outside training
Runtime and memory

Inference and fit section: 0.188 s.

Metric evaluation section: 0.0343 s.

Device: cpu. Batch size: 32.

These timings describe the recorded sections of this run, not total runtime or a general hardware benchmark. Peak memory is not reported.

Recorded results for this configuration
MetricValueCoverageUncertainty and source
AUC0.394

dimensionless · higher

2972/2972

Not reported

Result provenance
MCC-0.139

dimensionless · higher

2972/2972

Not reported

Result provenance
NDCG0.44

dimensionless · higher

2972/2972

Not reported

Result provenance
Spearman-0.209

dimensionless · higher

2972/2972

Not reported

Result provenance
Top_recall0.057

dimensionless · higher

2972/2972

Not reported

Result provenance

Uncertainty: Not estimated; one complete selected evaluation.

Evaluation methods, evidence and reproduction

Recipe: generate and evaluate predictions

This committed script regenerates the selected AMFR assay using the recorded checkpoint and masked-marginal scorer. It is not the complete ProteinGym track or a reproduction of a paper score.

Open the protocol's results and comparison checks →

Original execution documentation ↗

Mapping sources and review metadata

Mapping use-case-mapping-protein-stability-amfr-esm2 · revision 1

Initial bounded applicability review of the existing AMFR ESM-2 evaluation.

Reviewed evidence fingerprint 81883d91b243e8a9c958bca976ad951aa82ba5a130eff9263dc8f3a04707012e

ProteinGym v1.3 AMFR seeded-random control

Current source-reviewed mapping

Proxy evidence: transfer to this question is limited

A recorded null control helps inspect the assay and evaluation procedure. It is not a biological prediction method recommendation, a chance-performance interval or a matched comparison with the separately executed ESM-2 protocol.

Assessed endpoint
One fixed-seed random ranking of the mixed AMFR stability cohort.
Evaluation protocol
ProteinGym v1.3 AMFR seeded-random control
Computational task
A reviewed task relationship is not recorded for this protocol.
Input and population constraints
  • All 2,972 AMFR variants, mixing single and double substitutions; a single fixed seed of 0.
  • SHA256 ranking of prepared variant IDs; no biological prediction or label fitting.
  • Protocol-only mapping; keep its evaluation and configuration separate from ESM-2 and the planned singles comparison.

Limits on interpretation

  • One seed does not estimate chance variability or uncertainty. Do not subtract these separate protocol results to assert an evaluated winner.
  • Local input hashes do not independently authenticate assay bytes against an upstream published checksum.
  • This control supplies no pathogenicity or clinical suitability evidence; no human scientific review or independent replication is recorded.

Automated source review · 2026-09-25 · Codex research curation

Reviewed as a separate one-seed control, not clinical evidence or a cross-protocol comparison.

Evaluated configurations

Each configuration below belongs to this protocol. Inspect its inputs, population and scoring conditions before comparing it with another evaluation.

Seeded random ranking control (ProteinGym v1.3 AMFR substitution assay)

Rewire evaluation · Source checked

One seed-0 random ranking on all 2,972 AMFR variants. This is one complete assay, not the full ProteinGym track or an estimate of a chance-performance interval.

Population and split
2972/2972 · Full assay; no train/test split
Inputs and adaptation
protocol allowlisted biological inputs; see baseline configuration for extra inputs · none; seed fixed before scoring
Evaluation budget
one local CPU evaluation; no held-out hyperparameter selection
Runtime and memory

Inference and fit section: 0.00539 s.

Metric evaluation section: 0.0566 s.

Device: Not reported. Batch size: 32.

These timings describe the recorded sections of this run, not total runtime or a general hardware benchmark. Peak memory is not reported.

Recorded results for this configuration
MetricValueCoverageUncertainty and source
AUC0.514

dimensionless · higher

2972/2972

Not reported

Result provenance
MCC0.019

dimensionless · higher

2972/2972

Not reported

Result provenance
NDCG0.526

dimensionless · higher

2972/2972

Not reported

Result provenance
Spearman0.008

dimensionless · higher

2972/2972

Not reported

Result provenance
Top_recall0.087

dimensionless · higher

2972/2972

Not reported

Result provenance

Uncertainty: Not estimated; one selected evaluation.

Evaluation methods, evidence and reproduction

No execution recipe has been verified for this exact configuration and evaluation. Inspect its methods and original run documentation before attempting reproduction.

Open the protocol's results and comparison checks →

Original execution documentation ↗

Mapping sources and review metadata

Mapping use-case-mapping-protein-stability-amfr-random · revision 1

Initial review of the separate fixed-seed random control as limited context.

Reviewed evidence fingerprint 2fa413430c25c90a4184be20e096c2200185e566d1690953e29bc40d55c98ddf

What evidence is still missing?

  • The existing ESM-2 and fixed-seed random results use separate protocols. They are displayed separately and do not establish a matched cross-protocol winner.
  • No uncertainty intervals or seed-variability estimates are recorded for these completed evaluations. One random ranking is not a chance-performance interval.
  • The proposed ESM-2 versus EVCouplings site-independent and EVmutation comparison targets a frozen set of covered single substitutions. It has no completed results.
  • The planned comparison still needs a separate execution decision, a traceable real evolutionary model, the full-model adapter, resource logging, frozen populations and tested analysis code. Planning resource ceilings are not measured requirements.
  • The recorded ESM-2 timer excludes checkpoint loading, preparation and metrics; peak memory is unreported. It is not an end-to-end or cross-model speed comparison.
  • The existing AMFR protocols have no reviewed direct task-membership link. The mappings therefore reference the protocols only and do not infer membership from the ProteinGym suite.
  • Assay bytes were hashed locally without independent authentication against an upstream published checksum. Training overlap, independent reproduction and human scientific review remain unresolved.

Planned work

These plans do not contribute measured results or evaluated winners above.

Contribute evidence or propose a correction

Sources and review

Automated source review · 2026-09-25 · Codex research curation

Bounded review of pinned assay metadata, completed execution records and the planning artifact. The plan is not a result. No new model execution, human domain review, independent replication or clinical validation. Public planned-work navigation points to the exact reviewed source copy; the original private-repository link remains in source provenance.

Release provenance and downloads

Release 2026-09-25-8af07e960e5f

Use-case input digest d0c76e33f58fe8d4845cdba114d6bb6bea1920fc68c5a5e43922801c7e7ca27b

Download the mappings and review metadata (JSON) · Verify release checksums

Question use-case-protein-stability. Numerical values above come from this release's existing evaluation records.