PFMBench STABILITY: TAPE_Stability
TAPE_Stability. Scored with Spearman on TAPE_Stability. Fine-tuned with an adapter under the PFMBench harness; train, validation and test counts are in Table 1.
Overview
TAPE_Stability. Scored with Spearman on TAPE_Stability. Fine-tuned with an adapter under the PFMBench harness; train, validation and test counts are in Table 1.
Consult the linked sources for architecture or protocol details. Missing evidence is not evidence of a missing capability.
Evaluation design
Benchmarks bring together tasks and protocols. A task describes the biological question; a protocol defines a particular test.
Benchmarks
These source-backed links do not make different protocols or scores interchangeable.
Recorded evaluations
Each evaluation records what was tested and under which conditions.
- DPLM on PFMBench STABILITY: TAPE_Stability
- ESM-2 on PFMBench STABILITY: TAPE_Stability
- ESM-C on PFMBench STABILITY: TAPE_Stability
- ESM3 on PFMBench STABILITY: TAPE_Stability
- PGLM on PFMBench STABILITY: TAPE_Stability
- ProstT5 on PFMBench STABILITY: TAPE_Stability
- ProtGPT2 on PFMBench STABILITY: TAPE_Stability
- ProTrek on PFMBench STABILITY: TAPE_Stability
- ProtST on PFMBench STABILITY: TAPE_Stability
- ProtT5 on PFMBench STABILITY: TAPE_Stability
- SaProt on PFMBench STABILITY: TAPE_Stability
- VenusPLM on PFMBench STABILITY: TAPE_Stability
Run instructions
No runnable recipe has been reviewed for this task. Dataset access, model requirements, licences and compute requirements must be checked against its sources before execution.
A task describes a biological question. Choose a linked protocol to obtain concrete split and scoring instructions.
Published comparisons
Explore the results reported under one evaluation protocol. Each figure keeps its source, dataset and metric together; it is not a ranking across studies. The pooled view gathers every source table that reports the same metric and names what it does not hold constant.
PFMBench STABILITY: TAPE_Stability
spearman (correlation) · Higher values are better for this metric.
Every method PFMBench reports on TAPE_Stability, scored with Spearman on TAPE_Stability.
- ESM-2 · Configuration · Author-reported evaluation0.32112
- VenusPLM · Configuration · Author-reported evaluation0.33907
- ESM-C · Configuration · Author-reported evaluation0.29976
- ProtGPT2 · Configuration · Author-reported evaluation0.14803
- PGLM · Configuration · Author-reported evaluation0 .33127
- ProtT5 · Configuration · Author-reported evaluation0.18638
- DPLM · Configuration · Author-reported evaluation0.29440
- SaProt · Configuration · Author-reported evaluation0.24804
- ProstT5 · Configuration · Author-reported evaluation0.13032
- ProtST · Configuration · Author-reported evaluation0.06623
- ESM3 · Configuration · Author-reported evaluation0.15650
- ProTrek · Configuration · Author-reported evaluation0.04924
Source order is preserved. Plotted marks show point estimates; uncertainty, where reported, is retained in the printed values and table. Differences do not establish statistical significance.
pfmbench primary benchmark evidence · Table 1, row(TAPE_Stability)Values, uncertainty and evidence
| Tested entity | Printed value | Uncertainty | Evidence |
|---|---|---|---|
| ESM-2 · Configuration | 0.32112 correlation | Not reported | Author-reported evaluation · source checkedpfmbench primary benchmark evidence · Table 3, row(ESM-2 [ 35 ]), column(Stability) |
| VenusPLM · Configuration | 0.33907 correlation | Not reported | Author-reported evaluation · source checkedpfmbench primary benchmark evidence · Table 3, row(VenusPLM [ 59 ]), column(Stability) |
| ESM-C · Configuration | 0.29976 correlation | Not reported | Author-reported evaluation · source checkedpfmbench primary benchmark evidence · Table 3, row(ESM-C), column(Stability) |
| ProtGPT2 · Configuration | 0.14803 correlation | Not reported | Author-reported evaluation · source checkedpfmbench primary benchmark evidence · Table 3, row(ProtGPT2 [ 12 ]), column(Stability) |
| PGLM · Configuration | 0 .33127 correlation | Not reported | Author-reported evaluation · source checkedpfmbench primary benchmark evidence · Table 3, row(PGLM [ 5 ]), column(Stability) |
| ProtT5 · Configuration | 0.18638 correlation | Not reported | Author-reported evaluation · source checkedpfmbench primary benchmark evidence · Table 3, row(ProtT5 [ 11 ]), column(Stability) |
| DPLM · Configuration | 0.29440 correlation | Not reported | Author-reported evaluation · source checkedpfmbench primary benchmark evidence · Table 3, row(DPLM [ 68 ]), column(Stability) |
| SaProt · Configuration | 0.24804 correlation | Not reported | Author-reported evaluation · source checkedpfmbench primary benchmark evidence · Table 3, row(SaProt [ 55 ]), column(Stability) |
| ProstT5 · Configuration | 0.13032 correlation | Not reported | Author-reported evaluation · source checkedpfmbench primary benchmark evidence · Table 3, row(ProstT5 [ 21 ]), column(Stability) |
| ProtST · Configuration | 0.06623 correlation | Not reported | Author-reported evaluation · source checkedpfmbench primary benchmark evidence · Table 3, row(ProtST [ 71 ]), column(Stability) |
| ESM3 · Configuration | 0.15650 correlation | Not reported | Author-reported evaluation · source checkedpfmbench primary benchmark evidence · Table 3, row(ESM3 [ 18 ]), column(Stability) |
| ProTrek · Configuration | 0.04924 correlation | Not reported | Author-reported evaluation · source checkedpfmbench primary benchmark evidence · Table 3, row(ProTrek [ 56 ]), column(Stability) |
Scope and limitations
- Author-reported numbers, source checked but not independently reproduced.
- The metric differs by task, taken from Table 1, so these figures cannot be averaged into one score.
- Table 3 scores come from adapter fine-tuning; the ProteinGym figure is zero-shot and is not comparable to them.
Source transcription and grouping reviewed by automated source review on 2026-09-18. These experiments were not independently reproduced by rewire.
Tested entities and results
Release 2026-09-17-134cd1815de8 · 12 evaluations · 12 metric rows. Different protocols are not a single leaderboard. Where several source tables report the same metric, the published comparisons above offer a pooled view that names what it does not hold constant.
| Metric and finding | Coverage and uncertainty | Evidence |
|---|---|---|
| DPLM on PFMBench STABILITY: TAPE_Stability Configuration: DPLMTask: PFMBench STABILITY: TAPE_StabilityDataset subset: TAPE_Stability (PFMBench split) Fine-tuned with an adapter under the PFMBench harness; train, validation and test counts are in Table 1. Author-reported evaluation · Evaluation metadata: source checked | ||
| 0.29440 spearman Unit: correlation · Direction: higher | Uncertainty: Not reported Scored: Not reported · Eligible: Not reported | source checkedpfmbench primary benchmark evidence · Table 3, row(DPLM [ 68 ]), column(Stability) Source checking is not independent reproduction. |
| ESM-2 on PFMBench STABILITY: TAPE_Stability Configuration: ESM-2Task: PFMBench STABILITY: TAPE_StabilityDataset subset: TAPE_Stability (PFMBench split) Fine-tuned with an adapter under the PFMBench harness; train, validation and test counts are in Table 1. Author-reported evaluation · Evaluation metadata: source checked | ||
| 0.32112 spearman Unit: correlation · Direction: higher | Uncertainty: Not reported Scored: Not reported · Eligible: Not reported | source checkedpfmbench primary benchmark evidence · Table 3, row(ESM-2 [ 35 ]), column(Stability) Source checking is not independent reproduction. |
| ESM-C on PFMBench STABILITY: TAPE_Stability Configuration: ESM-CTask: PFMBench STABILITY: TAPE_StabilityDataset subset: TAPE_Stability (PFMBench split) Fine-tuned with an adapter under the PFMBench harness; train, validation and test counts are in Table 1. Author-reported evaluation · Evaluation metadata: source checked | ||
| 0.29976 spearman Unit: correlation · Direction: higher | Uncertainty: Not reported Scored: Not reported · Eligible: Not reported | source checkedpfmbench primary benchmark evidence · Table 3, row(ESM-C), column(Stability) Source checking is not independent reproduction. |
| ESM3 on PFMBench STABILITY: TAPE_Stability Configuration: ESM3Task: PFMBench STABILITY: TAPE_StabilityDataset subset: TAPE_Stability (PFMBench split) Fine-tuned with an adapter under the PFMBench harness; train, validation and test counts are in Table 1. Author-reported evaluation · Evaluation metadata: source checked | ||
| 0.15650 spearman Unit: correlation · Direction: higher | Uncertainty: Not reported Scored: Not reported · Eligible: Not reported | source checkedpfmbench primary benchmark evidence · Table 3, row(ESM3 [ 18 ]), column(Stability) Source checking is not independent reproduction. |
| PGLM on PFMBench STABILITY: TAPE_Stability Configuration: PGLMTask: PFMBench STABILITY: TAPE_StabilityDataset subset: TAPE_Stability (PFMBench split) Fine-tuned with an adapter under the PFMBench harness; train, validation and test counts are in Table 1. Author-reported evaluation · Evaluation metadata: source checked | ||
| 0 .33127 spearman Unit: correlation · Direction: higher | Uncertainty: Not reported Scored: Not reported · Eligible: Not reported | source checkedpfmbench primary benchmark evidence · Table 3, row(PGLM [ 5 ]), column(Stability) Source checking is not independent reproduction. |
| ProstT5 on PFMBench STABILITY: TAPE_Stability Configuration: ProstT5Task: PFMBench STABILITY: TAPE_StabilityDataset subset: TAPE_Stability (PFMBench split) Fine-tuned with an adapter under the PFMBench harness; train, validation and test counts are in Table 1. Author-reported evaluation · Evaluation metadata: source checked | ||
| 0.13032 spearman Unit: correlation · Direction: higher | Uncertainty: Not reported Scored: Not reported · Eligible: Not reported | source checkedpfmbench primary benchmark evidence · Table 3, row(ProstT5 [ 21 ]), column(Stability) Source checking is not independent reproduction. |
| ProtGPT2 on PFMBench STABILITY: TAPE_Stability Configuration: ProtGPT2Task: PFMBench STABILITY: TAPE_StabilityDataset subset: TAPE_Stability (PFMBench split) Fine-tuned with an adapter under the PFMBench harness; train, validation and test counts are in Table 1. Author-reported evaluation · Evaluation metadata: source checked | ||
| 0.14803 spearman Unit: correlation · Direction: higher | Uncertainty: Not reported Scored: Not reported · Eligible: Not reported | source checkedpfmbench primary benchmark evidence · Table 3, row(ProtGPT2 [ 12 ]), column(Stability) Source checking is not independent reproduction. |
| ProTrek on PFMBench STABILITY: TAPE_Stability Configuration: ProTrekTask: PFMBench STABILITY: TAPE_StabilityDataset subset: TAPE_Stability (PFMBench split) Fine-tuned with an adapter under the PFMBench harness; train, validation and test counts are in Table 1. Author-reported evaluation · Evaluation metadata: source checked | ||
| 0.04924 spearman Unit: correlation · Direction: higher | Uncertainty: Not reported Scored: Not reported · Eligible: Not reported | source checkedpfmbench primary benchmark evidence · Table 3, row(ProTrek [ 56 ]), column(Stability) Source checking is not independent reproduction. |
| ProtST on PFMBench STABILITY: TAPE_Stability Configuration: ProtSTTask: PFMBench STABILITY: TAPE_StabilityDataset subset: TAPE_Stability (PFMBench split) Fine-tuned with an adapter under the PFMBench harness; train, validation and test counts are in Table 1. Author-reported evaluation · Evaluation metadata: source checked | ||
| 0.06623 spearman Unit: correlation · Direction: higher | Uncertainty: Not reported Scored: Not reported · Eligible: Not reported | source checkedpfmbench primary benchmark evidence · Table 3, row(ProtST [ 71 ]), column(Stability) Source checking is not independent reproduction. |
| ProtT5 on PFMBench STABILITY: TAPE_Stability Configuration: ProtT5Task: PFMBench STABILITY: TAPE_StabilityDataset subset: TAPE_Stability (PFMBench split) Fine-tuned with an adapter under the PFMBench harness; train, validation and test counts are in Table 1. Author-reported evaluation · Evaluation metadata: source checked | ||
| 0.18638 spearman Unit: correlation · Direction: higher | Uncertainty: Not reported Scored: Not reported · Eligible: Not reported | source checkedpfmbench primary benchmark evidence · Table 3, row(ProtT5 [ 11 ]), column(Stability) Source checking is not independent reproduction. |
| SaProt on PFMBench STABILITY: TAPE_Stability Configuration: SaProtTask: PFMBench STABILITY: TAPE_StabilityDataset subset: TAPE_Stability (PFMBench split) Fine-tuned with an adapter under the PFMBench harness; train, validation and test counts are in Table 1. Author-reported evaluation · Evaluation metadata: source checked | ||
| 0.24804 spearman Unit: correlation · Direction: higher | Uncertainty: Not reported Scored: Not reported · Eligible: Not reported | source checkedpfmbench primary benchmark evidence · Table 3, row(SaProt [ 55 ]), column(Stability) Source checking is not independent reproduction. |
| VenusPLM on PFMBench STABILITY: TAPE_Stability Configuration: VenusPLMTask: PFMBench STABILITY: TAPE_StabilityDataset subset: TAPE_Stability (PFMBench split) Fine-tuned with an adapter under the PFMBench harness; train, validation and test counts are in Table 1. Author-reported evaluation · Evaluation metadata: source checked | ||
| 0.33907 spearman Unit: correlation · Direction: higher | Uncertainty: Not reported Scored: Not reported · Eligible: Not reported | source checkedpfmbench primary benchmark evidence · Table 3, row(VenusPLM [ 59 ]), column(Stability) Source checking is not independent reproduction. |
Evidence table
Inspect claims, sources and review details
Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.
One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.
1 evidence row matching the loaded filters
| Property and statement | Original source and location | Review and provenance |
|---|---|---|
| Relationship: part of discovery-benchmark-pfmbench Individual claims | pfmbench primary benchmark evidence Table 1, row(TAPE_Stability) Version: 2506.14796v1 | source checked automated source review · 2026-09-18 Audit detailsPrimary-source transcription with no human sign-off and no independent reproduction. Field: Claim: pfmbench-association-stability Source artifact SHA-256: Hash scope: Exact retrieved primary paper artifact bytes. |
Sources and history
Release 2026-09-17-134cd1815de8 · Record review: source checked
1 source records and release history
- pfmbench primary benchmark evidence · Original source · 2506.14796v1
Technical metadata and extraction receipts
Stable ID: pfmbench-task-stability
- areas
- proteins-complexes
- tasks
- TAPE_Stability
- metric
- Spearman
- metric direction
- higher
- dataset
- TAPE_Stability
- protocol
- Fine-tuned with an adapter under the PFMBench harness; train, validation and test counts are in Table 1.
- source locator
- Table 1, row(TAPE_Stability)
- comparison panels
- id: pfmbench-panel-stability; title: PFMBench STABILITY: TAPE_Stability; protocol id: pfmbench-task-stability; dataset id: pfmbench-dataset-tape-stability; metric: spearman; unit: correlation; direction: higher; result ids: pfmbench-result-esm-2-stability-spearman; pfmbench-result-venusplm-stability-spearman; pfmbench-result-esm-c-stability-spearman; pfmbench-result-protgpt2-stability-spearman; pfmbench-result-pglm-stability-spearman; pfmbench-result-prott5-stability-spearman; pfmbench-result-dplm-stability-spearman; pfmbench-result-saprot-stability-spearman; pfmbench-result-prostt5-stability-spearman; pfmbench-result-protst-stability-spearman; pfmbench-result-esm3-stability-spearman; pfmbench-result-protrek-stability-spearman; source ids: expansion-p3-pfmbench; source locator: Table 1, row(TAPE_Stability); context: Every method PFMBench reports on TAPE_Stability, scored with Spearman on TAPE_Stability.; caveats: Author-reported numbers, source checked but not independently reproduced.; The metric differs by task, taken from Table 1, so these figures cannot be averaged into one score.; Table 3 scores come from adapter fine-tuning; the ProteinGym figure is zero-shot and is not comparable to them.; review: method: automated_source_review; date: 2026-09-18
Related records
- part of: PFMBench
- subject: PFMBench STABILITY: part of discovery-benchmark-pfmbench
- benchmark: DPLM on PFMBench STABILITY: TAPE_Stability
- benchmark: ESM-2 on PFMBench STABILITY: TAPE_Stability
- benchmark: ESM-C on PFMBench STABILITY: TAPE_Stability
- benchmark: ESM3 on PFMBench STABILITY: TAPE_Stability
- benchmark: PGLM on PFMBench STABILITY: TAPE_Stability
- benchmark: ProstT5 on PFMBench STABILITY: TAPE_Stability
- benchmark: ProtGPT2 on PFMBench STABILITY: TAPE_Stability
- benchmark: ProTrek on PFMBench STABILITY: TAPE_Stability
- benchmark: ProtST on PFMBench STABILITY: TAPE_Stability
- benchmark: ProtT5 on PFMBench STABILITY: TAPE_Stability
- benchmark: SaProt on PFMBench STABILITY: TAPE_Stability
- benchmark: VenusPLM on PFMBench STABILITY: TAPE_Stability