rewire.it
Evaluator

scPertEval

scPertEval evaluates and calibrates scoring protocols for single-cell perturbation predictions.

SourcesVirtual-Cell-Research-Community/scPertEval official source · Pinned README: toolkit purpose; scoring/calibration/DE actions; protocol taxonomy

0 evaluations · 0 metric rows

At a glance

Inputs, training, access and other details

Explanatory profile: source reviewed · Automated source review, 2026-09-16. Review applies to the cited claims; unresolved fields are listed below. Numerical results retain their own review status.

Data, procedure and scoring
PropertyDescription and evidence
DatasetsPredicted and observed perturbation responses; the associated study evaluates protocols across public datasets.
SourcesVirtual-Cell-Research-Community/scPertEval official source · Pinned README: toolkit purpose; scoring/calibration/DE actions; protocol taxonomy
SplitsThe evaluator scores supplied outputs; predictor train/test partitions belong to the dataset/run being evaluated. · Not applicable
SourcesVirtual-Cell-Research-Community/scPertEval official source · Pinned README: toolkit purpose; scoring/calibration/DE actions; protocol taxonomy
MetricsProtocol scoring is separated from calibration against empirical positive/negative controls using DRF and BDS.
SourcesVirtual-Cell-Research-Community/scPertEval official source · Pinned README: toolkit purpose; scoring/calibration/DE actions; protocol taxonomy
BaselinesEmpirical controls calibrate how well a protocol distinguishes expected response quality.
SourcesVirtual-Cell-Research-Community/scPertEval official source · Pinned README: toolkit purpose; scoring/calibration/DE actions; protocol taxonomy
Leakage controlsThe metric implementation receives ground truth, predictions and an evaluation context; it does not construct model-training partitions or audit training data. Predictor leakage controls belong to the protocol and dataset used to produce the submitted predictions. · Not applicable
Sourcesscperteval0 primary benchmark evidence · Pinned src/scperteval/protocols/metrics.py: metric input contract and context
UncertaintyUncertainty across samples, datasets or training runs must be defined by the evaluation study; this evaluator entry does not fix one experiment. · Not applicable
SourcesVirtual-Cell-Research-Community/scPertEval official source · Pinned README: toolkit purpose; scoring/calibration/DE actions; protocol taxonomy
Entity typePerturbation scoring-protocol calibration toolkit.
SourcesVirtual-Cell-Research-Community/scPertEval official source · Pinned README: toolkit purpose; scoring/calibration/DE actions; protocol taxonomy
OrganismsOrganism scope belongs to the selected perturbation dataset. · Not applicable
SourcesVirtual-Cell-Research-Community/scPertEval official source · Pinned README: toolkit purpose; scoring/calibration/DE actions; protocol taxonomy
AssaysObserved single-cell perturbation responses and empirical positive/negative controls.
SourcesVirtual-Cell-Research-Community/scPertEval official source · Pinned README: toolkit purpose; scoring/calibration/DE actions; protocol taxonomy
Allowed inputsPredicted/observed responses plus a protocol specifying representation, metric, transformation and reporting.
SourcesVirtual-Cell-Research-Community/scPertEval official source · Pinned README: toolkit purpose; scoring/calibration/DE actions; protocol taxonomy
AdaptationThe toolkit scores and calibrates evaluation protocols; it does not impose predictor fine-tuning. · Not applicable
SourcesVirtual-Cell-Research-Community/scPertEval official source · Pinned README: toolkit purpose; scoring/calibration/DE actions; protocol taxonomy
ImplementationNot extracted or verified for this record.

How it works

How it worksEvaluation procedure
Evaluation procedure1. Allowed inputs: Predicted/observed responses plus a protocol specifying representation, metric, transformation and reporting.. Then: 2. Splits: The evaluator scores supplied outputs; predictor train/test partitions belong to the dataset/run being evaluated.. Then: 3. Metrics: Protocol scoring is separated from calibration against empirical positive/negative controls using DRF and BDS.Evaluation procedure1. Allowed inputs: Predicted/observed responses plus a protocol specifying representation, metric, transformation and reporting.. Then: 2. Splits: The evaluator scores supplied outputs; predictor train/test partitions belong to the dataset/run being evaluated.. Then: 3. Metrics: Protocol scoring is separated from calibration against empirical positive/negative controls using DRF and BDS.Evaluation procedure1. Allowed inputs: Predicted/observed responses plus a protocol specifying representation, metric, transformation and reporting.. Then: 2. Splits: The evaluator scores supplied outputs; predictor train/test partitions belong to the dataset/run being evaluated.. Then: 3. Metrics: Protocol scoring is separated from calibration against empirical positive/negative controls using DRF and BDS.

Conceptual procedure. Task variants and protocol versions retain their separate scoring conditions.

SourcesVirtual-Cell-Research-Community/scPertEval official source · Pinned README: toolkit purpose; scoring/calibration/DE actions; protocol taxonomy
Evaluation methodology

scPertEval supplies explicit scoring protocols for perturbation predictions. A metric can compare one perturbation or operate across the complete perturbation panel, depending on its declared scope. Ground-truth references and context are passed to the evaluator, so protocol and preprocessing choices remain part of the result.

Sourcesscperteval0 primary benchmark evidence · Pinned src/scperteval/protocols/metrics.py: metric input contract and context

Tested entities and results

Release 2026-09-17-d277315f7d76 · 0 evaluations · 0 metric rows. Different protocols are not a single leaderboard.

No evaluations linked in this release.

Papers and result coverage

Last literature check: 2026-09-17. Primary-paper discovery and source inspection. Source-checked results are not independently reproduced experiments.

Paper or primary resourceVersionReference
Towards Principled Evaluation of Single-Cell Perturbation Prediction ModelsbioRxiv 2026.07.23.740433v1Read source
DOI: 10.64898/2026.07.23.740433
scPertEval documentationDocumentation snapshot at retrieval; byte-pinned by SHA-256Read source

What is still missing

  • Supplementary Figure 1 reuses a separate Wei et al. benchmark; its origin must be retained rather than attributed as newly run scPertEval model results.
  • No universal default metric/normalization or scientific model winner is implied by the evaluator.
Search and extraction details

primary protocol screened

Searches

  • "PMC13182013"
  • "PMC12619997"
  • ProkBERT promoter benchmark
  • scPertEval benchmark paper
  • Boltz-2 affinity benchmark paper
  • ESMFold monomer structure benchmark Science 2023
  • "ProkBERT" "paper" "2024"
  • "scPertEval"
  • ProkBERT family prokaryotic language models microbe paper
  • Probabilistic harmonization annotation single-cell transcriptomics scANVI Nature Methods 2021
  • Evolutionary-scale prediction atomic-level protein structure language model ESMFold Science Lin 2023
  • Towards Principled Evaluation Single-Cell Perturbation Prediction Models Schäfer 2026

Evidence locations

  • Primary paper abstract, Sections 2–6 and Table 1
  • Supplementary Figure 1; Supplementary Table 1
  • Data/code availability

Strengths and limitations

Limitations and conditions

  • Metric computation alone cannot establish whether a predictor saw a held-out perturbation or cell context during training. That evidence must come from the model and dataset protocol.
    Sourcesscperteval0 primary benchmark evidence · Pinned src/scperteval/protocols/metrics.py: metric input contract and context
Profile review details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Stable record: discovery-benchmark-scperteval

Applicable tests and references

Applicability is distinct from a completed evaluation.

Evidence table

Inspect claims, sources and review details

Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.

One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.

18 evidence rows matching the loaded filters

Claims, original sources and review scope · Release 2026-09-17-d277315f7d76
Property and statementOriginal source and locationReview and provenance
Diagram caption

Conceptual procedure. Task variants and protocol versions retain their separate scoring conditions.

Individual claims
Virtual-Cell-Research-Community/scPertEval official source

Original source ↗

Pinned README: toolkit purpose; scoring/calibration/DE actions; protocol taxonomy

Version: 4685f11927e887745737600170da7a655b727553
Retrieved: 2026-09-16T10:30:23.079976+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.diagram.caption

Source artifact SHA-256: d8522e585806ec0008a36558e0dd1f6f0b0bb4deffb3f64a7ee4f09cd087e9f4

Hash scope: Hash scope not separately documented; inspect source record

Diagram steps

["Allowed inputs: Predicted/observed responses plus a protocol specifying representation, metric, transformation and reporting.","Splits: The evaluator scores supplied outputs; predictor train/test partitions belong to the dataset/run being evaluated.","Metrics: Protocol scoring is separated from calibration against empirical positive/negative controls using DRF and BDS."]

Individual claims
Virtual-Cell-Research-Community/scPertEval official source

Original source ↗

Pinned README: toolkit purpose; scoring/calibration/DE actions; protocol taxonomy

Version: 4685f11927e887745737600170da7a655b727553
Retrieved: 2026-09-16T10:30:23.079976+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.diagram.steps

Source artifact SHA-256: d8522e585806ec0008a36558e0dd1f6f0b0bb4deffb3f64a7ee4f09cd087e9f4

Hash scope: Hash scope not separately documented; inspect source record

Diagram title

Evaluation procedure

Individual claims
Virtual-Cell-Research-Community/scPertEval official source

Original source ↗

Pinned README: toolkit purpose; scoring/calibration/DE actions; protocol taxonomy

Version: 4685f11927e887745737600170da7a655b727553
Retrieved: 2026-09-16T10:30:23.079976+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.diagram.title

Source artifact SHA-256: d8522e585806ec0008a36558e0dd1f6f0b0bb4deffb3f64a7ee4f09cd087e9f4

Hash scope: Hash scope not separately documented; inspect source record

Datasets

Predicted and observed perturbation responses; the associated study evaluates protocols across public datasets.

Individual claims
Virtual-Cell-Research-Community/scPertEval official source

Original source ↗

Pinned README: toolkit purpose; scoring/calibration/DE actions; protocol taxonomy

Version: 4685f11927e887745737600170da7a655b727553
Retrieved: 2026-09-16T10:30:23.079976+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.facts.0.value

Source artifact SHA-256: d8522e585806ec0008a36558e0dd1f6f0b0bb4deffb3f64a7ee4f09cd087e9f4

Hash scope: Hash scope not separately documented; inspect source record

Splits

The evaluator scores supplied outputs; predictor train/test partitions belong to the dataset/run being evaluated.

Individual claims
Virtual-Cell-Research-Community/scPertEval official source

Original source ↗

Pinned README: toolkit purpose; scoring/calibration/DE actions; protocol taxonomy

Version: 4685f11927e887745737600170da7a655b727553
Retrieved: 2026-09-16T10:30:23.079976+00:00

inapplicable

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.facts.1.value

Source artifact SHA-256: d8522e585806ec0008a36558e0dd1f6f0b0bb4deffb3f64a7ee4f09cd087e9f4

Hash scope: Hash scope not separately documented; inspect source record

Adaptation

The toolkit scores and calibrates evaluation protocols; it does not impose predictor fine-tuning.

Individual claims
Virtual-Cell-Research-Community/scPertEval official source

Original source ↗

Pinned README: toolkit purpose; scoring/calibration/DE actions; protocol taxonomy

Version: 4685f11927e887745737600170da7a655b727553
Retrieved: 2026-09-16T10:30:23.079976+00:00

inapplicable

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.facts.10.value

Source artifact SHA-256: d8522e585806ec0008a36558e0dd1f6f0b0bb4deffb3f64a7ee4f09cd087e9f4

Hash scope: Hash scope not separately documented; inspect source record

Metrics

Protocol scoring is separated from calibration against empirical positive/negative controls using DRF and BDS.

Individual claims
Virtual-Cell-Research-Community/scPertEval official source

Original source ↗

Pinned README: toolkit purpose; scoring/calibration/DE actions; protocol taxonomy

Version: 4685f11927e887745737600170da7a655b727553
Retrieved: 2026-09-16T10:30:23.079976+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.facts.2.value

Source artifact SHA-256: d8522e585806ec0008a36558e0dd1f6f0b0bb4deffb3f64a7ee4f09cd087e9f4

Hash scope: Hash scope not separately documented; inspect source record

Baselines

Empirical controls calibrate how well a protocol distinguishes expected response quality.

Individual claims
Virtual-Cell-Research-Community/scPertEval official source

Original source ↗

Pinned README: toolkit purpose; scoring/calibration/DE actions; protocol taxonomy

Version: 4685f11927e887745737600170da7a655b727553
Retrieved: 2026-09-16T10:30:23.079976+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.facts.3.value

Source artifact SHA-256: d8522e585806ec0008a36558e0dd1f6f0b0bb4deffb3f64a7ee4f09cd087e9f4

Hash scope: Hash scope not separately documented; inspect source record

Leakage controls

The metric implementation receives ground truth, predictions and an evaluation context; it does not construct model-training partitions or audit training data. Predictor leakage controls belong to the protocol and dataset used to produce the submitted predictions.

Individual claims
scperteval0 primary benchmark evidence

Original source ↗

Pinned src/scperteval/protocols/metrics.py: metric input contract and context

Version: 4685f11927e887745737600170da7a655b727553:src/scperteval/protocols/metrics.py
Retrieved: 2026-09-16T21:11:45.124583+00:00

inapplicable

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.facts.4.value

Source artifact SHA-256: c4dbfdbc539ba78350ddba68ca4c03edfe89cf3625c6c699884bd3a85e47d929

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Uncertainty

Uncertainty across samples, datasets or training runs must be defined by the evaluation study; this evaluator entry does not fix one experiment.

Individual claims
Virtual-Cell-Research-Community/scPertEval official source

Original source ↗

Pinned README: toolkit purpose; scoring/calibration/DE actions; protocol taxonomy

Version: 4685f11927e887745737600170da7a655b727553
Retrieved: 2026-09-16T10:30:23.079976+00:00

inapplicable

automated source review · 2026-09-16

Audit details

Primary paper and/or task implementation reviewed for the explicitly cited methodology claims. Scope-limited absence is recorded only after the documented source search; no model runs or independent reproduction.

Field: attributes.profile.facts.5.value

Source artifact SHA-256: d8522e585806ec0008a36558e0dd1f6f0b0bb4deffb3f64a7ee4f09cd087e9f4

Hash scope: Hash scope not separately documented; inspect source record

Sources and history

Release 2026-09-17-d277315f7d76 · Record review: discovered

4 source records and release historyDownload this release
Technical metadata and extraction receipts

Stable ID: discovery-benchmark-scperteval

areas
single-cell
entity level
evaluator
scope note
Specialist molecular or omics evaluation; protocol details require review before numerical comparison.
task
Perturbation prediction metric calibration
version
Not reported
benchmark research
review date: 2026-09-17; status: primary_protocol_screened; primary sources: expansion-p3-scperteval-paper; expansion-p3-scperteval-docs; inspected locators: Primary paper abstract, Sections 2–6 and Table 1; Supplementary Figure 1; Supplementary Table 1; Data/code availability; searched queries: "PMC13182013"; "PMC12619997"; ProkBERT promoter benchmark; scPertEval benchmark paper; Boltz-2 affinity benchmark paper; ESMFold monomer structure benchmark Science 2023; "ProkBERT" "paper" "2024"; "scPertEval"; ProkBERT family prokaryotic language models microbe paper; Probabilistic harmonization annotation single-cell transcriptomics scANVI Nature Methods 2021; Evolutionary-scale prediction atomic-level protein structure language model ESMFold Science Lin 2023; Towards Principled Evaluation Single-Cell Perturbation Prediction Models Schäfer 2026; gaps: Supplementary Figure 1 reuses a separate Wei et al. benchmark; its origin must be retained rather than attributed as newly run scPertEval model results.; No universal default metric/normalization or scientific model winner is implied by the evaluator.; claim scope: Primary-paper discovery and source inspection. Source-checked results are not independently reproduced experiments.
historical missing metadata
dataset release: unextracted; metric implementation: unextracted; split manifest: unextracted; version: unextracted
metadata review scope
historical_missing_metadata preserves the original discovery state. Current descriptive evidence and missingness are recorded in profile.facts; numerical-result review is separate.
legacy kinds
benchmark
entity classification
review date: 2026-09-17; rationale: The cited profile describes scoring or assessment software/procedures applied to predictions, rather than the biological target or the underlying dataset.; source ids: src-discovery-virtual-cell-research-community-scperteval; source locator: Pinned README: toolkit purpose; scoring/calibration/DE actions; protocol taxonomy; ambiguities: None recorded
Related records

Suggest a correction