rewire.it
Task

human-versus-viral protein classification

This evaluation asks whether protein representations distinguish human and viral sequence labels. It is a classification task, not a direct test of immune function.

SourcesProtein Language Models Expose Viral Immune Mimicry · Abstract; Methods 2.1 Protein Datasets and 2.6 Model Performance; Table 1

8 evaluations · 32 metric rows

At a glance

Inputs, training, access and other details

Explanatory profile: limited source coverage · Automated source review, 2026-09-16. Review applies to the cited claims; unresolved fields are listed below. Numerical results retain their own review status.

Data, procedure and scoring
PropertyDescription and evidence
Entity typePaper-specific evaluation task; this profile is a descriptive evidence summary.
SourcesProtein Language Models Expose Viral Immune Mimicry · Abstract; Methods 2.1 Protein Datasets and 2.6 Model Performance; Table 1
DatasetsReviewed human and vertebrate-host viral protein records from Swiss-Prot/UniProtKB, with redundancy filtering described in Methods 2.1.
SourcesProtein Language Models Expose Viral Immune Mimicry · Abstract; Methods 2.1 Protein Datasets and 2.6 Model Performance; Table 1
OrganismsHuman proteins and proteins from viruses with a known vertebrate host.
SourcesProtein Language Models Expose Viral Immune Mimicry · Abstract; Methods 2.1 Protein Datasets and 2.6 Model Performance; Table 1
AssaysSequence-origin labels from curated database records. This classification endpoint is not an experimental immune-response assay.
SourcesProtein Language Models Expose Viral Immune Mimicry · Abstract; Methods 2.1 Protein Datasets and 2.6 Model Performance; Table 1
SplitsMethods 2.1 assigns whole UniRef50 clusters to training or test sets. This profile records the split principle, not a verified membership manifest.
SourcesProtein Language Models Expose Viral Immune Mimicry · Abstract; Methods 2.1 Protein Datasets and 2.6 Model Performance; Table 1
Allowed inputsProtein sequence representations paired with the study’s human/viral labels.
SourcesProtein Language Models Expose Viral Immune Mimicry · Abstract; Methods 2.1 Protein Datasets and 2.6 Model Performance; Table 1
AdaptationPretrained protein representations are evaluated through a study-specific classifier; a backbone name alone does not identify the full fitted pipeline.
SourcesProtein Language Models Expose Viral Immune Mimicry · Abstract; Methods 2.1 Protein Datasets and 2.6 Model Performance; Table 1
MetricsAUROC, log loss, accuracy, precision and recall are described; precision and recall use macro averaging. The linked result retains its original percentage unit.
SourcesProtein Language Models Expose Viral Immune Mimicry · Abstract; Methods 2.1 Protein Datasets and 2.6 Model Performance; Table 1
BaselinesTable 1 compares the reported protein-representation configurations. They are classification comparators, not experimental immune-function controls.
SourcesProtein Language Models Expose Viral Immune Mimicry · Abstract; Methods 2.1 Protein Datasets and 2.6 Model Performance; Table 1

How it works

How it worksConceptual assessment outline
Conceptual assessment outline1. Curated sequence-origin labels. Then: 2. Keep sequence clusters separate. Then: 3. Assess held-out classifications. Then: 4. Report classification metricsConceptual assessment outline1. Curated sequence-origin labels. Then: 2. Keep sequence clusters separate. Then: 3. Assess held-out classifications. Then: 4. Report classification metricsConceptual assessment outline1. Curated sequence-origin labels. Then: 2. Keep sequence clusters separate. Then: 3. Assess held-out classifications. Then: 4. Report classification metrics

Conceptual overview of the published statistical assessment.

SourcesProtein Language Models Expose Viral Immune Mimicry · Abstract; Methods 2.1 Protein Datasets and 2.6 Model Performance; Table 1
What the evaluation establishes

The source uses reviewed protein records and separates training and test data by sequence clusters. That reduces direct overlap between related examples under the stated clustering rule. It reports classification metrics for different representations; classification errors and biological explanations of those errors are separate claims.

SourcesProtein Language Models Expose Viral Immune Mimicry · Abstract; Methods 2.1 Protein Datasets and 2.6 Model Performance; Table 1

Evaluation design

Benchmarks bring together tasks and protocols. A task describes the biological question; a protocol defines a particular test.

These source-backed links do not make different protocols or scores interchangeable.

Published comparisons

Explore the results reported under one evaluation protocol. Each figure keeps its source, dataset and metric together; it is not a ranking across studies.

Held-out human-versus-virus protein classification · Table 1

AUROC (percent) · Higher values are better for this metric.

Compare protein origin classifiers; T 5 embeddings plus linear/tree models vs ESM 2 fine-tuning and simple controls. UniRef90 deduplication; proteins longer than1,600 residues excluded; UniRef50 clusters assigned80% training and20% test with no cluster shared. Separate four-fold error-analysis experiment not assigned toTable1.

Evaluation protocol · human and viral proteins

  1. BL Length · Configuration · Author-reported evaluation61.97
  2. AA n-grams · Configuration · Author-reported evaluation91.5
  3. ESM2 8M · Configuration · Author-reported evaluation98.09
  4. ESM2 35M · Configuration · Author-reported evaluation98.69
  5. ESM2 150M · Configuration · Author-reported evaluation99.26
  6. ESM2 650M · Configuration · Author-reported evaluation99.67
  7. Linear-T5 · Configuration · Author-reported evaluation99.56
  8. Tree-T5 · Configuration · Author-reported evaluation99.65

Source order is preserved. Plotted marks show point estimates; uncertainty, where reported, is retained in the printed values and table. Differences do not establish statistical significance.

Protein Language Models Expose Viral Immune Mimicry · Table 1: AUC (%), Held-out human-versus-virus protein classification
Values, uncertainty and evidence
AUROC: original source values
Tested entityPrinted valueUncertaintyEvidence
BL Length · Configuration61.97 percentNot reportedAuthor-reported evaluation · source checkedProtein Language Models Expose Viral Immune Mimicry · Table 1, row BL Length, column AUC (%); XML row2 column2
AA n-grams · Configuration91.5 percentNot reportedAuthor-reported evaluation · source checkedProtein Language Models Expose Viral Immune Mimicry · Table 1, row AA n-grams, column AUC (%); XML row3 column2
ESM2 8M · Configuration98.09 percentNot reportedAuthor-reported evaluation · source checkedProtein Language Models Expose Viral Immune Mimicry · Table 1, row ESM2 8M, column AUC (%); XML row4 column2
ESM2 35M · Configuration98.69 percentNot reportedAuthor-reported evaluation · source checkedProtein Language Models Expose Viral Immune Mimicry · Table 1, row ESM2 35M, column AUC (%); XML row5 column2
ESM2 150M · Configuration99.26 percentNot reportedAuthor-reported evaluation · source checkedProtein Language Models Expose Viral Immune Mimicry · Table 1, row ESM2 150M, column AUC (%); XML row6 column2
ESM2 650M · Configuration99.67 percentNot reportedAuthor-reported evaluation · source checkedProtein Language Models Expose Viral Immune Mimicry; Protein Language Models Expose Viral Immune Mimicry · Table 1, ESM2 650M row, AUC (%) column
Linear-T5 · Configuration99.56 percentNot reportedAuthor-reported evaluation · source checkedProtein Language Models Expose Viral Immune Mimicry · Table 1, row Linear-T5, column AUC (%); XML row8 column2
Tree-T5 · Configuration99.65 percentNot reportedAuthor-reported evaluation · source checkedProtein Language Models Expose Viral Immune Mimicry · Table 1, row Tree-T5, column AUC (%); XML row9 column2
Scope and limitations
  • Origin classification does not establish immune mimicry.
  • Table footnote says all values are percent but log-loss values are dimensionless; log-loss units quarantined.
  • Narrative n-gram AUC91.9 differs from Table1 value91.5; preserve table value.
  • No interval assigned unless printed in source cell.

Source transcription and grouping reviewed by automated source review on 2026-09-17. These experiments were not independently reproduced by rewire.

Tested entities and results

Release 2026-09-17-d277315f7d76 · 8 evaluations · 32 metric rows. Different protocols are not a single leaderboard.

Results grouped by the exact reported evaluation
Metric and findingCoverage and uncertaintyEvidence
ESM2 650M: human-versus-viral protein classification

Compare protein origin classifiers; T 5 embeddings plus linear/tree models vs ESM 2 fine-tuning and simple controls. UniRef90 deduplication; proteins longer than1,600 residues excluded; UniRef50 clusters assigned80% training and20% test with no cluster shared. Separate four-fold error-analysis experiment not assigned toTable1.

Author-reported evaluation · Evaluation metadata: needs review

99.67% AUROC

Unit: percent · Direction: higher

Uncertainty: not reported in legacy extract

Scored: Not reported · Eligible: Not reported

source checkedProtein Language Models Expose Viral Immune Mimicry; Protein Language Models Expose Viral Immune Mimicry · Table 1, ESM2 650M row, AUC (%) column

Source checking is not independent reproduction.

96.85% Prec.

Unit: percent · Direction: higher

Uncertainty: unreported

Scored: Not reported · Eligible: Not reported

source checkedProtein Language Models Expose Viral Immune Mimicry · Table 1, row ESM2 650M, column Prec.; XML row7 column4

Source checking is not independent reproduction.

97.86% Accur.

Unit: percent · Direction: higher

Uncertainty: unreported

Scored: Not reported · Eligible: Not reported

source checkedProtein Language Models Expose Viral Immune Mimicry · Table 1, row ESM2 650M, column Accur.; XML row7 column3

Source checking is not independent reproduction.

96.68% Recall

Unit: percent · Direction: higher

Uncertainty: unreported

Scored: Not reported · Eligible: Not reported

source checkedProtein Language Models Expose Viral Immune Mimicry · Table 1, row ESM2 650M, column Recall; XML row7 column5

Source checking is not independent reproduction.

ESM2 8M: Held-out human-versus-virus protein classification

Compare protein origin classifiers; T 5 embeddings plus linear/tree models vs ESM 2 fine-tuning and simple controls. UniRef90 deduplication; proteins longer than1,600 residues excluded; UniRef50 clusters assigned80% training and20% test with no cluster shared. Separate four-fold error-analysis experiment not assigned toTable1.

Author-reported evaluation · Evaluation metadata: needs review

92.15% Prec.

Unit: percent · Direction: higher

Uncertainty: unreported

Scored: Not reported · Eligible: Not reported

source checkedProtein Language Models Expose Viral Immune Mimicry · Table 1, row ESM2 8M, column Prec.; XML row4 column4

Source checking is not independent reproduction.

92.33% Recall

Unit: percent · Direction: higher

Uncertainty: unreported

Scored: Not reported · Eligible: Not reported

source checkedProtein Language Models Expose Viral Immune Mimicry · Table 1, row ESM2 8M, column Recall; XML row4 column5

Source checking is not independent reproduction.

98.09% AUROC

Unit: percent · Direction: higher

Uncertainty: unreported

Scored: Not reported · Eligible: Not reported

source checkedProtein Language Models Expose Viral Immune Mimicry · Table 1, row ESM2 8M, column AUC (%); XML row4 column2

Source checking is not independent reproduction.

AA n-grams: Held-out human-versus-virus protein classification

Compare protein origin classifiers; T 5 embeddings plus linear/tree models vs ESM 2 fine-tuning and simple controls. UniRef90 deduplication; proteins longer than1,600 residues excluded; UniRef50 clusters assigned80% training and20% test with no cluster shared. Separate four-fold error-analysis experiment not assigned toTable1.

Author-reported evaluation · Evaluation metadata: needs review

91.5% AUROC

Unit: percent · Direction: higher

Uncertainty: unreported

Scored: Not reported · Eligible: Not reported

source checkedProtein Language Models Expose Viral Immune Mimicry · Table 1, row AA n-grams, column AUC (%); XML row3 column2

Source checking is not independent reproduction.

88.49% Prec.

Unit: percent · Direction: higher

Uncertainty: unreported

Scored: Not reported · Eligible: Not reported

source checkedProtein Language Models Expose Viral Immune Mimicry · Table 1, row AA n-grams, column Prec.; XML row3 column4

Source checking is not independent reproduction.

ESM2 35M: Held-out human-versus-virus protein classification

Compare protein origin classifiers; T 5 embeddings plus linear/tree models vs ESM 2 fine-tuning and simple controls. UniRef90 deduplication; proteins longer than1,600 residues excluded; UniRef50 clusters assigned80% training and20% test with no cluster shared. Separate four-fold error-analysis experiment not assigned toTable1.

Author-reported evaluation · Evaluation metadata: needs review

98.69% AUROC

Unit: percent · Direction: higher

Uncertainty: unreported

Scored: Not reported · Eligible: Not reported

source checkedProtein Language Models Expose Viral Immune Mimicry · Table 1, row ESM2 35M, column AUC (%); XML row5 column2

Source checking is not independent reproduction.

93.81% Prec.

Unit: percent · Direction: higher

Uncertainty: unreported

Scored: Not reported · Eligible: Not reported

source checkedProtein Language Models Expose Viral Immune Mimicry · Table 1, row ESM2 35M, column Prec.; XML row5 column4

Source checking is not independent reproduction.

93.92% Recall

Unit: percent · Direction: higher

Uncertainty: unreported

Scored: Not reported · Eligible: Not reported

source checkedProtein Language Models Expose Viral Immune Mimicry · Table 1, row ESM2 35M, column Recall; XML row5 column5

Source checking is not independent reproduction.

95.83% Accur.

Unit: percent · Direction: higher

Uncertainty: unreported

Scored: Not reported · Eligible: Not reported

source checkedProtein Language Models Expose Viral Immune Mimicry · Table 1, row ESM2 35M, column Accur.; XML row5 column3

Source checking is not independent reproduction.

ESM2 150M: Held-out human-versus-virus protein classification

Compare protein origin classifiers; T 5 embeddings plus linear/tree models vs ESM 2 fine-tuning and simple controls. UniRef90 deduplication; proteins longer than1,600 residues excluded; UniRef50 clusters assigned80% training and20% test with no cluster shared. Separate four-fold error-analysis experiment not assigned toTable1.

Author-reported evaluation · Evaluation metadata: needs review

95.54% Prec.

Unit: percent · Direction: higher

Uncertainty: unreported

Scored: Not reported · Eligible: Not reported

source checkedProtein Language Models Expose Viral Immune Mimicry · Table 1, row ESM2 150M, column Prec.; XML row6 column4

Source checking is not independent reproduction.

96.99% Accur.

Unit: percent · Direction: higher

Uncertainty: unreported

Scored: Not reported · Eligible: Not reported

source checkedProtein Language Models Expose Viral Immune Mimicry · Table 1, row ESM2 150M, column Accur.; XML row6 column3

Source checking is not independent reproduction.

95.48% Recall

Unit: percent · Direction: higher

Uncertainty: unreported

Scored: Not reported · Eligible: Not reported

source checkedProtein Language Models Expose Viral Immune Mimicry · Table 1, row ESM2 150M, column Recall; XML row6 column5

Source checking is not independent reproduction.

99.26% AUROC

Unit: percent · Direction: higher

Uncertainty: unreported

Scored: Not reported · Eligible: Not reported

source checkedProtein Language Models Expose Viral Immune Mimicry · Table 1, row ESM2 150M, column AUC (%); XML row6 column2

Source checking is not independent reproduction.

Tree-T5: Held-out human-versus-virus protein classification

Compare protein origin classifiers; T 5 embeddings plus linear/tree models vs ESM 2 fine-tuning and simple controls. UniRef90 deduplication; proteins longer than1,600 residues excluded; UniRef50 clusters assigned80% training and20% test with no cluster shared. Separate four-fold error-analysis experiment not assigned toTable1.

Author-reported evaluation · Evaluation metadata: needs review

97.7% Recall

Unit: percent · Direction: higher

Uncertainty: unreported

Scored: Not reported · Eligible: Not reported

source checkedProtein Language Models Expose Viral Immune Mimicry · Table 1, row Tree-T5, column Recall; XML row9 column5

Source checking is not independent reproduction.

97.7% Accur.

Unit: percent · Direction: higher

Uncertainty: unreported

Scored: Not reported · Eligible: Not reported

source checkedProtein Language Models Expose Viral Immune Mimicry · Table 1, row Tree-T5, column Accur.; XML row9 column3

Source checking is not independent reproduction.

99.65% AUROC

Unit: percent · Direction: higher

Uncertainty: unreported

Scored: Not reported · Eligible: Not reported

source checkedProtein Language Models Expose Viral Immune Mimicry · Table 1, row Tree-T5, column AUC (%); XML row9 column2

Source checking is not independent reproduction.

Linear-T5: Held-out human-versus-virus protein classification

Compare protein origin classifiers; T 5 embeddings plus linear/tree models vs ESM 2 fine-tuning and simple controls. UniRef90 deduplication; proteins longer than1,600 residues excluded; UniRef50 clusters assigned80% training and20% test with no cluster shared. Separate four-fold error-analysis experiment not assigned toTable1.

Author-reported evaluation · Evaluation metadata: needs review

97.57% Recall

Unit: percent · Direction: higher

Uncertainty: unreported

Scored: Not reported · Eligible: Not reported

source checkedProtein Language Models Expose Viral Immune Mimicry · Table 1, row Linear-T5, column Recall; XML row8 column5

Source checking is not independent reproduction.

97.57% Accur.

Unit: percent · Direction: higher

Uncertainty: unreported

Scored: Not reported · Eligible: Not reported

source checkedProtein Language Models Expose Viral Immune Mimicry · Table 1, row Linear-T5, column Accur.; XML row8 column3

Source checking is not independent reproduction.

BL Length: Held-out human-versus-virus protein classification

Compare protein origin classifiers; T 5 embeddings plus linear/tree models vs ESM 2 fine-tuning and simple controls. UniRef90 deduplication; proteins longer than1,600 residues excluded; UniRef50 clusters assigned80% training and20% test with no cluster shared. Separate four-fold error-analysis experiment not assigned toTable1.

Author-reported evaluation · Evaluation metadata: needs review

78.5% Accur.

Unit: percent · Direction: higher

Uncertainty: unreported

Scored: Not reported · Eligible: Not reported

source checkedProtein Language Models Expose Viral Immune Mimicry · Table 1, row BL Length, column Accur.; XML row2 column3

Source checking is not independent reproduction.

78.5% Prec.

Unit: percent · Direction: higher

Uncertainty: unreported

Scored: Not reported · Eligible: Not reported

source checkedProtein Language Models Expose Viral Immune Mimicry · Table 1, row BL Length, column Prec.; XML row2 column4

Source checking is not independent reproduction.

78.5% Recall

Unit: percent · Direction: higher

Uncertainty: unreported

Scored: Not reported · Eligible: Not reported

source checkedProtein Language Models Expose Viral Immune Mimicry · Table 1, row BL Length, column Recall; XML row2 column5

Source checking is not independent reproduction.

Papers and result coverage

Last literature check: 2026-09-17. Primary-source discovery and table/protocol screening; source checked is not independently reproduced. Raw acquisitions not automatically numerical publication approval.

Paper or primary resourceVersionReference
Protein Language Models Expose Viral Immune Mimicryversion of recordRead source
DOI: 10.3390/v17091199

What is still missing

  • Independent batch review before import; preserve existing observation identities.
Search and extraction details

complete comparison tables extracted pending publication review

Searches

  • Protein Language Models Expose Viral Immune Mimicry primary paper benchmark results

Evidence locations

  • Table1; held-out protein-origin classification; separate error-analysis CV

Strengths and limitations

Strengths and considerations

Limitations and conditions

  • Distinguishing sequence origin does not establish an immune mechanism. Database selection, similarity filtering and label composition constrain generalisation.
    SourcesProtein Language Models Expose Viral Immune Mimicry · Abstract; Methods 2.1 Protein Datasets and 2.6 Model Performance; Table 1
Profile review details

Inspected the cited primary abstract and descriptive computational-evaluation sections. Review covers the descriptive claims shown; no executable protocol was reconstructed. Original numerical records retain their prior transcription review.

Stable record: reported-task-53506fe386e4a1

Evidence table

Inspect claims, sources and review details

Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.

One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.

16 evidence rows matching the loaded filters

Claims, original sources and review scope · Release 2026-09-17-d277315f7d76
Property and statementOriginal source and locationReview and provenance
Diagram caption

Conceptual overview of the published statistical assessment.

Individual claims
Protein Language Models Expose Viral Immune Mimicry

Original source ↗

Abstract; Methods 2.1 Protein Datasets and 2.6 Model Performance; Table 1

Version: version of record
Retrieved: 2026-09-16T10:33:57.274Z

source checked

automated source review · 2026-09-16

Audit details

Inspected the cited primary abstract and descriptive computational-evaluation sections. Review covers the descriptive claims shown; no executable protocol was reconstructed. Original numerical records retain their prior transcription review.

Field: attributes.profile.diagram.caption

Source artifact SHA-256: 15250af2f75f70e2b6a3725d00bc7e276ed9d2d4a54f0ae9d6eaabf6be13e4a1

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Diagram steps

["Curated sequence-origin labels","Keep sequence clusters separate","Assess held-out classifications","Report classification metrics"]

Individual claims
Protein Language Models Expose Viral Immune Mimicry

Original source ↗

Abstract; Methods 2.1 Protein Datasets and 2.6 Model Performance; Table 1

Version: version of record
Retrieved: 2026-09-16T10:33:57.274Z

source checked

automated source review · 2026-09-16

Audit details

Inspected the cited primary abstract and descriptive computational-evaluation sections. Review covers the descriptive claims shown; no executable protocol was reconstructed. Original numerical records retain their prior transcription review.

Field: attributes.profile.diagram.steps

Source artifact SHA-256: 15250af2f75f70e2b6a3725d00bc7e276ed9d2d4a54f0ae9d6eaabf6be13e4a1

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Diagram title

Conceptual assessment outline

Individual claims
Protein Language Models Expose Viral Immune Mimicry

Original source ↗

Abstract; Methods 2.1 Protein Datasets and 2.6 Model Performance; Table 1

Version: version of record
Retrieved: 2026-09-16T10:33:57.274Z

source checked

automated source review · 2026-09-16

Audit details

Inspected the cited primary abstract and descriptive computational-evaluation sections. Review covers the descriptive claims shown; no executable protocol was reconstructed. Original numerical records retain their prior transcription review.

Field: attributes.profile.diagram.title

Source artifact SHA-256: 15250af2f75f70e2b6a3725d00bc7e276ed9d2d4a54f0ae9d6eaabf6be13e4a1

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Entity type

Paper-specific evaluation task; this profile is a descriptive evidence summary.

Individual claims
Protein Language Models Expose Viral Immune Mimicry

Original source ↗

Abstract; Methods 2.1 Protein Datasets and 2.6 Model Performance; Table 1

Version: version of record
Retrieved: 2026-09-16T10:33:57.274Z

source checked

automated source review · 2026-09-16

Audit details

Inspected the cited primary abstract and descriptive computational-evaluation sections. Review covers the descriptive claims shown; no executable protocol was reconstructed. Original numerical records retain their prior transcription review.

Field: attributes.profile.facts.0.value

Source artifact SHA-256: 15250af2f75f70e2b6a3725d00bc7e276ed9d2d4a54f0ae9d6eaabf6be13e4a1

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Datasets

Reviewed human and vertebrate-host viral protein records from Swiss-Prot/UniProtKB, with redundancy filtering described in Methods 2.1.

Individual claims
Protein Language Models Expose Viral Immune Mimicry

Original source ↗

Abstract; Methods 2.1 Protein Datasets and 2.6 Model Performance; Table 1

Version: version of record
Retrieved: 2026-09-16T10:33:57.274Z

source checked

automated source review · 2026-09-16

Audit details

Inspected the cited primary abstract and descriptive computational-evaluation sections. Review covers the descriptive claims shown; no executable protocol was reconstructed. Original numerical records retain their prior transcription review.

Field: attributes.profile.facts.1.value

Source artifact SHA-256: 15250af2f75f70e2b6a3725d00bc7e276ed9d2d4a54f0ae9d6eaabf6be13e4a1

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Organisms

Human proteins and proteins from viruses with a known vertebrate host.

Individual claims
Protein Language Models Expose Viral Immune Mimicry

Original source ↗

Abstract; Methods 2.1 Protein Datasets and 2.6 Model Performance; Table 1

Version: version of record
Retrieved: 2026-09-16T10:33:57.274Z

source checked

automated source review · 2026-09-16

Audit details

Inspected the cited primary abstract and descriptive computational-evaluation sections. Review covers the descriptive claims shown; no executable protocol was reconstructed. Original numerical records retain their prior transcription review.

Field: attributes.profile.facts.2.value

Source artifact SHA-256: 15250af2f75f70e2b6a3725d00bc7e276ed9d2d4a54f0ae9d6eaabf6be13e4a1

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Assays

Sequence-origin labels from curated database records. This classification endpoint is not an experimental immune-response assay.

Individual claims
Protein Language Models Expose Viral Immune Mimicry

Original source ↗

Abstract; Methods 2.1 Protein Datasets and 2.6 Model Performance; Table 1

Version: version of record
Retrieved: 2026-09-16T10:33:57.274Z

source checked

automated source review · 2026-09-16

Audit details

Inspected the cited primary abstract and descriptive computational-evaluation sections. Review covers the descriptive claims shown; no executable protocol was reconstructed. Original numerical records retain their prior transcription review.

Field: attributes.profile.facts.3.value

Source artifact SHA-256: 15250af2f75f70e2b6a3725d00bc7e276ed9d2d4a54f0ae9d6eaabf6be13e4a1

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Splits

Methods 2.1 assigns whole UniRef50 clusters to training or test sets. This profile records the split principle, not a verified membership manifest.

Individual claims
Protein Language Models Expose Viral Immune Mimicry

Original source ↗

Abstract; Methods 2.1 Protein Datasets and 2.6 Model Performance; Table 1

Version: version of record
Retrieved: 2026-09-16T10:33:57.274Z

source checked

automated source review · 2026-09-16

Audit details

Inspected the cited primary abstract and descriptive computational-evaluation sections. Review covers the descriptive claims shown; no executable protocol was reconstructed. Original numerical records retain their prior transcription review.

Field: attributes.profile.facts.4.value

Source artifact SHA-256: 15250af2f75f70e2b6a3725d00bc7e276ed9d2d4a54f0ae9d6eaabf6be13e4a1

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Allowed inputs

Protein sequence representations paired with the study’s human/viral labels.

Individual claims
Protein Language Models Expose Viral Immune Mimicry

Original source ↗

Abstract; Methods 2.1 Protein Datasets and 2.6 Model Performance; Table 1

Version: version of record
Retrieved: 2026-09-16T10:33:57.274Z

source checked

automated source review · 2026-09-16

Audit details

Inspected the cited primary abstract and descriptive computational-evaluation sections. Review covers the descriptive claims shown; no executable protocol was reconstructed. Original numerical records retain their prior transcription review.

Field: attributes.profile.facts.5.value

Source artifact SHA-256: 15250af2f75f70e2b6a3725d00bc7e276ed9d2d4a54f0ae9d6eaabf6be13e4a1

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Adaptation

Pretrained protein representations are evaluated through a study-specific classifier; a backbone name alone does not identify the full fitted pipeline.

Individual claims
Protein Language Models Expose Viral Immune Mimicry

Original source ↗

Abstract; Methods 2.1 Protein Datasets and 2.6 Model Performance; Table 1

Version: version of record
Retrieved: 2026-09-16T10:33:57.274Z

source checked

automated source review · 2026-09-16

Audit details

Inspected the cited primary abstract and descriptive computational-evaluation sections. Review covers the descriptive claims shown; no executable protocol was reconstructed. Original numerical records retain their prior transcription review.

Field: attributes.profile.facts.6.value

Source artifact SHA-256: 15250af2f75f70e2b6a3725d00bc7e276ed9d2d4a54f0ae9d6eaabf6be13e4a1

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Sources and history

Release 2026-09-17-d277315f7d76 · Record review: needs review

2 source records and release historyDownload this release
Technical metadata and extraction receipts

Stable ID: reported-task-53506fe386e4a1

areas
proteins-complexes
tasks
human-versus-viral protein classification
entity level
task
version
Not reported
task
human-versus-viral protein classification
scope note
Paper-specific evaluation task; protocol completeness requires further extraction.
comparison panels
id: part2-viral-immune-mimicry-2025-viruses-17-01199-t001-3a83e9bebb; title: Held-out human-versus-virus protein classification · Table 1; protocol id: paper-protocol-1352834b9391c1dacb; dataset id: reported-dataset-43f24c4dfb7351; metric: AUROC; unit: percent; direction: higher; result ids: paper-result-e2941183edbcf0f806; paper-result-088aaf71a44cdc5709; paper-result-a19e0533c00420b50c; paper-result-13ca62957ffa02eef1; paper-result-a541b9f3a77eec62e2; lit-b4-009; paper-result-b740d08b806206779f; paper-result-6536626a03a3ef62f0; source ids: part2-viral-immune-mimicry-2025; source locator: Table 1: AUC (%), Held-out human-versus-virus protein classification; context: Compare protein origin classifiers; T 5 embeddings plus linear/tree models vs ESM 2 fine-tuning and simple controls. UniRef90 deduplication; proteins longer than1,600 residues excluded; UniRef50 clusters assigned80% training and20% test with no cluster shared. Separate four-fold error-analysis experiment not assigned toTable1.; caveats: Origin classification does not establish immune mimicry.; Table footnote says all values are percent but log-loss values are dimensionless; log-loss units quarantined.; Narrative n-gram AUC91.9 differs from Table1 value91.5; preserve table value.; No interval assigned unless printed in source cell.; review: method: automated_source_review; date: 2026-09-17; id: part2-viral-immune-mimicry-2025-viruses-17-01199-t001-668a472b03; title: Held-out human-versus-virus protein classification · Table 1; protocol id: paper-protocol-1352834b9391c1dacb; dataset id: reported-dataset-43f24c4dfb7351; metric: Accur.; unit: percent; direction: higher; result ids: paper-result-91e7ac901444849ecd; paper-result-ca12e3e749c612d6d8; paper-result-ee7f4e8a864f523eaf; paper-result-544d2c1cd2bd46af9b; paper-result-32da435487ad9e4e7b; paper-result-7ffc4775000c61b1e4; paper-result-87b3b45c365edc42f3; paper-result-439ed50c8779070823; source ids: part2-viral-immune-mimicry-2025; source locator: Table 1: Accur., Held-out human-versus-virus protein classification; context: Compare protein origin classifiers; T 5 embeddings plus linear/tree models vs ESM 2 fine-tuning and simple controls. UniRef90 deduplication; proteins longer than1,600 residues excluded; UniRef50 clusters assigned80% training and20% test with no cluster shared. Separate four-fold error-analysis experiment not assigned toTable1.; caveats: Origin classification does not establish immune mimicry.; Table footnote says all values are percent but log-loss values are dimensionless; log-loss units quarantined.; Narrative n-gram AUC91.9 differs from Table1 value91.5; preserve table value.; No interval assigned unless printed in source cell.; review: method: automated_source_review; date: 2026-09-17; id: part2-viral-immune-mimicry-2025-viruses-17-01199-t001-e488608662; title: Held-out human-versus-virus protein classification · Table 1; protocol id: paper-protocol-1352834b9391c1dacb; dataset id: reported-dataset-43f24c4dfb7351; metric: Prec.; unit: percent; direction: higher; result ids: paper-result-a950e05fb8768ed4e8; paper-result-0c1a3836d475c5a506; paper-result-00f7697f28c706f26b; paper-result-1f9cd7bc0d3be4259f; paper-result-24e80b6223e7bd1109; paper-result-0c5e3578112cc7247a; paper-result-c583dfa48d139dadc8; paper-result-f67a4402bb1e2c96ba; source ids: part2-viral-immune-mimicry-2025; source locator: Table 1: Prec., Held-out human-versus-virus protein classification; context: Compare protein origin classifiers; T 5 embeddings plus linear/tree models vs ESM 2 fine-tuning and simple controls. UniRef90 deduplication; proteins longer than1,600 residues excluded; UniRef50 clusters assigned80% training and20% test with no cluster shared. Separate four-fold error-analysis experiment not assigned toTable1.; caveats: Origin classification does not establish immune mimicry.; Table footnote says all values are percent but log-loss values are dimensionless; log-loss units quarantined.; Narrative n-gram AUC91.9 differs from Table1 value91.5; preserve table value.; No interval assigned unless printed in source cell.; review: method: automated_source_review; date: 2026-09-17; id: part2-viral-immune-mimicry-2025-viruses-17-01199-t001-a4943fb2ad; title: Held-out human-versus-virus protein classification · Table 1; protocol id: paper-protocol-1352834b9391c1dacb; dataset id: reported-dataset-43f24c4dfb7351; metric: Recall; unit: percent; direction: higher; result ids: paper-result-b36e1a943ebc5dfb9a; paper-result-ccac699408a1229570; paper-result-4f2c34cb0e09f0167a; paper-result-4014edf9066fc226f1; paper-result-9d9d31080b47613e46; paper-result-a42799c550a1ae61e7; paper-result-67bb8fdb6692f4336f; paper-result-2aafdf0e911ae9cff5; source ids: part2-viral-immune-mimicry-2025; source locator: Table 1: Recall, Held-out human-versus-virus protein classification; context: Compare protein origin classifiers; T 5 embeddings plus linear/tree models vs ESM 2 fine-tuning and simple controls. UniRef90 deduplication; proteins longer than1,600 residues excluded; UniRef50 clusters assigned80% training and20% test with no cluster shared. Separate four-fold error-analysis experiment not assigned toTable1.; caveats: Origin classification does not establish immune mimicry.; Table footnote says all values are percent but log-loss values are dimensionless; log-loss units quarantined.; Narrative n-gram AUC91.9 differs from Table1 value91.5; preserve table value.; No interval assigned unless printed in source cell.; review: method: automated_source_review; date: 2026-09-17
benchmark research
review date: 2026-09-17; status: complete_comparison_tables_extracted_pending_publication_review; primary sources: part2-viral-immune-mimicry-2025; inspected locators: Table1; held-out protein-origin classification; separate error-analysis CV; searched queries: Protein Language Models Expose Viral Immune Mimicry primary paper benchmark results; gaps: Independent batch review before import; preserve existing observation identities.; claim scope: Primary-source discovery and table/protocol screening; source checked is not independently reproduced. Raw acquisitions not automatically numerical publication approval.
historical missing metadata
protocol version: not_reported_in_legacy_extract; split: not_reported_in_legacy_extract
metadata review scope
historical_missing_metadata preserves the original discovery state. Current descriptive evidence and missingness are recorded in profile.facts; numerical-result review is separate.
legacy kinds
benchmark
entity classification
review date: 2026-09-17; rationale: This source-scoped record identifies the biological prediction task and holds its paper context. Preserve the existing task identity; exact split, model adaptation and scoring remain in linked evaluations or separate protocol records.; source ids: viral-immune-mimicry-2025; source locator: Abstract; Methods 2.1 Protein Datasets and 2.6 Model Performance; Table 1; ambiguities: A paper- or suite-specific task may constrain some inputs or metrics; that alone does not make it interchangeable with a complete versioned protocol. No protocol equivalence is inferred.; Some legacy profile Entity type facts use the generic phrase computational evaluation protocol. That boilerplate is not sufficient to establish a single fixed protocol identity or to merge this task with another protocol record.
Related records

Suggest a correction