rewire.it
Configuration

ESM2 650M

This ESM-2 650M configuration is fine-tuned with LoRA to distinguish human from viral protein sequences.

SourcesProtein Language Models Expose Viral Immune Mimicry · Section 2.3 Human-Virus Model Training and Implementation, paragraphs 1–3; Section 2.4 Finding and Analyzing Model Mistakes, paragraphs 1–2; Table 1, ESM2 650M row

1 evaluation · 4 metric rows

How it worksEvaluated procedure (conceptual)
Evaluated procedure (conceptual)1. Protein sequences. Then: 2. ESM-2 650M fine-tuned with LoRA. Then: 3. Human-versus-viral classification scoresEvaluated procedure (conceptual)1. Protein sequences. Then: 2. ESM-2 650M fine-tuned with LoRA. Then: 3. Human-versus-viral classification scoresEvaluated procedure (conceptual)1. Protein sequences. Then: 2. ESM-2 650M fine-tuned with LoRA. Then: 3. Human-versus-viral classification scores

Conceptual input–method–output guide. Check the procedure text and linked evaluation for fitted components, additional inputs and exact settings.

SourcesProtein Language Models Expose Viral Immune Mimicry · Section 2.3 Human-Virus Model Training and Implementation, paragraphs 1–3; Section 2.4 Finding and Analyzing Model Mistakes, paragraphs 1–2; Table 1, ESM2 650M row

At a glance

Model type

Protein sequence transformer; this record is the paper-specific evaluated configuration.

Sourcesfacebookresearch/esm README.md · README.md model description

Biological inputs

Protein sequences represented by pretrained embeddings

SourcesProtein Language Models Expose Viral Immune Mimicry · 2. Materials and Methods/2.2. Pretrained Deep Language Models (ESM, T5) (paragraph 1); 3. Results/3.4. Latent Structure Embeddings Clustering (paragraph 1)

limited source coverage · Automated source review, 2026-09-17. All specifications and missing details

Evaluations and results

Release 2026-09-17-d277315f7d76 · 1 evaluation · 4 metric rows. Different protocols are not a single leaderboard.

Results grouped by the exact reported evaluation
Metric and findingCoverage and uncertaintyEvidence
ESM2 650M: human-versus-viral protein classification

Compare protein origin classifiers; T 5 embeddings plus linear/tree models vs ESM 2 fine-tuning and simple controls. UniRef90 deduplication; proteins longer than1,600 residues excluded; UniRef50 clusters assigned80% training and20% test with no cluster shared. Separate four-fold error-analysis experiment not assigned toTable1.

Author-reported evaluation · Evaluation metadata: needs review

99.67% AUROC

Unit: percent · Direction: higher

Uncertainty: not reported in legacy extract

Scored: Not reported · Eligible: Not reported

source checkedProtein Language Models Expose Viral Immune Mimicry; Protein Language Models Expose Viral Immune Mimicry · Table 1, ESM2 650M row, AUC (%) column

Source checking is not independent reproduction.

96.85% Prec.

Unit: percent · Direction: higher

Uncertainty: unreported

Scored: Not reported · Eligible: Not reported

source checkedProtein Language Models Expose Viral Immune Mimicry · Table 1, row ESM2 650M, column Prec.; XML row7 column4

Source checking is not independent reproduction.

97.86% Accur.

Unit: percent · Direction: higher

Uncertainty: unreported

Scored: Not reported · Eligible: Not reported

source checkedProtein Language Models Expose Viral Immune Mimicry · Table 1, row ESM2 650M, column Accur.; XML row7 column3

Source checking is not independent reproduction.

96.68% Recall

Unit: percent · Direction: higher

Uncertainty: unreported

Scored: Not reported · Eligible: Not reported

source checkedProtein Language Models Expose Viral Immune Mimicry · Table 1, row ESM2 650M, column Recall; XML row7 column5

Source checking is not independent reproduction.

How it works

How the evaluated method works

The ESM-2 650M model is fine-tuned with LoRA on protein sequences to classify human versus viral origin. The separate embedding-based linear/tree comparisons use T5; the paper’s immune-mimicry error analysis uses Linear-T5, not this ESM-2 configuration.

SourcesProtein Language Models Expose Viral Immune Mimicry · Section 2.3 Human-Virus Model Training and Implementation, paragraphs 1–3; Section 2.4 Finding and Analyzing Model Mistakes, paragraphs 1–2; Table 1, ESM2 650M row
Underlying method and version boundaries

ESM-2 is a transformer protein language-model family. The official repository exposes residue embeddings, sequence-level pooling and models at several sizes; the study configuration determines which of these is evaluated.

Sourcesfacebookresearch/esm README.md · README.md; introduction, model description, pretrained-model and usage sections at pinned revision
What was evaluated

The linked evaluation record identifies ESM2 650M: human-versus-viral protein classification. Its dataset, split, adaptation and evidence origin remain attached to the reported results.

SourcesProtein Language Models Expose Viral Immune Mimicry · The named evaluation’s methods and comparison table; exact preserved evaluation IDs: evaluation-lit-b4-009

Strengths and limitations

Strengths and considerations

  • Cluster-based splitting reduces close-sequence overlap between train and test.
    SourcesProtein Language Models Expose Viral Immune Mimicry · 2. Materials and Methods/2.1. Protein Datasets (paragraph 2); 2. Materials and Methods/2.4. Finding and Analyzing Model Mistakes (paragraph 1)

Limitations and conditions

Profile review details

Previous profile source review retained. On 2026-09-17, an independent agent and root agent checked Methods 2.3–2.4 and corrected the ESM-2 LoRA mechanism, diagram and parameter wording. This correction does not constitute human review or benchmark reproduction.

Stable record: reported-model-4c73500c39e9d0

Specifications

Inputs, training, access and other details

Explanatory profile: limited source coverage · Automated source review, 2026-09-17. Review applies to the cited claims; unresolved fields are listed below. Numerical results retain their own review status.

Inputs, outputs and configuration
PropertyDescription and evidence
Model typeProtein sequence transformer; this record is the paper-specific evaluated configuration.
Sourcesfacebookresearch/esm README.md · README.md model description
Architecture / procedureThe ESM-2 650M model is fine-tuned with LoRA on protein sequences to classify human versus viral origin. The separate embedding-based linear/tree comparisons use T5; the paper’s immune-mimicry error analysis uses Linear-T5, not this ESM-2 configuration.
SourcesProtein Language Models Expose Viral Immune Mimicry · Section 2.3 Human-Virus Model Training and Implementation, paragraphs 1–3; Section 2.4 Finding and Analyzing Model Mistakes, paragraphs 1–2; Table 1, ESM2 650M row
Biological inputsProtein sequences represented by pretrained embeddings
SourcesProtein Language Models Expose Viral Immune Mimicry · 2. Materials and Methods/2.2. Pretrained Deep Language Models (ESM, T5) (paragraph 1); 3. Results/3.4. Latent Structure Embeddings Clustering (paragraph 1)
OutputsViral-versus-human classification scores
SourcesProtein Language Models Expose Viral Immune Mimicry · 3. Results/3.5. Immunogenicity Analysis (paragraph 1); 3. Results (paragraph 1)
ParametersThe ESM-2 backbone is labelled 650M parameters. The paper uses LoRA with rank 8 and scaling factor 8; it does not report an aggregate trained-configuration parameter count.
SourcesProtein Language Models Expose Viral Immune Mimicry · Section 2.3 Human-Virus Model Training and Implementation, paragraphs 1–3; Section 2.4 Finding and Analyzing Model Mistakes, paragraphs 1–2; Table 1, ESM2 650M row
Known versions / configurationESM2 650M is the comparison-table label; that label does not specify an immutable weight revision. · Not reported in inspected sources
SourcesProtein Language Models Expose Viral Immune Mimicry · Model identification in the comparison table and corresponding Methods; immutable checkpoint revision is not supplied by the table label.
Training data / fittingUniRef90 duplicate removal and a UniRef50-clustered 80:20 split; sequences above 1,600 residues are excluded.
SourcesProtein Language Models Expose Viral Immune Mimicry · 2. Materials and Methods/2.1. Protein Datasets (paragraph 2); 2. Materials and Methods/2.1. Protein Datasets (paragraph 1)
Context limitsThe study excludes proteins longer than 1,600 residues.
SourcesProtein Language Models Expose Viral Immune Mimicry · 2. Materials and Methods/2.1. Protein Datasets (paragraph 2); 3. Results/3.3. Virus Errors Analysis (paragraph 2)
AccessOfficial upstream implementation and usage documentation: https://github.com/facebookresearch/esm/blob/2b369911bb5b4b0dda914521b9475cad1656b2ac/README.md. This pinned documentation revision is not automatically the evaluated weight revision.
Sourcesfacebookresearch/esm README.md · README.md; installation, model download and usage instructions
Code licenceMIT (upstream repository code at the cited revision; this does not establish every dependency or historical checkpoint licence).
Sourcesfacebookresearch/esm LICENSE · LICENSE; complete licence text
Weights licenceThe inspected model-access documentation does not explicitly identify terms for this exact evaluated checkpoint or fitted head; repository code terms are shown separately. · Not reported in inspected sources
Sourcesfacebookresearch/esm README.md · README.md; checkpoint/access documentation and licence scope

Evidence table

Inspect claims, sources and review details

Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.

One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.

20 evidence rows matching the loaded filters

Claims, original sources and review scope · Release 2026-09-17-d277315f7d76
Property and statementOriginal source and locationReview and provenance
Diagram caption

Conceptual input–method–output guide. Check the procedure text and linked evaluation for fitted components, additional inputs and exact settings.

Individual claims
Protein Language Models Expose Viral Immune Mimicry

Original source ↗

Section 2.3 Human-Virus Model Training and Implementation, paragraphs 1–3; Section 2.4 Finding and Analyzing Model Mistakes, paragraphs 1–2; Table 1, ESM2 650M row

Version: version of record
Retrieved: 2026-09-16T10:33:57.274Z

source checked

automated source review · 2026-09-17

Audit details

Previous profile source review retained. On 2026-09-17, an independent agent and root agent checked Methods 2.3–2.4 and corrected the ESM-2 LoRA mechanism, diagram and parameter wording. This correction does not constitute human review or benchmark reproduction.

Field: attributes.profile.diagram.caption

Source artifact SHA-256: 15250af2f75f70e2b6a3725d00bc7e276ed9d2d4a54f0ae9d6eaabf6be13e4a1

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Diagram steps

["Protein sequences","ESM-2 650M fine-tuned with LoRA","Human-versus-viral classification scores"]

Individual claims
Protein Language Models Expose Viral Immune Mimicry

Original source ↗

Section 2.3 Human-Virus Model Training and Implementation, paragraphs 1–3; Section 2.4 Finding and Analyzing Model Mistakes, paragraphs 1–2; Table 1, ESM2 650M row

Version: version of record
Retrieved: 2026-09-16T10:33:57.274Z

source checked

automated source review · 2026-09-17

Audit details

Previous profile source review retained. On 2026-09-17, an independent agent and root agent checked Methods 2.3–2.4 and corrected the ESM-2 LoRA mechanism, diagram and parameter wording. This correction does not constitute human review or benchmark reproduction.

Field: attributes.profile.diagram.steps

Source artifact SHA-256: 15250af2f75f70e2b6a3725d00bc7e276ed9d2d4a54f0ae9d6eaabf6be13e4a1

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Diagram title

Evaluated procedure (conceptual)

Individual claims
Protein Language Models Expose Viral Immune Mimicry

Original source ↗

Section 2.3 Human-Virus Model Training and Implementation, paragraphs 1–3; Section 2.4 Finding and Analyzing Model Mistakes, paragraphs 1–2; Table 1, ESM2 650M row

Version: version of record
Retrieved: 2026-09-16T10:33:57.274Z

source checked

automated source review · 2026-09-17

Audit details

Previous profile source review retained. On 2026-09-17, an independent agent and root agent checked Methods 2.3–2.4 and corrected the ESM-2 LoRA mechanism, diagram and parameter wording. This correction does not constitute human review or benchmark reproduction.

Field: attributes.profile.diagram.title

Source artifact SHA-256: 15250af2f75f70e2b6a3725d00bc7e276ed9d2d4a54f0ae9d6eaabf6be13e4a1

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Model type

Protein sequence transformer; this record is the paper-specific evaluated configuration.

Individual claims
facebookresearch/esm README.md

Original source ↗

README.md model description

Version: 2b369911bb5b4b0dda914521b9475cad1656b2ac
Retrieved: 2026-09-16T20:00:00.816433+00:00

source checked

automated source review · 2026-09-17

Audit details

Previous profile source review retained. On 2026-09-17, an independent agent and root agent checked Methods 2.3–2.4 and corrected the ESM-2 LoRA mechanism, diagram and parameter wording. This correction does not constitute human review or benchmark reproduction.

Field: attributes.profile.facts.0.value

Source artifact SHA-256: 8b273c21a322fc9473d1b68d0dd40c8166ab2f89e4a190aa26ca87251b97cba9

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Architecture / procedure

The ESM-2 650M model is fine-tuned with LoRA on protein sequences to classify human versus viral origin. The separate embedding-based linear/tree comparisons use T5; the paper’s immune-mimicry error analysis uses Linear-T5, not this ESM-2 configuration.

Individual claims
Protein Language Models Expose Viral Immune Mimicry

Original source ↗

Section 2.3 Human-Virus Model Training and Implementation, paragraphs 1–3; Section 2.4 Finding and Analyzing Model Mistakes, paragraphs 1–2; Table 1, ESM2 650M row

Version: version of record
Retrieved: 2026-09-16T10:33:57.274Z

source checked

automated source review · 2026-09-17

Audit details

Previous profile source review retained. On 2026-09-17, an independent agent and root agent checked Methods 2.3–2.4 and corrected the ESM-2 LoRA mechanism, diagram and parameter wording. This correction does not constitute human review or benchmark reproduction.

Field: attributes.profile.facts.1.value

Source artifact SHA-256: 15250af2f75f70e2b6a3725d00bc7e276ed9d2d4a54f0ae9d6eaabf6be13e4a1

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Weights licence

The inspected model-access documentation does not explicitly identify terms for this exact evaluated checkpoint or fitted head; repository code terms are shown separately.

Individual claims
facebookresearch/esm README.md

Original source ↗

README.md; checkpoint/access documentation and licence scope

Version: 2b369911bb5b4b0dda914521b9475cad1656b2ac
Retrieved: 2026-09-16T20:00:00.816433+00:00

unreported

automated source review · 2026-09-17

Audit details

Previous profile source review retained. On 2026-09-17, an independent agent and root agent checked Methods 2.3–2.4 and corrected the ESM-2 LoRA mechanism, diagram and parameter wording. This correction does not constitute human review or benchmark reproduction.

Field: attributes.profile.facts.10.value

Source artifact SHA-256: 8b273c21a322fc9473d1b68d0dd40c8166ab2f89e4a190aa26ca87251b97cba9

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Biological inputs

Protein sequences represented by pretrained embeddings

Individual claims
Protein Language Models Expose Viral Immune Mimicry

Original source ↗

2. Materials and Methods/2.2. Pretrained Deep Language Models (ESM, T5) (paragraph 1); 3. Results/3.4. Latent Structure Embeddings Clustering (paragraph 1)

Version: version of record
Retrieved: 2026-09-16T10:33:57.274Z

source checked

automated source review · 2026-09-17

Audit details

Previous profile source review retained. On 2026-09-17, an independent agent and root agent checked Methods 2.3–2.4 and corrected the ESM-2 LoRA mechanism, diagram and parameter wording. This correction does not constitute human review or benchmark reproduction.

Field: attributes.profile.facts.2.value

Source artifact SHA-256: 15250af2f75f70e2b6a3725d00bc7e276ed9d2d4a54f0ae9d6eaabf6be13e4a1

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Outputs

Viral-versus-human classification scores

Individual claims
Protein Language Models Expose Viral Immune Mimicry

Original source ↗

3. Results/3.5. Immunogenicity Analysis (paragraph 1); 3. Results (paragraph 1)

Version: version of record
Retrieved: 2026-09-16T10:33:57.274Z

source checked

automated source review · 2026-09-17

Audit details

Previous profile source review retained. On 2026-09-17, an independent agent and root agent checked Methods 2.3–2.4 and corrected the ESM-2 LoRA mechanism, diagram and parameter wording. This correction does not constitute human review or benchmark reproduction.

Field: attributes.profile.facts.3.value

Source artifact SHA-256: 15250af2f75f70e2b6a3725d00bc7e276ed9d2d4a54f0ae9d6eaabf6be13e4a1

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Parameters

The ESM-2 backbone is labelled 650M parameters. The paper uses LoRA with rank 8 and scaling factor 8; it does not report an aggregate trained-configuration parameter count.

Individual claims
Protein Language Models Expose Viral Immune Mimicry

Original source ↗

Section 2.3 Human-Virus Model Training and Implementation, paragraphs 1–3; Section 2.4 Finding and Analyzing Model Mistakes, paragraphs 1–2; Table 1, ESM2 650M row

Version: version of record
Retrieved: 2026-09-16T10:33:57.274Z

source checked

automated source review · 2026-09-17

Audit details

Previous profile source review retained. On 2026-09-17, an independent agent and root agent checked Methods 2.3–2.4 and corrected the ESM-2 LoRA mechanism, diagram and parameter wording. This correction does not constitute human review or benchmark reproduction.

Field: attributes.profile.facts.4.value

Source artifact SHA-256: 15250af2f75f70e2b6a3725d00bc7e276ed9d2d4a54f0ae9d6eaabf6be13e4a1

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Known versions / configuration

ESM2 650M is the comparison-table label; that label does not specify an immutable weight revision.

Individual claims
Protein Language Models Expose Viral Immune Mimicry

Original source ↗

Model identification in the comparison table and corresponding Methods; immutable checkpoint revision is not supplied by the table label.

Version: version of record
Retrieved: 2026-09-16T10:33:57.274Z

unreported

automated source review · 2026-09-17

Audit details

Previous profile source review retained. On 2026-09-17, an independent agent and root agent checked Methods 2.3–2.4 and corrected the ESM-2 LoRA mechanism, diagram and parameter wording. This correction does not constitute human review or benchmark reproduction.

Field: attributes.profile.facts.5.value

Source artifact SHA-256: 15250af2f75f70e2b6a3725d00bc7e276ed9d2d4a54f0ae9d6eaabf6be13e4a1

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Sources and history

Release 2026-09-17-d277315f7d76 · Record review: needs review

3 source records and release historyDownload this release
Technical metadata and extraction receipts

Stable ID: reported-model-4c73500c39e9d0

areas
proteins-complexes
entity level
method
version
Not reported
reported name
ESM2 650M
historical missing metadata
version: not_reported_in_legacy_extract; checkpoint revision: not_reported_in_legacy_extract; training data: not_reported_in_legacy_extract; licence: not_reported_in_legacy_extract
metadata review scope
historical_missing_metadata preserves the original discovery state. Current descriptive evidence and missingness are recorded in profile.facts; numerical-result review is separate.
legacy kinds
model
entity classification
review date: 2026-09-17; rationale: This source-scoped entry preserves the method/configuration actually named in an evaluation. It is neither a global family identity nor proof of an immutable checkpoint; the linked evaluation retains adaptation, fitting and scoring details.; source ids: viral-immune-mimicry-2025; evidence-reported-base-esm-readme-md; source locator: 3. Results/3.2. Error Analysis Models Insights (paragraph 1); Abstract (paragraph 1) | README.md model description | 3. Results (paragraph 1); 3. Results/3.4. Latent Structure Embeddings Clustering (paragraph 1); ambiguities: Configuration means the source-labelled evaluated identity. It does not establish missing checkpoint hashes, default settings or equivalence to same-named records in other papers.
profile correction
review date: 2026-09-17; review method: Independent agent inspection of cached complete primary article Methods2.3–2.4 and Table1, cross-checked against preserved evaluation-lit-b4-009; rationale: The current mechanism and diagram conflate the ESM2 Table 1 fine-tuning configuration with T5 static-embedding classifiers and the separate Linear-T5 error analysis. Sections2.3–2.4 explicitly distinguish these. Preserve configuration kind and all numerical results; correct only explanatory content.; source ids: viral-immune-mimicry-2025; source locator: Section 2.3 Human-Virus Model Training and Implementation, paragraphs 1–3; Section 2.4 Finding and Analyzing Model Mistakes, paragraphs 1–2; Table 1, ESM2 650M row
Related records

Suggest a correction