rewire.it
Configuration

C2S (GPT-2 Large)

Cell2Sentence adapts GPT-2 to single-cell transcriptomics by writing cells as expression-ranked gene-name sequences.

SourcesCell2Sentence: Teaching Large Language Models the Language of Biology · Methods/Data transformation (paragraph 8); Introduction (paragraph 2)

1 evaluation · 1 metric row

How it worksEvaluated procedure (conceptual)
Evaluated procedure (conceptual)1. Single-cell gene-expression profiles converted to ranked gene names. Then: 2. C2S (GPT-2 Large). Then: 3. Cell-type annotations or generated ranked gene sequencesEvaluated procedure (conceptual)1. Single-cell gene-expression profiles converted to ranked gene names. Then: 2. C2S (GPT-2 Large). Then: 3. Cell-type annotations or generated ranked gene sequencesEvaluated procedure (conceptual)1. Single-cell gene-expression profiles converted to ranked gene names. Then: 2. C2S (GPT-2 Large). Then: 3. Cell-type annotations or generated ranked gene sequences

Conceptual input–method–output guide. Check the procedure text and linked evaluation for fitted components, additional inputs and exact settings.

SourcesCell2Sentence: Teaching Large Language Models the Language of Biology · Methods (paragraph 1); Experiments/Fine-Tuning Datasets (paragraph 3)

At a glance

limited source coverage · Automated source review, 2026-09-16. All specifications and missing details

Evaluations and results

Release 2026-09-17-d277315f7d76 · 1 evaluation · 1 metric row. Different protocols are not a single leaderboard.

Results grouped by the exact reported evaluation
Metric and findingCoverage and uncertaintyEvidence
C2S (GPT-2 Large): Combinatorial cell-label classification

Partial-credit labels including cell type, perturbation, and dose.

Author-reported evaluation · Evaluation metadata: needs review

0.631 Partial-label accuracy

Unit: unitless · Direction: unknown

Uncertainty: ± 0.0031

Scored: Not reported · Eligible: Not reported

source checkedCell2Sentence: Teaching Large Language Models the Language of Biology · Table 3, Partial label / C2S (GPT-2 Large) row, L1000 Acc column

Source checking is not independent reproduction.

How it works

How the evaluated method works

Genes are ordered by decreasing transcript abundance to create cell sentences. GPT-2 Large is fine-tuned on this representation for cell-type-conditioned generation and cell-type annotation.

SourcesCell2Sentence: Teaching Large Language Models the Language of Biology · Methods (paragraph 1); Experiments/Fine-Tuning Datasets (paragraph 3)
What was evaluated

The linked evaluation record identifies C2S (GPT-2 Large): Combinatorial cell-label classification. Its dataset, split, adaptation and evidence origin remain attached to the reported results.

SourcesCell2Sentence: Teaching Large Language Models the Language of Biology · The named evaluation’s methods and comparison table; exact preserved evaluation IDs: evaluation-lit-027

Strengths and limitations

Profile review details

Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.

Stable record: reported-model-ab02228f50a37c

Specifications

Inputs, training, access and other details

Explanatory profile: limited source coverage · Automated source review, 2026-09-16. Review applies to the cited claims; unresolved fields are listed below. Numerical results retain their own review status.

Inputs, outputs and configuration
PropertyDescription and evidence
Model typeStudy-specific predictive method; this record is the paper-specific evaluated configuration.
SourcesCell2Sentence: Teaching Large Language Models the Language of Biology · Methods (paragraph 1); Experiments/Fine-Tuning Datasets (paragraph 3)
Architecture / procedureGenes are ordered by decreasing transcript abundance to create cell sentences. GPT-2 Large is fine-tuned on this representation for cell-type-conditioned generation and cell-type annotation.
SourcesCell2Sentence: Teaching Large Language Models the Language of Biology · Methods (paragraph 1); Experiments/Fine-Tuning Datasets (paragraph 3)
Biological inputsSingle-cell gene-expression profiles converted to ranked gene names
SourcesCell2Sentence: Teaching Large Language Models the Language of Biology · Methods/Data transformation (paragraph 8); Methods/Data transformation (paragraph 3)
OutputsCell-type annotations or generated ranked gene sequences
SourcesCell2Sentence: Teaching Large Language Models the Language of Biology · Methods/Data transformation (paragraph 7); Inference Details (paragraph 2)
Parameters774,030,080 parameters for the model checkpoint used in the Cell2Sentence L1000 experiment; this is not a claim about every checkpoint in the family.
SourcesCell2Sentence: Teaching Large Language Models the Language of Biology · Table 5, Comparison of required compute on the L1000 dataset; # Parameters column and caption
Known versions / configurationGPT-2 Large
SourcesCell2Sentence: Teaching Large Language Models the Language of Biology · Experiments/Experiment 3: abstract summary generation/Objective: (paragraph 1); Experiments/Fine-Tuning Datasets (paragraph 2)
Training data / fittingTask-specific single-cell expression datasets described in the preprint; gene ranking replaces direct numeric expression input.
SourcesCell2Sentence: Teaching Large Language Models the Language of Biology · Methods/Data transformation (paragraph 8); Experimental Details/Evaluation Datasets (paragraph 1)
Context limitsGPT-2 configurations use 1,024 tokens; the separately evaluated Pythia-160m configuration uses 9,200 tokens.
SourcesCell2Sentence: Teaching Large Language Models the Language of Biology · Training Details (paragraph 1); Experiments (paragraph 1)
AccessOfficial study implementation and usage documentation: https://github.com/vandijklab/cell2sentence/blob/a6efaf079f98491d4723ced44b929936b94368aa/README.md. This pinned documentation revision is not automatically the evaluated weight revision.
Sourcesvandijklab/cell2sentence README.md · README.md; installation, model download and usage instructions
Code licenceApache 2.0 (study repository code at the cited revision; this does not establish every dependency or historical checkpoint licence).
Sourcesvandijklab/cell2sentence LICENSE · LICENSE; complete licence text
Weights licenceThe inspected model-access documentation does not explicitly identify terms for this exact evaluated checkpoint or fitted head; repository code terms are shown separately. · Not reported in inspected sources
Sourcesvandijklab/cell2sentence README.md · README.md; checkpoint/access documentation and licence scope

Evidence table

Inspect claims, sources and review details

Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.

One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.

19 evidence rows matching the loaded filters

Claims, original sources and review scope · Release 2026-09-17-d277315f7d76
Property and statementOriginal source and locationReview and provenance
Diagram caption

Conceptual input–method–output guide. Check the procedure text and linked evaluation for fitted components, additional inputs and exact settings.

Individual claims
Cell2Sentence: Teaching Large Language Models the Language of Biology

Original source ↗

Methods (paragraph 1); Experiments/Fine-Tuning Datasets (paragraph 3)

Version: preprint archived 2024-10-29
Retrieved: 2026-09-16T10:41:16.533640+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.

Field: attributes.profile.diagram.caption

Source artifact SHA-256: e088727d6e04857fccb7033a9b074e1850f775e86e7d2e99e603dde09558ab02

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Diagram steps

["Single-cell gene-expression profiles converted to ranked gene names","C2S (GPT-2 Large)","Cell-type annotations or generated ranked gene sequences"]

Individual claims
Cell2Sentence: Teaching Large Language Models the Language of Biology

Original source ↗

Methods (paragraph 1); Experiments/Fine-Tuning Datasets (paragraph 3)

Version: preprint archived 2024-10-29
Retrieved: 2026-09-16T10:41:16.533640+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.

Field: attributes.profile.diagram.steps

Source artifact SHA-256: e088727d6e04857fccb7033a9b074e1850f775e86e7d2e99e603dde09558ab02

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Diagram title

Evaluated procedure (conceptual)

Individual claims
Cell2Sentence: Teaching Large Language Models the Language of Biology

Original source ↗

Methods (paragraph 1); Experiments/Fine-Tuning Datasets (paragraph 3)

Version: preprint archived 2024-10-29
Retrieved: 2026-09-16T10:41:16.533640+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.

Field: attributes.profile.diagram.title

Source artifact SHA-256: e088727d6e04857fccb7033a9b074e1850f775e86e7d2e99e603dde09558ab02

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Model type

Study-specific predictive method; this record is the paper-specific evaluated configuration.

Individual claims
Cell2Sentence: Teaching Large Language Models the Language of Biology

Original source ↗

Methods (paragraph 1); Experiments/Fine-Tuning Datasets (paragraph 3)

Version: preprint archived 2024-10-29
Retrieved: 2026-09-16T10:41:16.533640+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.

Field: attributes.profile.facts.0.value

Source artifact SHA-256: e088727d6e04857fccb7033a9b074e1850f775e86e7d2e99e603dde09558ab02

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Architecture / procedure

Genes are ordered by decreasing transcript abundance to create cell sentences. GPT-2 Large is fine-tuned on this representation for cell-type-conditioned generation and cell-type annotation.

Individual claims
Cell2Sentence: Teaching Large Language Models the Language of Biology

Original source ↗

Methods (paragraph 1); Experiments/Fine-Tuning Datasets (paragraph 3)

Version: preprint archived 2024-10-29
Retrieved: 2026-09-16T10:41:16.533640+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.

Field: attributes.profile.facts.1.value

Source artifact SHA-256: e088727d6e04857fccb7033a9b074e1850f775e86e7d2e99e603dde09558ab02

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Weights licence

The inspected model-access documentation does not explicitly identify terms for this exact evaluated checkpoint or fitted head; repository code terms are shown separately.

Individual claims
vandijklab/cell2sentence README.md

Original source ↗

README.md; checkpoint/access documentation and licence scope

Version: a6efaf079f98491d4723ced44b929936b94368aa
Retrieved: 2026-09-16T20:42:56.481821+00:00

unreported

automated source review · 2026-09-16

Audit details

Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.

Field: attributes.profile.facts.10.value

Source artifact SHA-256: 60b183e5311a46eff139d728d860875874ff255242bcd7a237789b893ffc262c

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Biological inputs

Single-cell gene-expression profiles converted to ranked gene names

Individual claims
Cell2Sentence: Teaching Large Language Models the Language of Biology

Original source ↗

Methods/Data transformation (paragraph 8); Methods/Data transformation (paragraph 3)

Version: preprint archived 2024-10-29
Retrieved: 2026-09-16T10:41:16.533640+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.

Field: attributes.profile.facts.2.value

Source artifact SHA-256: e088727d6e04857fccb7033a9b074e1850f775e86e7d2e99e603dde09558ab02

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Outputs

Cell-type annotations or generated ranked gene sequences

Individual claims
Cell2Sentence: Teaching Large Language Models the Language of Biology

Original source ↗

Methods/Data transformation (paragraph 7); Inference Details (paragraph 2)

Version: preprint archived 2024-10-29
Retrieved: 2026-09-16T10:41:16.533640+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.

Field: attributes.profile.facts.3.value

Source artifact SHA-256: e088727d6e04857fccb7033a9b074e1850f775e86e7d2e99e603dde09558ab02

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Parameters

774,030,080 parameters for the model checkpoint used in the Cell2Sentence L1000 experiment; this is not a claim about every checkpoint in the family.

Individual claims
Cell2Sentence: Teaching Large Language Models the Language of Biology

Original source ↗

Table 5, Comparison of required compute on the L1000 dataset; # Parameters column and caption

Version: preprint archived 2024-10-29
Retrieved: 2026-09-16T10:41:16.533640+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.

Field: attributes.profile.facts.4.value

Source artifact SHA-256: e088727d6e04857fccb7033a9b074e1850f775e86e7d2e99e603dde09558ab02

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Known versions / configuration

GPT-2 Large

Individual claims
Cell2Sentence: Teaching Large Language Models the Language of Biology

Original source ↗

Experiments/Experiment 3: abstract summary generation/Objective: (paragraph 1); Experiments/Fine-Tuning Datasets (paragraph 2)

Version: preprint archived 2024-10-29
Retrieved: 2026-09-16T10:41:16.533640+00:00

source checked

automated source review · 2026-09-16

Audit details

Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.

Field: attributes.profile.facts.5.value

Source artifact SHA-256: e088727d6e04857fccb7033a9b074e1850f775e86e7d2e99e603dde09558ab02

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Sources and history

Release 2026-09-17-d277315f7d76 · Record review: needs review

3 source records and release historyDownload this release
Technical metadata and extraction receipts

Stable ID: reported-model-ab02228f50a37c

areas
cells-tissues
entity level
method
version
GPT-2 Large
reported name
C2S (GPT-2 Large)
historical missing metadata
checkpoint revision: not_reported_in_legacy_extract; training data: not_reported_in_legacy_extract; licence: not_reported_in_legacy_extract
metadata review scope
historical_missing_metadata preserves the original discovery state. Current descriptive evidence and missingness are recorded in profile.facts; numerical-result review is separate.
legacy kinds
model
entity classification
review date: 2026-09-17; rationale: This source-scoped entry preserves the method/configuration actually named in an evaluation. It is neither a global family identity nor proof of an immutable checkpoint; the linked evaluation retains adaptation, fitting and scoring details.; source ids: cell2sentence-2024; source locator: Methods (paragraph 1); Experiments/Fine-Tuning Datasets (paragraph 3) | Methods/Data transformation (paragraph 8); Introduction (paragraph 2); ambiguities: Configuration means the source-labelled evaluated identity. It does not establish missing checkpoint hashes, default settings or equivalence to same-named records in other papers.
Related records

Suggest a correction