Model type
Study-specific predictive method; this record is the paper-specific evaluated configuration.
Cell2Sentence adapts GPT-2 to single-cell transcriptomics by writing cells as expression-ranked gene-name sequences.
Conceptual input–method–output guide. Check the procedure text and linked evaluation for fitted components, additional inputs and exact settings.
Study-specific predictive method; this record is the paper-specific evaluated configuration.
Single-cell gene-expression profiles converted to ranked gene names
Cell-type annotations or generated ranked gene sequences
limited source coverage · Automated source review, 2026-09-16. All specifications and missing details
Release 2026-09-17-d277315f7d76 · 1 evaluation · 1 metric row. Different protocols are not a single leaderboard.
| Metric and finding | Coverage and uncertainty | Evidence |
|---|---|---|
| C2S (GPT-2 Large): Combinatorial cell-label classification Partial-credit labels including cell type, perturbation, and dose. Author-reported evaluation · Evaluation metadata: needs review | ||
| 0.631 Partial-label accuracy Unit: unitless · Direction: unknown | Uncertainty: ± 0.0031 Scored: Not reported · Eligible: Not reported | source checkedCell2Sentence: Teaching Large Language Models the Language of Biology · Table 3, Partial label / C2S (GPT-2 Large) row, L1000 Acc column Source checking is not independent reproduction. |
Genes are ordered by decreasing transcript abundance to create cell sentences. GPT-2 Large is fine-tuned on this representation for cell-type-conditioned generation and cell-type annotation.
The linked evaluation record identifies C2S (GPT-2 Large): Combinatorial cell-label classification. Its dataset, split, adaptation and evidence origin remain attached to the reported results.
Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.
Stable record: reported-model-ab02228f50a37cExplanatory profile: limited source coverage · Automated source review, 2026-09-16. Review applies to the cited claims; unresolved fields are listed below. Numerical results retain their own review status.
| Property | Description and evidence |
|---|---|
| Model type | Study-specific predictive method; this record is the paper-specific evaluated configuration.SourcesCell2Sentence: Teaching Large Language Models the Language of Biology · Methods (paragraph 1); Experiments/Fine-Tuning Datasets (paragraph 3) |
| Architecture / procedure | Genes are ordered by decreasing transcript abundance to create cell sentences. GPT-2 Large is fine-tuned on this representation for cell-type-conditioned generation and cell-type annotation.SourcesCell2Sentence: Teaching Large Language Models the Language of Biology · Methods (paragraph 1); Experiments/Fine-Tuning Datasets (paragraph 3) |
| Biological inputs | Single-cell gene-expression profiles converted to ranked gene namesSourcesCell2Sentence: Teaching Large Language Models the Language of Biology · Methods/Data transformation (paragraph 8); Methods/Data transformation (paragraph 3) |
| Outputs | Cell-type annotations or generated ranked gene sequencesSourcesCell2Sentence: Teaching Large Language Models the Language of Biology · Methods/Data transformation (paragraph 7); Inference Details (paragraph 2) |
| Parameters | 774,030,080 parameters for the model checkpoint used in the Cell2Sentence L1000 experiment; this is not a claim about every checkpoint in the family.SourcesCell2Sentence: Teaching Large Language Models the Language of Biology · Table 5, Comparison of required compute on the L1000 dataset; # Parameters column and caption |
| Known versions / configuration | GPT-2 LargeSourcesCell2Sentence: Teaching Large Language Models the Language of Biology · Experiments/Experiment 3: abstract summary generation/Objective: (paragraph 1); Experiments/Fine-Tuning Datasets (paragraph 2) |
| Training data / fitting | Task-specific single-cell expression datasets described in the preprint; gene ranking replaces direct numeric expression input.SourcesCell2Sentence: Teaching Large Language Models the Language of Biology · Methods/Data transformation (paragraph 8); Experimental Details/Evaluation Datasets (paragraph 1) |
| Context limits | GPT-2 configurations use 1,024 tokens; the separately evaluated Pythia-160m configuration uses 9,200 tokens.SourcesCell2Sentence: Teaching Large Language Models the Language of Biology · Training Details (paragraph 1); Experiments (paragraph 1) |
| Access | Official study implementation and usage documentation: https://github.com/vandijklab/cell2sentence/blob/a6efaf079f98491d4723ced44b929936b94368aa/README.md. This pinned documentation revision is not automatically the evaluated weight revision.Sourcesvandijklab/cell2sentence README.md · README.md; installation, model download and usage instructions |
| Code licence | Apache 2.0 (study repository code at the cited revision; this does not establish every dependency or historical checkpoint licence).Sourcesvandijklab/cell2sentence LICENSE · LICENSE; complete licence text |
| Weights licence | The inspected model-access documentation does not explicitly identify terms for this exact evaluated checkpoint or fitted head; repository code terms are shown separately. · Not reported in inspected sourcesSourcesvandijklab/cell2sentence README.md · README.md; checkpoint/access documentation and licence scope |
Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.
One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.
19 evidence rows matching the loaded filters
| Property and statement | Original source and location | Review and provenance |
|---|---|---|
| Diagram caption Conceptual input–method–output guide. Check the procedure text and linked evaluation for fitted components, additional inputs and exact settings. Individual claims | Cell2Sentence: Teaching Large Language Models the Language of Biology Methods (paragraph 1); Experiments/Fine-Tuning Datasets (paragraph 3) Version: preprint archived 2024-10-29 | source checked automated source review · 2026-09-16 Audit detailsPrimary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Diagram steps ["Single-cell gene-expression profiles converted to ranked gene names","C2S (GPT-2 Large)","Cell-type annotations or generated ranked gene sequences"] Individual claims | Cell2Sentence: Teaching Large Language Models the Language of Biology Methods (paragraph 1); Experiments/Fine-Tuning Datasets (paragraph 3) Version: preprint archived 2024-10-29 | source checked automated source review · 2026-09-16 Audit detailsPrimary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Diagram title Evaluated procedure (conceptual) Individual claims | Cell2Sentence: Teaching Large Language Models the Language of Biology Methods (paragraph 1); Experiments/Fine-Tuning Datasets (paragraph 3) Version: preprint archived 2024-10-29 | source checked automated source review · 2026-09-16 Audit detailsPrimary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Model type Study-specific predictive method; this record is the paper-specific evaluated configuration. Individual claims | Cell2Sentence: Teaching Large Language Models the Language of Biology Methods (paragraph 1); Experiments/Fine-Tuning Datasets (paragraph 3) Version: preprint archived 2024-10-29 | source checked automated source review · 2026-09-16 Audit detailsPrimary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Architecture / procedure Genes are ordered by decreasing transcript abundance to create cell sentences. GPT-2 Large is fine-tuned on this representation for cell-type-conditioned generation and cell-type annotation. Individual claims | Cell2Sentence: Teaching Large Language Models the Language of Biology Methods (paragraph 1); Experiments/Fine-Tuning Datasets (paragraph 3) Version: preprint archived 2024-10-29 | source checked automated source review · 2026-09-16 Audit detailsPrimary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Weights licence The inspected model-access documentation does not explicitly identify terms for this exact evaluated checkpoint or fitted head; repository code terms are shown separately. Individual claims | vandijklab/cell2sentence README.md README.md; checkpoint/access documentation and licence scope Version: a6efaf079f98491d4723ced44b929936b94368aa | unreported automated source review · 2026-09-16 Audit detailsPrimary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Biological inputs Single-cell gene-expression profiles converted to ranked gene names Individual claims | Cell2Sentence: Teaching Large Language Models the Language of Biology Methods/Data transformation (paragraph 8); Methods/Data transformation (paragraph 3) Version: preprint archived 2024-10-29 | source checked automated source review · 2026-09-16 Audit detailsPrimary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Outputs Cell-type annotations or generated ranked gene sequences Individual claims | Cell2Sentence: Teaching Large Language Models the Language of Biology Methods/Data transformation (paragraph 7); Inference Details (paragraph 2) Version: preprint archived 2024-10-29 | source checked automated source review · 2026-09-16 Audit detailsPrimary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Parameters 774,030,080 parameters for the model checkpoint used in the Cell2Sentence L1000 experiment; this is not a claim about every checkpoint in the family. Individual claims | Cell2Sentence: Teaching Large Language Models the Language of Biology Table 5, Comparison of required compute on the L1000 dataset; # Parameters column and caption Version: preprint archived 2024-10-29 | source checked automated source review · 2026-09-16 Audit detailsPrimary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Known versions / configuration GPT-2 Large Individual claims | Cell2Sentence: Teaching Large Language Models the Language of Biology Experiments/Experiment 3: abstract summary generation/Objective: (paragraph 1); Experiments/Fine-Tuning Datasets (paragraph 2) Version: preprint archived 2024-10-29 | source checked automated source review · 2026-09-16 Audit detailsPrimary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
Release 2026-09-17-d277315f7d76 · Record review: needs review
Stable ID: reported-model-ab02228f50a37c