rewire.it
Model

Nucleotide Transformer v2

Nucleotide Transformer v2 represents DNA using an encoder pretrained on multiple species.

Sources (7)InstaDeepAI/nucleotide-transformer-v2-50m-multi-species: README.md; InstaDeepAI/nucleotide-transformer-v2-50m-multi-species: config.json; instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; nt: Journal full-text XML · README.md: Model Summary, Training data and licence metadata; config.json; Nucleotide Transformer paper Methods: Architecture and Training (v2); source conflict with 50M card retained

44 evaluations · 44 metric rows · 1 evaluated configuration using this model

How it worksNucleotide Transformer v2 workflow
Nucleotide Transformer v2 workflow1. DNA sequence. Then: 2. 6-mer tokenizer. Then: 3. NT-v2 encoder. Then: 4. Embeddings or task adaptationNucleotide Transformer v2 workflow1. DNA sequence. Then: 2. 6-mer tokenizer. Then: 3. NT-v2 encoder. Then: 4. Embeddings or task adaptationNucleotide Transformer v2 workflow1. DNA sequence. Then: 2. 6-mer tokenizer. Then: 3. NT-v2 encoder. Then: 4. Embeddings or task adaptation

Conceptual summary of the documented data flow; optional inputs and configured downstream stages must be reported for a reproducible evaluation.

Sources (7)InstaDeepAI/nucleotide-transformer-v2-50m-multi-species: README.md; InstaDeepAI/nucleotide-transformer-v2-50m-multi-species: config.json; instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; nt: Journal full-text XML · README.md: Model Summary, Training data and licence metadata; config.json; Nucleotide Transformer paper Methods: Architecture and Training (v2); source conflict with 50M card retained

Overview

Model type

DNA transformer encoder

Sources (7)InstaDeepAI/nucleotide-transformer-v2-50m-multi-species: README.md; InstaDeepAI/nucleotide-transformer-v2-50m-multi-species: config.json; instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; nt: Journal full-text XML · README.md: Model Summary, Training data and licence metadata; config.json; Nucleotide Transformer paper Methods: Architecture and Training (v2); source conflict with 50M card retained

Inputs

DNA sequences with tokenization determined by 6-mers and individual ambiguous/remainder bases.

Sources (7)InstaDeepAI/nucleotide-transformer-v2-50m-multi-species: README.md; InstaDeepAI/nucleotide-transformer-v2-50m-multi-species: config.json; instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; nt: Journal full-text XML · README.md: Model Summary, Training data and licence metadata; config.json; Nucleotide Transformer paper Methods: Architecture and Training (v2); source conflict with 50M card retained

Outputs

Contextual DNA embeddings and masked-token probabilities; downstream tasks need adaptation.

Sources (7)InstaDeepAI/nucleotide-transformer-v2-50m-multi-species: README.md; InstaDeepAI/nucleotide-transformer-v2-50m-multi-species: config.json; instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; nt: Journal full-text XML · README.md: Model Summary, Training data and licence metadata; config.json; Nucleotide Transformer paper Methods: Architecture and Training (v2); source conflict with 50M card retained

Access

Official downloadable model/card and usage examples: https://huggingface.co/InstaDeepAI/nucleotide-transformer-v2-50m-multi-species

Sources (7)InstaDeepAI/nucleotide-transformer-v2-50m-multi-species: README.md; InstaDeepAI/nucleotide-transformer-v2-50m-multi-species: config.json; instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; nt: Journal full-text XML · README.md: Model Summary, Training data and licence metadata; config.json; Nucleotide Transformer paper Methods: Architecture and Training (v2); source conflict with 50M card retained

limited source coverage · Automated source review, 2026-09-16. All specifications and missing details

Evaluations and results

44 evaluations · 44 metric rows. Different protocols are not a single leaderboard.

Filter evaluations

Applied filters: All linked evaluations

Exact evaluated configurations and original reported results
Tested configurationProtocol and datasetFindingEvidence and details
Configuration: NT-v2-50M-MSTask: GENEB LINEAR-PROBE: Average macro-MCC across the 13 representative tasks, linear probe
Dataset subset: GENEB representative task subset (GENEB split)
0.511 macro_mcc
correlation · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · source checked
Methods, coverage and source

NT-v2-50M-MS on GENEB LINEAR-PROBE: Average macro-MCC across the 13 representative tasks, linear probe

Frozen embeddings scored with a probe over the thirteen representative tasks, one from each functional category (Table 7).

Aggregation: Not reported

GENEB: Why Genomic Models Are Hard to Compare · Table 8, row(NT-v2-50M-MS), column(Linear MCC)
Configuration: NT-v2-50M-MSTask: GENEB MLP-PROBE: Average macro-MCC across the 13 representative tasks, MLP probe
Dataset subset: GENEB representative task subset (GENEB split)
0.521 macro_mcc
correlation · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · source checked
Methods, coverage and source

NT-v2-50M-MS on GENEB MLP-PROBE: Average macro-MCC across the 13 representative tasks, MLP probe

Frozen embeddings scored with a probe over the thirteen representative tasks, one from each functional category (Table 7).

Aggregation: Not reported

GENEB: Why Genomic Models Are Hard to Compare · Table 8, row(NT-v2-50M-MS), column(MLP MCC)
Configuration: N.T.-v2-50mTask: NABench CCV-CORR-APTAMER: Fitness prediction on aptamer assays, supervised, contiguous cross validation
Dataset subset: NABench aptamer assays (NABench split)
0.056 spearman
correlation · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · source checked
Methods, coverage and source

N.T.-v2-50m on NABench CCV-CORR-APTAMER: Fitness prediction on aptamer assays, supervised, contiguous cross validation

Scored supervised, contiguous cross validation across the NABench aptamer assays.

Aggregation: Not reported

NABench: Large-Scale Benchmarks of Nucleotide Foundation Models for Fitness Prediction · Table 11, row(N.T.-v2-50m), column(aptamer)
Configuration: N.T.-v2-50mTask: NABench CCV-CORR-ENHANCER: Fitness prediction on enhancer assays, supervised, contiguous cross validation
Dataset subset: NABench enhancer assays (NABench split)
0.116 spearman
correlation · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · source checked
Methods, coverage and source

N.T.-v2-50m on NABench CCV-CORR-ENHANCER: Fitness prediction on enhancer assays, supervised, contiguous cross validation

Scored supervised, contiguous cross validation across the NABench enhancer assays.

Aggregation: Not reported

NABench: Large-Scale Benchmarks of Nucleotide Foundation Models for Fitness Prediction · Table 11, row(N.T.-v2-50m), column(enhancer)
Configuration: N.T.-v2-50mTask: NABench CCV-CORR-MRNA: Fitness prediction on mRNA assays, supervised, contiguous cross validation
Dataset subset: NABench mRNA assays (NABench split)
0.073 spearman
correlation · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · source checked
Methods, coverage and source

N.T.-v2-50m on NABench CCV-CORR-MRNA: Fitness prediction on mRNA assays, supervised, contiguous cross validation

Scored supervised, contiguous cross validation across the NABench mRNA assays.

Aggregation: Not reported

NABench: Large-Scale Benchmarks of Nucleotide Foundation Models for Fitness Prediction · Table 11, row(N.T.-v2-50m), column(mRNA)
Configuration: N.T.-v2-50mTask: NABench CCV-CORR-PROMOTER: Fitness prediction on promoter assays, supervised, contiguous cross validation
Dataset subset: NABench promoter assays (NABench split)
0.220 spearman
correlation · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · source checked
Methods, coverage and source

N.T.-v2-50m on NABench CCV-CORR-PROMOTER: Fitness prediction on promoter assays, supervised, contiguous cross validation

Scored supervised, contiguous cross validation across the NABench promoter assays.

Aggregation: Not reported

NABench: Large-Scale Benchmarks of Nucleotide Foundation Models for Fitness Prediction · Table 11, row(N.T.-v2-50m), column(promoter)
Configuration: N.T.-v2-50mTask: NABench CCV-CORR-RIBOZYME: Fitness prediction on ribozyme assays, supervised, contiguous cross validation
Dataset subset: NABench ribozyme assays (NABench split)
0.300 spearman
correlation · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · source checked
Methods, coverage and source

N.T.-v2-50m on NABench CCV-CORR-RIBOZYME: Fitness prediction on ribozyme assays, supervised, contiguous cross validation

Scored supervised, contiguous cross validation across the NABench ribozyme assays.

Aggregation: Not reported

NABench: Large-Scale Benchmarks of Nucleotide Foundation Models for Fitness Prediction · Table 11, row(N.T.-v2-50m), column(ribozyme)
Configuration: N.T.-v2-50mTask: NABench DMS-CCV: Overall fitness prediction on NABench deep mutational scanning assays, Contiguous cross validation Spearman ρ
Dataset subset: NABench deep mutational scanning assays (NABench split)
0.220 spearman
correlation · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · source checked
Methods, coverage and source

N.T.-v2-50m on NABench DMS-CCV: Overall fitness prediction on NABench deep mutational scanning assays, Contiguous cross validation Spearman ρ

Aggregated by the NABench authors across NABench deep mutational scanning assays.

Aggregation: Not reported

NABench: Large-Scale Benchmarks of Nucleotide Foundation Models for Fitness Prediction · Table 6, row(N.T.-v2-50m), column(Contiguous cross validation Spearman ρ)
Configuration: N.T.-v2-50mTask: NABench DMS-FS: Overall fitness prediction on NABench deep mutational scanning assays, Few-shot Spearman ρ
Dataset subset: NABench deep mutational scanning assays (NABench split)
0.110 spearman
correlation · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · source checked
Methods, coverage and source

N.T.-v2-50m on NABench DMS-FS: Overall fitness prediction on NABench deep mutational scanning assays, Few-shot Spearman ρ

Aggregated by the NABench authors across NABench deep mutational scanning assays.

Aggregation: Not reported

NABench: Large-Scale Benchmarks of Nucleotide Foundation Models for Fitness Prediction · Table 6, row(N.T.-v2-50m), column(Few-shot Spearman ρ)
Configuration: N.T.-v2-50mTask: NABench DMS-RCV: Overall fitness prediction on NABench deep mutational scanning assays, Random cross validation Spearman ρ
Dataset subset: NABench deep mutational scanning assays (NABench split)
0.465 spearman
correlation · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · source checked
Methods, coverage and source

N.T.-v2-50m on NABench DMS-RCV: Overall fitness prediction on NABench deep mutational scanning assays, Random cross validation Spearman ρ

Aggregated by the NABench authors across NABench deep mutational scanning assays.

Aggregation: Not reported

NABench: Large-Scale Benchmarks of Nucleotide Foundation Models for Fitness Prediction · Table 6, row(N.T.-v2-50m), column(Random cross validation Spearman ρ)
Configuration: N.T.-v2-50mTask: NABench DMS-ZS-AUC: Overall fitness prediction on NABench deep mutational scanning assays, Zero-shot AUC
Dataset subset: NABench deep mutational scanning assays (NABench split)
0.512 auc
fraction · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · source checked
Methods, coverage and source

N.T.-v2-50m on NABench DMS-ZS-AUC: Overall fitness prediction on NABench deep mutational scanning assays, Zero-shot AUC

Aggregated by the NABench authors across NABench deep mutational scanning assays.

Aggregation: Not reported

NABench: Large-Scale Benchmarks of Nucleotide Foundation Models for Fitness Prediction · Table 6, row(N.T.-v2-50m), column(Zero-shot AUC)
Configuration: N.T.-v2-50mTask: NABench DMS-ZS-CORR: Overall fitness prediction on NABench deep mutational scanning assays, Zero-shot Spearman ρ
Dataset subset: NABench deep mutational scanning assays (NABench split)
0.097 spearman
correlation · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · source checked
Methods, coverage and source

N.T.-v2-50m on NABench DMS-ZS-CORR: Overall fitness prediction on NABench deep mutational scanning assays, Zero-shot Spearman ρ

Aggregated by the NABench authors across NABench deep mutational scanning assays.

Aggregation: Not reported

NABench: Large-Scale Benchmarks of Nucleotide Foundation Models for Fitness Prediction · Table 6, row(N.T.-v2-50m), column(Zero-shot Spearman ρ)
Configuration: N.T.-v2-50mTask: NABench DMS-ZS-MCC: Overall fitness prediction on NABench deep mutational scanning assays, Zero-shot MCC
Dataset subset: NABench deep mutational scanning assays (NABench split)
0.044 mcc
correlation · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · source checked
Methods, coverage and source

N.T.-v2-50m on NABench DMS-ZS-MCC: Overall fitness prediction on NABench deep mutational scanning assays, Zero-shot MCC

Aggregated by the NABench authors across NABench deep mutational scanning assays.

Aggregation: Not reported

NABench: Large-Scale Benchmarks of Nucleotide Foundation Models for Fitness Prediction · Table 6, row(N.T.-v2-50m), column(Zero-shot MCC)
Configuration: N.T.-v2-50mTask: NABench DMS-ZS-NDCG: Overall fitness prediction on NABench deep mutational scanning assays, Zero-shot NDCG
Dataset subset: NABench deep mutational scanning assays (NABench split)
0.349 ndcg
fraction · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · source checked
Methods, coverage and source

N.T.-v2-50m on NABench DMS-ZS-NDCG: Overall fitness prediction on NABench deep mutational scanning assays, Zero-shot NDCG

Aggregated by the NABench authors across NABench deep mutational scanning assays.

Aggregation: Not reported

NABench: Large-Scale Benchmarks of Nucleotide Foundation Models for Fitness Prediction · Table 6, row(N.T.-v2-50m), column(Zero-shot NDCG)
Configuration: N.T.-v2-50mTask: NABench FS-CORR-APTAMER: Fitness prediction on aptamer assays, few-shot
Dataset subset: NABench aptamer assays (NABench split)
0.230 spearman
correlation · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · source checked
Methods, coverage and source

N.T.-v2-50m on NABench FS-CORR-APTAMER: Fitness prediction on aptamer assays, few-shot

Scored few-shot across the NABench aptamer assays.

Aggregation: Not reported

NABench: Large-Scale Benchmarks of Nucleotide Foundation Models for Fitness Prediction · Table 9, row(N.T.-v2-50m), column(aptamer)
Configuration: N.T.-v2-50mTask: NABench FS-CORR-ENHANCER: Fitness prediction on enhancer assays, few-shot
Dataset subset: NABench enhancer assays (NABench split)
0.026 spearman
correlation · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · source checked
Methods, coverage and source

N.T.-v2-50m on NABench FS-CORR-ENHANCER: Fitness prediction on enhancer assays, few-shot

Scored few-shot across the NABench enhancer assays.

Aggregation: Not reported

NABench: Large-Scale Benchmarks of Nucleotide Foundation Models for Fitness Prediction · Table 9, row(N.T.-v2-50m), column(enhancer)
Configuration: N.T.-v2-50mTask: NABench FS-CORR-MRNA: Fitness prediction on mRNA assays, few-shot
Dataset subset: NABench mRNA assays (NABench split)
0.214 spearman
correlation · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · source checked
Methods, coverage and source

N.T.-v2-50m on NABench FS-CORR-MRNA: Fitness prediction on mRNA assays, few-shot

Scored few-shot across the NABench mRNA assays.

Aggregation: Not reported

NABench: Large-Scale Benchmarks of Nucleotide Foundation Models for Fitness Prediction · Table 9, row(N.T.-v2-50m), column(mRNA)
Configuration: N.T.-v2-50mTask: NABench FS-CORR-PROMOTER: Fitness prediction on promoter assays, few-shot
Dataset subset: NABench promoter assays (NABench split)
0.061 spearman
correlation · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · source checked
Methods, coverage and source

N.T.-v2-50m on NABench FS-CORR-PROMOTER: Fitness prediction on promoter assays, few-shot

Scored few-shot across the NABench promoter assays.

Aggregation: Not reported

NABench: Large-Scale Benchmarks of Nucleotide Foundation Models for Fitness Prediction · Table 9, row(N.T.-v2-50m), column(promoter)
Configuration: N.T.-v2-50mTask: NABench FS-CORR-RIBOZYME: Fitness prediction on ribozyme assays, few-shot
Dataset subset: NABench ribozyme assays (NABench split)
0.090 spearman
correlation · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · source checked
Methods, coverage and source

N.T.-v2-50m on NABench FS-CORR-RIBOZYME: Fitness prediction on ribozyme assays, few-shot

Scored few-shot across the NABench ribozyme assays.

Aggregation: Not reported

NABench: Large-Scale Benchmarks of Nucleotide Foundation Models for Fitness Prediction · Table 9, row(N.T.-v2-50m), column(ribozyme)
Configuration: N.T.-v2-50mTask: NABench FS-CORR-TRNA: Fitness prediction on tRNA assays, few-shot
Dataset subset: NABench tRNA assays (NABench split)
0.327 spearman
correlation · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · source checked
Methods, coverage and source

N.T.-v2-50m on NABench FS-CORR-TRNA: Fitness prediction on tRNA assays, few-shot

Scored few-shot across the NABench tRNA assays.

Aggregation: Not reported

NABench: Large-Scale Benchmarks of Nucleotide Foundation Models for Fitness Prediction · Table 9, row(N.T.-v2-50m), column(tRNA)
Configuration: N.T.-v2-50mTask: NABench RCV-CORR-APTAMER: Fitness prediction on aptamer assays, supervised, random cross validation
Dataset subset: NABench aptamer assays (NABench split)
0.462 spearman
correlation · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · source checked
Methods, coverage and source

N.T.-v2-50m on NABench RCV-CORR-APTAMER: Fitness prediction on aptamer assays, supervised, random cross validation

Scored supervised, random cross validation across the NABench aptamer assays.

Aggregation: Not reported

NABench: Large-Scale Benchmarks of Nucleotide Foundation Models for Fitness Prediction · Table 10, row(N.T.-v2-50m), column(aptamer)
Configuration: N.T.-v2-50mTask: NABench RCV-CORR-ENHANCER: Fitness prediction on enhancer assays, supervised, random cross validation
Dataset subset: NABench enhancer assays (NABench split)
0.148 spearman
correlation · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · source checked
Methods, coverage and source

N.T.-v2-50m on NABench RCV-CORR-ENHANCER: Fitness prediction on enhancer assays, supervised, random cross validation

Scored supervised, random cross validation across the NABench enhancer assays.

Aggregation: Not reported

NABench: Large-Scale Benchmarks of Nucleotide Foundation Models for Fitness Prediction · Table 10, row(N.T.-v2-50m), column(enhancer)
Configuration: N.T.-v2-50mTask: NABench RCV-CORR-MRNA: Fitness prediction on mRNA assays, supervised, random cross validation
Dataset subset: NABench mRNA assays (NABench split)
0.570 spearman
correlation · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · source checked
Methods, coverage and source

N.T.-v2-50m on NABench RCV-CORR-MRNA: Fitness prediction on mRNA assays, supervised, random cross validation

Scored supervised, random cross validation across the NABench mRNA assays.

Aggregation: Not reported

NABench: Large-Scale Benchmarks of Nucleotide Foundation Models for Fitness Prediction · Table 10, row(N.T.-v2-50m), column(mRNA)
Configuration: N.T.-v2-50mTask: NABench RCV-CORR-PROMOTER: Fitness prediction on promoter assays, supervised, random cross validation
Dataset subset: NABench promoter assays (NABench split)
0.632 spearman
correlation · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · source checked
Methods, coverage and source

N.T.-v2-50m on NABench RCV-CORR-PROMOTER: Fitness prediction on promoter assays, supervised, random cross validation

Scored supervised, random cross validation across the NABench promoter assays.

Aggregation: Not reported

NABench: Large-Scale Benchmarks of Nucleotide Foundation Models for Fitness Prediction · Table 10, row(N.T.-v2-50m), column(promoter)
Configuration: N.T.-v2-50mTask: NABench RCV-CORR-RIBOZYME: Fitness prediction on ribozyme assays, supervised, random cross validation
Dataset subset: NABench ribozyme assays (NABench split)
0.464 spearman
correlation · higher

Uncertainty: Not reported

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · source checked
Methods, coverage and source

N.T.-v2-50m on NABench RCV-CORR-RIBOZYME: Fitness prediction on ribozyme assays, supervised, random cross validation

Scored supervised, random cross validation across the NABench ribozyme assays.

Aggregation: Not reported

NABench: Large-Scale Benchmarks of Nucleotide Foundation Models for Fitness Prediction · Table 10, row(N.T.-v2-50m), column(ribozyme)

Source checking is not independent reproduction. Release 2026-09-23-2b89723c6dd9.

Related configurations, pipelines and services

These configurations, services and pipelines use this model within their own configurations. Their results, where available, are not assigned to the underlying model.

Use this model

How it works, versions and access

Related profile: Nucleotide Transformer. This page retains the exact record and its evaluation context.

Versions and evaluated configurations

How it works

How it works

Nucleotide Transformer v2 represents DNA using an encoder pretrained on multiple species. Encoder-only transformer with 6-mer tokenization, rotary position embeddings and SwiGLU feed-forward layers. The documented inputs are DNA sequences with tokenization determined by 6-mers and individual ambiguous/remainder bases. The output consists of contextual DNA embeddings and masked-token probabilities; downstream tasks need adaptation.

Sources (7)InstaDeepAI/nucleotide-transformer-v2-50m-multi-species: README.md; InstaDeepAI/nucleotide-transformer-v2-50m-multi-species: config.json; instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; nt: Journal full-text XML · README.md: Model Summary, Training data and licence metadata; config.json; Nucleotide Transformer paper Methods: Architecture and Training (v2); source conflict with 50M card retained
Versions and reproducibility

nucleotide-transformer-v2-50m-multi-species is the linked checkpoint; distinguish it from other family scales. Source conflict retained: the model card describes 1,000-token pretraining, while the official NT-v2 documentation describes 2,048-token capacity. Token and base counts must be stated separately.

Sources (7)InstaDeepAI/nucleotide-transformer-v2-50m-multi-species: README.md; InstaDeepAI/nucleotide-transformer-v2-50m-multi-species: config.json; instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; nt: Journal full-text XML · README.md: Model Summary, Training data and licence metadata; config.json; Nucleotide Transformer paper Methods: Architecture and Training (v2); source conflict with 50M card retained
Strengths, limitations and unresolved questions

Strengths and limitations

Strengths and considerations

Limitations and conditions

Profile review details

Inspected pinned official documentation, relevant implementation files and named primary-paper sections. Claims are limited to those artifacts. Remaining field extraction and identity conflicts are explicit; no new performance claims, model runs or human review are implied.

Stable record: catalog-model-nt-v2

Specifications

Inputs, training, access and other details

Explanatory profile: limited source coverage · Automated source review, 2026-09-16. Review applies to the cited claims; unresolved fields are listed below. Numerical results retain their own review status.

Inputs, outputs and configuration
PropertyDescription and evidence
Model typeDNA transformer encoder
Sources (7)InstaDeepAI/nucleotide-transformer-v2-50m-multi-species: README.md; InstaDeepAI/nucleotide-transformer-v2-50m-multi-species: config.json; instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; nt: Journal full-text XML · README.md: Model Summary, Training data and licence metadata; config.json; Nucleotide Transformer paper Methods: Architecture and Training (v2); source conflict with 50M card retained
ArchitectureEncoder-only transformer with 6-mer tokenization, rotary position embeddings and SwiGLU feed-forward layers.
Sources (7)InstaDeepAI/nucleotide-transformer-v2-50m-multi-species: README.md; InstaDeepAI/nucleotide-transformer-v2-50m-multi-species: config.json; instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; nt: Journal full-text XML · README.md: Model Summary, Training data and licence metadata; config.json; Nucleotide Transformer paper Methods: Architecture and Training (v2); source conflict with 50M card retained
InputsDNA sequences with tokenization determined by 6-mers and individual ambiguous/remainder bases.
Sources (7)InstaDeepAI/nucleotide-transformer-v2-50m-multi-species: README.md; InstaDeepAI/nucleotide-transformer-v2-50m-multi-species: config.json; instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; nt: Journal full-text XML · README.md: Model Summary, Training data and licence metadata; config.json; Nucleotide Transformer paper Methods: Architecture and Training (v2); source conflict with 50M card retained
OutputsContextual DNA embeddings and masked-token probabilities; downstream tasks need adaptation.
Sources (7)InstaDeepAI/nucleotide-transformer-v2-50m-multi-species: README.md; InstaDeepAI/nucleotide-transformer-v2-50m-multi-species: config.json; instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; nt: Journal full-text XML · README.md: Model Summary, Training data and licence metadata; config.json; Nucleotide Transformer paper Methods: Architecture and Training (v2); source conflict with 50M card retained
ParametersThe linked checkpoint is 50M; v2 family also includes 100M, 250M and 500M.
Sources (7)InstaDeepAI/nucleotide-transformer-v2-50m-multi-species: README.md; InstaDeepAI/nucleotide-transformer-v2-50m-multi-species: config.json; instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; nt: Journal full-text XML · README.md: Model Summary, Training data and licence metadata; config.json; Nucleotide Transformer paper Methods: Architecture and Training (v2); source conflict with 50M card retained
Known versionsnucleotide-transformer-v2-50m-multi-species is the linked checkpoint; distinguish it from other family scales.
Sources (7)InstaDeepAI/nucleotide-transformer-v2-50m-multi-species: README.md; InstaDeepAI/nucleotide-transformer-v2-50m-multi-species: config.json; instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; nt: Journal full-text XML · README.md: Model Summary, Training data and licence metadata; config.json; Nucleotide Transformer paper Methods: Architecture and Training (v2); source conflict with 50M card retained
Training dataThe linked 50M card reports 850 reference genomes, excluding plants and viruses; 174B source nucleotides and 300B training tokens. Source corpus size is distinct from repeated training exposure.
Sources (7)InstaDeepAI/nucleotide-transformer-v2-50m-multi-species: README.md; InstaDeepAI/nucleotide-transformer-v2-50m-multi-species: config.json; instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; nt: Journal full-text XML · README.md: Model Summary, Training data and licence metadata; config.json; Nucleotide Transformer paper Methods: Architecture and Training (v2); source conflict with 50M card retained
Training cutoffThe paper specifies human-reference, 1000 Genomes and multispecies training collections by variant. A single latest-deposition date for all sequences is not supplied in the inspected pretraining-data section. · Not reported in inspected sources
Sources (7)InstaDeepAI/nucleotide-transformer-v2-50m-multi-species: README.md; InstaDeepAI/nucleotide-transformer-v2-50m-multi-species: config.json; instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; nt: Journal full-text XML · README.md: Model Summary, Training data and licence metadata; config.json; Nucleotide Transformer paper Methods: Architecture and Training (v2); source conflict with 50M card retained
Context limitsThe inspected 50M model card describes 1,000-token pretraining, whereas the NT-v2 paper and documentation describe 2,048-token capacity. Preserve this source discrepancy and the selected checkpoint; tokens and bases differ.
Sources (7)InstaDeepAI/nucleotide-transformer-v2-50m-multi-species: README.md; InstaDeepAI/nucleotide-transformer-v2-50m-multi-species: config.json; instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; nt: Journal full-text XML · README.md: Model Summary, Training data and licence metadata; config.json; Nucleotide Transformer paper Methods: Architecture and Training (v2); source conflict with 50M card retained
Weights licenceCC-BY-NC-SA-4.0 as declared in the official model card.
Sources (7)InstaDeepAI/nucleotide-transformer-v2-50m-multi-species: README.md; InstaDeepAI/nucleotide-transformer-v2-50m-multi-species: config.json; instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; nt: Journal full-text XML · README.md: Model Summary, Training data and licence metadata; config.json; Nucleotide Transformer paper Methods: Architecture and Training (v2); source conflict with 50M card retained
AccessOfficial downloadable model/card and usage examples: https://huggingface.co/InstaDeepAI/nucleotide-transformer-v2-50m-multi-species
Sources (7)InstaDeepAI/nucleotide-transformer-v2-50m-multi-species: README.md; InstaDeepAI/nucleotide-transformer-v2-50m-multi-species: config.json; instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; nt: Journal full-text XML · README.md: Model Summary, Training data and licence metadata; config.json; Nucleotide Transformer paper Methods: Architecture and Training (v2); source conflict with 50M card retained
Code licenceCC-BY-NC-SA-4.0
Sourcesinstadeepai/nucleotide-transformer: LICENSE.md · LICENSE.md: licence text

Evidence

Source checking verifies the cited claim or transcription. It does not establish independent reproduction.

Evidence table

Inspect claims, sources and review details

Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.

One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.

142 evidence rows matching the loaded filters

Claims, original sources and review scope · Release 2026-09-23-2b89723c6dd9
Property and statementOriginal source and locationReview and provenance
Diagram caption
Conceptual summary of the documented data flow; optional inputs and configured downstream stages must be reported for a reproducible evaluation.
Individual claims
instadeepai/nucleotide-transformer: README.md

Original source ↗

README.md: Model Summary, Training data and licence metadata; config.json; Nucleotide Transformer paper Methods: Architecture and Training (v2); source conflict with 50M card retained

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: 2dc37b86e16a6970fbc731751f7719d9f676f7f9
Retrieved: 2026-09-16T19:46:19.364532+00:00

source checked

automated source review · 2026-09-16

Audit details

Inspected pinned official documentation, relevant implementation files and named primary-paper sections. Claims are limited to those artifacts. Remaining field extraction and identity conflicts are explicit; no new performance claims, model runs or human review are implied.

Field: attributes.profile.diagram.caption

Source artifact SHA-256: 9f51bbb20c4c5c36e77fb03ca1c5c36236e287c48a1ee31f53150545d421ec25

Hash scope: SHA-256 of retrieved original artifact bytes

Format: original_artifact

Inspected artifact

Diagram caption
Conceptual summary of the documented data flow; optional inputs and configured downstream stages must be reported for a reproducible evaluation.
Individual claims
instadeepai/nucleotide-transformer: docs/segment_nt.md

Original source ↗

README.md: Model Summary, Training data and licence metadata; config.json; Nucleotide Transformer paper Methods: Architecture and Training (v2); source conflict with 50M card retained

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: 2dc37b86e16a6970fbc731751f7719d9f676f7f9
Retrieved: 2026-09-16T19:46:19.364532+00:00

source checked

automated source review · 2026-09-16

Audit details

Inspected pinned official documentation, relevant implementation files and named primary-paper sections. Claims are limited to those artifacts. Remaining field extraction and identity conflicts are explicit; no new performance claims, model runs or human review are implied.

Field: attributes.profile.diagram.caption

Source artifact SHA-256: 8eec4580ba64ab944fb9b42674be70ffe793135f3503cce8640f5b08f8290f7a

Hash scope: SHA-256 of retrieved original artifact bytes

Format: original_artifact

Inspected artifact

Diagram caption
Conceptual summary of the documented data flow; optional inputs and configured downstream stages must be reported for a reproducible evaluation.
Individual claims
InstaDeepAI/nucleotide-transformer-v2-50m-multi-species: README.md

Original source ↗

README.md: Model Summary, Training data and licence metadata; config.json; Nucleotide Transformer paper Methods: Architecture and Training (v2); source conflict with 50M card retained

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: 81b29e5786726d891dbf929404ef20adca5b36f1
Retrieved: 2026-09-16T19:46:20.607357+00:00

source checked

automated source review · 2026-09-16

Audit details

Inspected pinned official documentation, relevant implementation files and named primary-paper sections. Claims are limited to those artifacts. Remaining field extraction and identity conflicts are explicit; no new performance claims, model runs or human review are implied.

Field: attributes.profile.diagram.caption

Source artifact SHA-256: e526d7b98f106bc2ca9ba73fa166ff5fd62853812e9757a6692eeedf25e42923

Hash scope: SHA-256 of retrieved original artifact bytes

Format: original_artifact

Inspected artifact

Diagram caption
Conceptual summary of the documented data flow; optional inputs and configured downstream stages must be reported for a reproducible evaluation.
Individual claims
instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md

Original source ↗

README.md: Model Summary, Training data and licence metadata; config.json; Nucleotide Transformer paper Methods: Architecture and Training (v2); source conflict with 50M card retained

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: 2dc37b86e16a6970fbc731751f7719d9f676f7f9
Retrieved: 2026-09-16T19:46:19.364532+00:00

source checked

automated source review · 2026-09-16

Audit details

Inspected pinned official documentation, relevant implementation files and named primary-paper sections. Claims are limited to those artifacts. Remaining field extraction and identity conflicts are explicit; no new performance claims, model runs or human review are implied.

Field: attributes.profile.diagram.caption

Source artifact SHA-256: 23e27d40473bbabadac45c56e8e282349f0b053da85fda89ff2125b5fe381fc6

Hash scope: SHA-256 of retrieved original artifact bytes

Format: original_artifact

Inspected artifact

Diagram caption
Conceptual summary of the documented data flow; optional inputs and configured downstream stages must be reported for a reproducible evaluation.
Individual claims
instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md

Original source ↗

README.md: Model Summary, Training data and licence metadata; config.json; Nucleotide Transformer paper Methods: Architecture and Training (v2); source conflict with 50M card retained

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: 2dc37b86e16a6970fbc731751f7719d9f676f7f9
Retrieved: 2026-09-16T19:46:19.364532+00:00

source checked

automated source review · 2026-09-16

Audit details

Inspected pinned official documentation, relevant implementation files and named primary-paper sections. Claims are limited to those artifacts. Remaining field extraction and identity conflicts are explicit; no new performance claims, model runs or human review are implied.

Field: attributes.profile.diagram.caption

Source artifact SHA-256: ab16d582de98652526b5cebb120eec969328f9db29dc741826bcd81c397e0672

Hash scope: SHA-256 of retrieved original artifact bytes

Format: original_artifact

Inspected artifact

Diagram caption
Conceptual summary of the documented data flow; optional inputs and configured downstream stages must be reported for a reproducible evaluation.
Individual claims
InstaDeepAI/nucleotide-transformer-v2-50m-multi-species: config.json

Original source ↗

README.md: Model Summary, Training data and licence metadata; config.json; Nucleotide Transformer paper Methods: Architecture and Training (v2); source conflict with 50M card retained

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: 81b29e5786726d891dbf929404ef20adca5b36f1
Retrieved: 2026-09-16T19:46:20.607357+00:00

source checked

automated source review · 2026-09-16

Audit details

Inspected pinned official documentation, relevant implementation files and named primary-paper sections. Claims are limited to those artifacts. Remaining field extraction and identity conflicts are explicit; no new performance claims, model runs or human review are implied.

Field: attributes.profile.diagram.caption

Source artifact SHA-256: e20f497248c7cb264c7cd4582dbcfd52dc4cbf74a97fc711559b8c8f71c635db

Hash scope: SHA-256 of retrieved original artifact bytes

Format: original_artifact

Inspected artifact

Diagram caption
Conceptual summary of the documented data flow; optional inputs and configured downstream stages must be reported for a reproducible evaluation.
Individual claims
nt: Journal full-text XML

Original source ↗

README.md: Model Summary, Training data and licence metadata; config.json; Nucleotide Transformer paper Methods: Architecture and Training (v2); source conflict with 50M card retained

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: Retrieved page snapshot; no immutable publisher revision supplied
Retrieved: 2026-09-16T20:16:14.422628+00:00

source checked

automated source review · 2026-09-16

Audit details

Inspected pinned official documentation, relevant implementation files and named primary-paper sections. Claims are limited to those artifacts. Remaining field extraction and identity conflicts are explicit; no new performance claims, model runs or human review are implied.

Field: attributes.profile.diagram.caption

Source artifact SHA-256: 7c3b78a4f38ef053a08e466222e5662dd9711d91c535512d0b35d459d1fc7249

Hash scope: SHA-256 of retrieved original artifact bytes

Format: original_artifact

Inspected artifact

Diagram steps
  • DNA sequence
  • 6-mer tokenizer
  • NT-v2 encoder
  • Embeddings or task adaptation
Individual claims
instadeepai/nucleotide-transformer: README.md

Original source ↗

README.md: Model Summary, Training data and licence metadata; config.json; Nucleotide Transformer paper Methods: Architecture and Training (v2); source conflict with 50M card retained

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: 2dc37b86e16a6970fbc731751f7719d9f676f7f9
Retrieved: 2026-09-16T19:46:19.364532+00:00

source checked

automated source review · 2026-09-16

Audit details

Inspected pinned official documentation, relevant implementation files and named primary-paper sections. Claims are limited to those artifacts. Remaining field extraction and identity conflicts are explicit; no new performance claims, model runs or human review are implied.

Field: attributes.profile.diagram.steps

Source artifact SHA-256: 9f51bbb20c4c5c36e77fb03ca1c5c36236e287c48a1ee31f53150545d421ec25

Hash scope: SHA-256 of retrieved original artifact bytes

Format: original_artifact

Inspected artifact

Diagram steps
  • DNA sequence
  • 6-mer tokenizer
  • NT-v2 encoder
  • Embeddings or task adaptation
Individual claims
instadeepai/nucleotide-transformer: docs/segment_nt.md

Original source ↗

README.md: Model Summary, Training data and licence metadata; config.json; Nucleotide Transformer paper Methods: Architecture and Training (v2); source conflict with 50M card retained

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: 2dc37b86e16a6970fbc731751f7719d9f676f7f9
Retrieved: 2026-09-16T19:46:19.364532+00:00

source checked

automated source review · 2026-09-16

Audit details

Inspected pinned official documentation, relevant implementation files and named primary-paper sections. Claims are limited to those artifacts. Remaining field extraction and identity conflicts are explicit; no new performance claims, model runs or human review are implied.

Field: attributes.profile.diagram.steps

Source artifact SHA-256: 8eec4580ba64ab944fb9b42674be70ffe793135f3503cce8640f5b08f8290f7a

Hash scope: SHA-256 of retrieved original artifact bytes

Format: original_artifact

Inspected artifact

Diagram steps
  • DNA sequence
  • 6-mer tokenizer
  • NT-v2 encoder
  • Embeddings or task adaptation
Individual claims
InstaDeepAI/nucleotide-transformer-v2-50m-multi-species: README.md

Original source ↗

README.md: Model Summary, Training data and licence metadata; config.json; Nucleotide Transformer paper Methods: Architecture and Training (v2); source conflict with 50M card retained

Shared locator for this statement’s cited sources; not a separate locator for each citation.

Version: 81b29e5786726d891dbf929404ef20adca5b36f1
Retrieved: 2026-09-16T19:46:20.607357+00:00

source checked

automated source review · 2026-09-16

Audit details

Inspected pinned official documentation, relevant implementation files and named primary-paper sections. Claims are limited to those artifacts. Remaining field extraction and identity conflicts are explicit; no new performance claims, model runs or human review are implied.

Field: attributes.profile.diagram.steps

Source artifact SHA-256: e526d7b98f106bc2ca9ba73fa166ff5fd62853812e9757a6692eeedf25e42923

Hash scope: SHA-256 of retrieved original artifact bytes

Format: original_artifact

Inspected artifact

Sources and history

View linked audit checks and correction history

Release 2026-09-23-2b89723c6dd9 · Record review: discovered

9 source records and release historyDownload this release
Technical metadata and extraction receipts

Stable ID: catalog-model-nt-v2

areas
dna-genomes
method types
foundation model
entity level
family
version
50M multi-species
reported name
Nucleotide Transformer v2
access
Public checkpoint.
method type
foundation model
historical missing metadata
checkpoint revision: not_yet_extracted; training data: not_yet_extracted; licence: not_yet_extracted
metadata review scope
historical_missing_metadata preserves the original discovery state. Current descriptive evidence and missingness are recorded in profile.facts; numerical-result review is separate.
entity classification
review date: 2026-09-17; rationale: The cited profile describes a named learned biological predictor or representation model/family. Preserve this identity separately from task-specific fitting, individual checkpoints, pipelines and hosted access.; source ids: evidence-official-4abb9affb1a9e5438c91; evidence-official-6b79ddfdd330693bf4fb; evidence-official-18d8f4d5fc3f6616922d; evidence-official-657e83427ab59f3aec83; evidence-official-56f02d45976d011d80aa; evidence-official-4920952f9b3c4b8909a0; evidence-official-aadbeb0f10ec55d34f5f; source locator: README.md: Model Summary, Training data and licence metadata; config.json; Nucleotide Transformer paper Methods: Architecture and Training (v2); source conflict with 50M card retained; ambiguities: None recorded
Related records

Suggest a correction