rewire.it
Configuration

SegmentNT-3kb (NTv1 human; 2.5B)

SegmentNT labels genomic elements at individual nucleotide positions using a pretrained DNA backbone.

Sources (7)instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; InstaDeepAI/segment_nt: README.md; InstaDeepAI/segment_nt: config.json; segmentnt: Journal full-text XML · SegmentNT paper Methods: SegmentNT architecture, Model ablations, Human genomic elements and Multispecies training; official model-card Training data and licence metadata

14 evaluations · 28 metric rows

How it worksSegmentNT workflow
SegmentNT workflow1. DNA sequence. Then: 2. NT backbone with position rescaling. Then: 3. 1D U-Net head. Then: 4. Per-base element predictionsSegmentNT workflow1. DNA sequence. Then: 2. NT backbone with position rescaling. Then: 3. 1D U-Net head. Then: 4. Per-base element predictionsSegmentNT workflow1. DNA sequence. Then: 2. NT backbone with position rescaling. Then: 3. 1D U-Net head. Then: 4. Per-base element predictions

Conceptual summary of the documented data flow; optional inputs and configured downstream stages must be reported for a reproducible evaluation.

Sources (7)instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; InstaDeepAI/segment_nt: README.md; InstaDeepAI/segment_nt: config.json; segmentnt: Journal full-text XML · SegmentNT paper Methods: SegmentNT architecture, Model ablations, Human genomic elements and Multispecies training; official model-card Training data and licence metadata

Overview

Model type

DNA encoder with nucleotide-level segmentation head

Sources (7)instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; InstaDeepAI/segment_nt: README.md; InstaDeepAI/segment_nt: config.json; segmentnt: Journal full-text XML · SegmentNT paper Methods: SegmentNT architecture, Model ablations, Human genomic elements and Multispecies training; official model-card Training data and licence metadata

Inputs

DNA sequences without N bases, tokenized into 6-mers under the documented length constraints.

Sources (7)instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; InstaDeepAI/segment_nt: README.md; InstaDeepAI/segment_nt: config.json; segmentnt: Journal full-text XML · SegmentNT paper Methods: SegmentNT architecture, Model ablations, Human genomic elements and Multispecies training; official model-card Training data and licence metadata

Outputs

Per-base probabilities for 14 genomic-element classes.

Sources (7)instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; InstaDeepAI/segment_nt: README.md; InstaDeepAI/segment_nt: config.json; segmentnt: Journal full-text XML · SegmentNT paper Methods: SegmentNT architecture, Model ablations, Human genomic elements and Multispecies training; official model-card Training data and licence metadata

Access

Official project documentation and implementation: https://github.com/instadeepai/nucleotide-transformer

Sources (7)instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; InstaDeepAI/segment_nt: README.md; InstaDeepAI/segment_nt: config.json; segmentnt: Journal full-text XML · SegmentNT paper Methods: SegmentNT architecture, Model ablations, Human genomic elements and Multispecies training; official model-card Training data and licence metadata

Source reviewed · Automated source review, 2026-09-16. All specifications and missing details

Evaluations and results

14 evaluations · 28 metric rows. Different protocols are not a single leaderboard.

Filter evaluations

Applied filters: All linked evaluations

Exact evaluated configurations and original reported results
Tested configurationProtocol and datasetFindingEvidence and details
Configuration: SegmentNT-3kb (NTv1 human; 2.5B)Protocol: SegmentNT human genome annotation 3UTR auPRC: 3UTR: per-nucleotide annotation (auPRC)
Dataset subset: Human genome 3UTR test chromosomes 20 and 21 (SegmentNT human genome annotation split)
0.36 (± 0.014) auprc
dimensionless · higher

Uncertainty: type: standard_deviation; value: 0.014

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · source checked
Methods, coverage and source

SegmentNT-3kb (NTv1 human; 2.5B) on SegmentNT human genome annotation 3UTR auPRC: 3UTR: per-nucleotide annotation (auPRC)

Test chromosomes 20 and 21; validation chromosome 22; training remaining chromosomes. Test chunks with genes homologous to train/validation genes excluded using Ensembl BioMart accessed 2024-05-08; homologous distal regulatory elements not excluded. Ten test-set samplings with sliding windows beginning at different genomic starting positions. Mean plus reported standard deviation; not ten training seeds. Best validation checkpoint by average MCC across 14 elements. Per-nucleotide predictions pooled across test sequences separately for each genomic element. MCC and auPRC are separate metrics, not cross-element or cross-table pooled comparisons.

Aggregation: Not reported

SegmentNT supplementary information: complete Tables 2 and 3 · Supplementary Table 3; PDF page 5 (printed 4); block 1; data row 12; model SegmentNT-3kb (NTv1 human; 2.5B); column 3UTR
Configuration: SegmentNT-3kb (NTv1 human; 2.5B)Protocol: SegmentNT human genome annotation 3UTR MCC: 3UTR: per-nucleotide annotation (MCC)
Dataset subset: Human genome 3UTR test chromosomes 20 and 21 (SegmentNT human genome annotation split)
0.42 (± 0.015) mcc
dimensionless · higher

Uncertainty: type: standard_deviation; value: 0.015

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · source checked
Methods, coverage and source

SegmentNT-3kb (NTv1 human; 2.5B) on SegmentNT human genome annotation 3UTR MCC: 3UTR: per-nucleotide annotation (MCC)

Test chromosomes 20 and 21; validation chromosome 22; training remaining chromosomes. Test chunks with genes homologous to train/validation genes excluded using Ensembl BioMart accessed 2024-05-08; homologous distal regulatory elements not excluded. Ten test-set samplings with sliding windows beginning at different genomic starting positions. Mean plus reported standard deviation; not ten training seeds. Best validation checkpoint by average MCC across 14 elements. Per-nucleotide predictions pooled across test sequences separately for each genomic element. MCC and auPRC are separate metrics, not cross-element or cross-table pooled comparisons.

Aggregation: Not reported

SegmentNT supplementary information: complete Tables 2 and 3 · Supplementary Table 2; PDF page 4 (printed 3); block 1; data row 12; model SegmentNT-3kb (NTv1 human; 2.5B); column 3UTR
Configuration: SegmentNT-3kb (NTv1 human; 2.5B)Protocol: SegmentNT human genome annotation 5UTR auPRC: 5UTR: per-nucleotide annotation (auPRC)
Dataset subset: Human genome 5UTR test chromosomes 20 and 21 (SegmentNT human genome annotation split)
0.31 (± 0.005) auprc
dimensionless · higher

Uncertainty: type: standard_deviation; value: 0.005

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · source checked
Methods, coverage and source

SegmentNT-3kb (NTv1 human; 2.5B) on SegmentNT human genome annotation 5UTR auPRC: 5UTR: per-nucleotide annotation (auPRC)

Test chromosomes 20 and 21; validation chromosome 22; training remaining chromosomes. Test chunks with genes homologous to train/validation genes excluded using Ensembl BioMart accessed 2024-05-08; homologous distal regulatory elements not excluded. Ten test-set samplings with sliding windows beginning at different genomic starting positions. Mean plus reported standard deviation; not ten training seeds. Best validation checkpoint by average MCC across 14 elements. Per-nucleotide predictions pooled across test sequences separately for each genomic element. MCC and auPRC are separate metrics, not cross-element or cross-table pooled comparisons.

Aggregation: Not reported

SegmentNT supplementary information: complete Tables 2 and 3 · Supplementary Table 3; PDF page 5 (printed 4); block 1; data row 12; model SegmentNT-3kb (NTv1 human; 2.5B); column 5UTR
Configuration: SegmentNT-3kb (NTv1 human; 2.5B)Protocol: SegmentNT human genome annotation 5UTR MCC: 5UTR: per-nucleotide annotation (MCC)
Dataset subset: Human genome 5UTR test chromosomes 20 and 21 (SegmentNT human genome annotation split)
0.38 (± 0.006) mcc
dimensionless · higher

Uncertainty: type: standard_deviation; value: 0.006

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · source checked
Methods, coverage and source

SegmentNT-3kb (NTv1 human; 2.5B) on SegmentNT human genome annotation 5UTR MCC: 5UTR: per-nucleotide annotation (MCC)

Test chromosomes 20 and 21; validation chromosome 22; training remaining chromosomes. Test chunks with genes homologous to train/validation genes excluded using Ensembl BioMart accessed 2024-05-08; homologous distal regulatory elements not excluded. Ten test-set samplings with sliding windows beginning at different genomic starting positions. Mean plus reported standard deviation; not ten training seeds. Best validation checkpoint by average MCC across 14 elements. Per-nucleotide predictions pooled across test sequences separately for each genomic element. MCC and auPRC are separate metrics, not cross-element or cross-table pooled comparisons.

Aggregation: Not reported

SegmentNT supplementary information: complete Tables 2 and 3 · Supplementary Table 2; PDF page 4 (printed 3); block 1; data row 12; model SegmentNT-3kb (NTv1 human; 2.5B); column 5UTR
Configuration: SegmentNT-3kb (NTv1 human; 2.5B)Protocol: SegmentNT human genome annotation CTCF-bound auPRC: CTCF-bound: per-nucleotide annotation (auPRC)
Dataset subset: Human genome CTCF-bound test chromosomes 20 and 21 (SegmentNT human genome annotation split)
0.04 (± 0.002) auprc
dimensionless · higher

Uncertainty: type: standard_deviation; value: 0.002

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · source checked
Methods, coverage and source

SegmentNT-3kb (NTv1 human; 2.5B) on SegmentNT human genome annotation CTCF-bound auPRC: CTCF-bound: per-nucleotide annotation (auPRC)

Test chromosomes 20 and 21; validation chromosome 22; training remaining chromosomes. Test chunks with genes homologous to train/validation genes excluded using Ensembl BioMart accessed 2024-05-08; homologous distal regulatory elements not excluded. Ten test-set samplings with sliding windows beginning at different genomic starting positions. Mean plus reported standard deviation; not ten training seeds. Best validation checkpoint by average MCC across 14 elements. Per-nucleotide predictions pooled across test sequences separately for each genomic element. MCC and auPRC are separate metrics, not cross-element or cross-table pooled comparisons.

Aggregation: Not reported

SegmentNT supplementary information: complete Tables 2 and 3 · Supplementary Table 3; PDF page 5 (printed 4); block 1; data row 12; model SegmentNT-3kb (NTv1 human; 2.5B); column CTCF-bound
Configuration: SegmentNT-3kb (NTv1 human; 2.5B)Protocol: SegmentNT human genome annotation CTCF-bound MCC: CTCF-bound: per-nucleotide annotation (MCC)
Dataset subset: Human genome CTCF-bound test chromosomes 20 and 21 (SegmentNT human genome annotation split)
0.08 (± 0.006) mcc
dimensionless · higher

Uncertainty: type: standard_deviation; value: 0.006

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · source checked
Methods, coverage and source

SegmentNT-3kb (NTv1 human; 2.5B) on SegmentNT human genome annotation CTCF-bound MCC: CTCF-bound: per-nucleotide annotation (MCC)

Test chromosomes 20 and 21; validation chromosome 22; training remaining chromosomes. Test chunks with genes homologous to train/validation genes excluded using Ensembl BioMart accessed 2024-05-08; homologous distal regulatory elements not excluded. Ten test-set samplings with sliding windows beginning at different genomic starting positions. Mean plus reported standard deviation; not ten training seeds. Best validation checkpoint by average MCC across 14 elements. Per-nucleotide predictions pooled across test sequences separately for each genomic element. MCC and auPRC are separate metrics, not cross-element or cross-table pooled comparisons.

Aggregation: Not reported

SegmentNT supplementary information: complete Tables 2 and 3 · Supplementary Table 2; PDF page 4 (printed 3); block 1; data row 12; model SegmentNT-3kb (NTv1 human; 2.5B); column CTCF-bound
Configuration: SegmentNT-3kb (NTv1 human; 2.5B)Protocol: SegmentNT human genome annotation enhancer tissue-invariant auPRC: enhancer tissue-invariant: per-nucleotide annotation (auPRC)
Dataset subset: Human genome enhancer tissue-invariant test chromosomes 20 and 21 (SegmentNT human genome annotation split)
0.10 (± 0.003) auprc
dimensionless · higher

Uncertainty: type: standard_deviation; value: 0.003

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · source checked
Methods, coverage and source

SegmentNT-3kb (NTv1 human; 2.5B) on SegmentNT human genome annotation enhancer tissue-invariant auPRC: enhancer tissue-invariant: per-nucleotide annotation (auPRC)

Test chromosomes 20 and 21; validation chromosome 22; training remaining chromosomes. Test chunks with genes homologous to train/validation genes excluded using Ensembl BioMart accessed 2024-05-08; homologous distal regulatory elements not excluded. Ten test-set samplings with sliding windows beginning at different genomic starting positions. Mean plus reported standard deviation; not ten training seeds. Best validation checkpoint by average MCC across 14 elements. Per-nucleotide predictions pooled across test sequences separately for each genomic element. MCC and auPRC are separate metrics, not cross-element or cross-table pooled comparisons.

Aggregation: Not reported

SegmentNT supplementary information: complete Tables 2 and 3 · Supplementary Table 3; PDF page 5 (printed 4); block 1; data row 12; model SegmentNT-3kb (NTv1 human; 2.5B); column enhancer tissue-invariant
Configuration: SegmentNT-3kb (NTv1 human; 2.5B)Protocol: SegmentNT human genome annotation enhancer tissue-invariant MCC: enhancer tissue-invariant: per-nucleotide annotation (MCC)
Dataset subset: Human genome enhancer tissue-invariant test chromosomes 20 and 21 (SegmentNT human genome annotation split)
0.13 (± 0.006) mcc
dimensionless · higher

Uncertainty: type: standard_deviation; value: 0.006

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · source checked
Methods, coverage and source

SegmentNT-3kb (NTv1 human; 2.5B) on SegmentNT human genome annotation enhancer tissue-invariant MCC: enhancer tissue-invariant: per-nucleotide annotation (MCC)

Test chromosomes 20 and 21; validation chromosome 22; training remaining chromosomes. Test chunks with genes homologous to train/validation genes excluded using Ensembl BioMart accessed 2024-05-08; homologous distal regulatory elements not excluded. Ten test-set samplings with sliding windows beginning at different genomic starting positions. Mean plus reported standard deviation; not ten training seeds. Best validation checkpoint by average MCC across 14 elements. Per-nucleotide predictions pooled across test sequences separately for each genomic element. MCC and auPRC are separate metrics, not cross-element or cross-table pooled comparisons.

Aggregation: Not reported

SegmentNT supplementary information: complete Tables 2 and 3 · Supplementary Table 2; PDF page 4 (printed 3); block 1; data row 12; model SegmentNT-3kb (NTv1 human; 2.5B); column enhancer tissue-invariant
Configuration: SegmentNT-3kb (NTv1 human; 2.5B)Protocol: SegmentNT human genome annotation enhancer tissue-specific auPRC: enhancer tissue-specific: per-nucleotide annotation (auPRC)
Dataset subset: Human genome enhancer tissue-specific test chromosomes 20 and 21 (SegmentNT human genome annotation split)
0.30 (± 0.002) auprc
dimensionless · higher

Uncertainty: type: standard_deviation; value: 0.002

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · source checked
Methods, coverage and source

SegmentNT-3kb (NTv1 human; 2.5B) on SegmentNT human genome annotation enhancer tissue-specific auPRC: enhancer tissue-specific: per-nucleotide annotation (auPRC)

Test chromosomes 20 and 21; validation chromosome 22; training remaining chromosomes. Test chunks with genes homologous to train/validation genes excluded using Ensembl BioMart accessed 2024-05-08; homologous distal regulatory elements not excluded. Ten test-set samplings with sliding windows beginning at different genomic starting positions. Mean plus reported standard deviation; not ten training seeds. Best validation checkpoint by average MCC across 14 elements. Per-nucleotide predictions pooled across test sequences separately for each genomic element. MCC and auPRC are separate metrics, not cross-element or cross-table pooled comparisons.

Aggregation: Not reported

SegmentNT supplementary information: complete Tables 2 and 3 · Supplementary Table 3; PDF page 5 (printed 4); block 1; data row 12; model SegmentNT-3kb (NTv1 human; 2.5B); column enhancer tissue-specific
Configuration: SegmentNT-3kb (NTv1 human; 2.5B)Protocol: SegmentNT human genome annotation enhancer tissue-specific MCC: enhancer tissue-specific: per-nucleotide annotation (MCC)
Dataset subset: Human genome enhancer tissue-specific test chromosomes 20 and 21 (SegmentNT human genome annotation split)
0.20 (± 0.002) mcc
dimensionless · higher

Uncertainty: type: standard_deviation; value: 0.002

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · source checked
Methods, coverage and source

SegmentNT-3kb (NTv1 human; 2.5B) on SegmentNT human genome annotation enhancer tissue-specific MCC: enhancer tissue-specific: per-nucleotide annotation (MCC)

Test chromosomes 20 and 21; validation chromosome 22; training remaining chromosomes. Test chunks with genes homologous to train/validation genes excluded using Ensembl BioMart accessed 2024-05-08; homologous distal regulatory elements not excluded. Ten test-set samplings with sliding windows beginning at different genomic starting positions. Mean plus reported standard deviation; not ten training seeds. Best validation checkpoint by average MCC across 14 elements. Per-nucleotide predictions pooled across test sequences separately for each genomic element. MCC and auPRC are separate metrics, not cross-element or cross-table pooled comparisons.

Aggregation: Not reported

SegmentNT supplementary information: complete Tables 2 and 3 · Supplementary Table 2; PDF page 4 (printed 3); block 1; data row 12; model SegmentNT-3kb (NTv1 human; 2.5B); column enhancer tissue-specific
Configuration: SegmentNT-3kb (NTv1 human; 2.5B)Protocol: SegmentNT human genome annotation exon auPRC: exon: per-nucleotide annotation (auPRC)
Dataset subset: Human genome exon test chromosomes 20 and 21 (SegmentNT human genome annotation split)
0.45 (± 0.005) auprc
dimensionless · higher

Uncertainty: type: standard_deviation; value: 0.005

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · source checked
Methods, coverage and source

SegmentNT-3kb (NTv1 human; 2.5B) on SegmentNT human genome annotation exon auPRC: exon: per-nucleotide annotation (auPRC)

Test chromosomes 20 and 21; validation chromosome 22; training remaining chromosomes. Test chunks with genes homologous to train/validation genes excluded using Ensembl BioMart accessed 2024-05-08; homologous distal regulatory elements not excluded. Ten test-set samplings with sliding windows beginning at different genomic starting positions. Mean plus reported standard deviation; not ten training seeds. Best validation checkpoint by average MCC across 14 elements. Per-nucleotide predictions pooled across test sequences separately for each genomic element. MCC and auPRC are separate metrics, not cross-element or cross-table pooled comparisons.

Aggregation: Not reported

SegmentNT supplementary information: complete Tables 2 and 3 · Supplementary Table 3; PDF page 5 (printed 4); block 1; data row 12; model SegmentNT-3kb (NTv1 human; 2.5B); column exon
Configuration: SegmentNT-3kb (NTv1 human; 2.5B)Protocol: SegmentNT human genome annotation exon MCC: exon: per-nucleotide annotation (MCC)
Dataset subset: Human genome exon test chromosomes 20 and 21 (SegmentNT human genome annotation split)
0.48 (± 0.006) mcc
dimensionless · higher

Uncertainty: type: standard_deviation; value: 0.006

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · source checked
Methods, coverage and source

SegmentNT-3kb (NTv1 human; 2.5B) on SegmentNT human genome annotation exon MCC: exon: per-nucleotide annotation (MCC)

Test chromosomes 20 and 21; validation chromosome 22; training remaining chromosomes. Test chunks with genes homologous to train/validation genes excluded using Ensembl BioMart accessed 2024-05-08; homologous distal regulatory elements not excluded. Ten test-set samplings with sliding windows beginning at different genomic starting positions. Mean plus reported standard deviation; not ten training seeds. Best validation checkpoint by average MCC across 14 elements. Per-nucleotide predictions pooled across test sequences separately for each genomic element. MCC and auPRC are separate metrics, not cross-element or cross-table pooled comparisons.

Aggregation: Not reported

SegmentNT supplementary information: complete Tables 2 and 3 · Supplementary Table 2; PDF page 4 (printed 3); block 1; data row 12; model SegmentNT-3kb (NTv1 human; 2.5B); column exon
Configuration: SegmentNT-3kb (NTv1 human; 2.5B)Protocol: SegmentNT human genome annotation intron auPRC: intron: per-nucleotide annotation (auPRC)
Dataset subset: Human genome intron test chromosomes 20 and 21 (SegmentNT human genome annotation split)
0.69 (± 0.001) auprc
dimensionless · higher

Uncertainty: type: standard_deviation; value: 0.001

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · source checked
Methods, coverage and source

SegmentNT-3kb (NTv1 human; 2.5B) on SegmentNT human genome annotation intron auPRC: intron: per-nucleotide annotation (auPRC)

Test chromosomes 20 and 21; validation chromosome 22; training remaining chromosomes. Test chunks with genes homologous to train/validation genes excluded using Ensembl BioMart accessed 2024-05-08; homologous distal regulatory elements not excluded. Ten test-set samplings with sliding windows beginning at different genomic starting positions. Mean plus reported standard deviation; not ten training seeds. Best validation checkpoint by average MCC across 14 elements. Per-nucleotide predictions pooled across test sequences separately for each genomic element. MCC and auPRC are separate metrics, not cross-element or cross-table pooled comparisons.

Aggregation: Not reported

SegmentNT supplementary information: complete Tables 2 and 3 · Supplementary Table 3; PDF page 5 (printed 4); block 1; data row 12; model SegmentNT-3kb (NTv1 human; 2.5B); column intron
Configuration: SegmentNT-3kb (NTv1 human; 2.5B)Protocol: SegmentNT human genome annotation intron MCC: intron: per-nucleotide annotation (MCC)
Dataset subset: Human genome intron test chromosomes 20 and 21 (SegmentNT human genome annotation split)
0.29 (± 0.003) mcc
dimensionless · higher

Uncertainty: type: standard_deviation; value: 0.003

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · source checked
Methods, coverage and source

SegmentNT-3kb (NTv1 human; 2.5B) on SegmentNT human genome annotation intron MCC: intron: per-nucleotide annotation (MCC)

Test chromosomes 20 and 21; validation chromosome 22; training remaining chromosomes. Test chunks with genes homologous to train/validation genes excluded using Ensembl BioMart accessed 2024-05-08; homologous distal regulatory elements not excluded. Ten test-set samplings with sliding windows beginning at different genomic starting positions. Mean plus reported standard deviation; not ten training seeds. Best validation checkpoint by average MCC across 14 elements. Per-nucleotide predictions pooled across test sequences separately for each genomic element. MCC and auPRC are separate metrics, not cross-element or cross-table pooled comparisons.

Aggregation: Not reported

SegmentNT supplementary information: complete Tables 2 and 3 · Supplementary Table 2; PDF page 4 (printed 3); block 1; data row 12; model SegmentNT-3kb (NTv1 human; 2.5B); column intron
Configuration: SegmentNT-3kb (NTv1 human; 2.5B)Protocol: SegmentNT human genome annotation lncRNA auPRC: lncRNA: per-nucleotide annotation (auPRC)
Dataset subset: Human genome lncRNA test chromosomes 20 and 21 (SegmentNT human genome annotation split)
0.21 (± 0.002) auprc
dimensionless · higher

Uncertainty: type: standard_deviation; value: 0.002

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · source checked
Methods, coverage and source

SegmentNT-3kb (NTv1 human; 2.5B) on SegmentNT human genome annotation lncRNA auPRC: lncRNA: per-nucleotide annotation (auPRC)

Test chromosomes 20 and 21; validation chromosome 22; training remaining chromosomes. Test chunks with genes homologous to train/validation genes excluded using Ensembl BioMart accessed 2024-05-08; homologous distal regulatory elements not excluded. Ten test-set samplings with sliding windows beginning at different genomic starting positions. Mean plus reported standard deviation; not ten training seeds. Best validation checkpoint by average MCC across 14 elements. Per-nucleotide predictions pooled across test sequences separately for each genomic element. MCC and auPRC are separate metrics, not cross-element or cross-table pooled comparisons.

Aggregation: Not reported

SegmentNT supplementary information: complete Tables 2 and 3 · Supplementary Table 3; PDF page 5 (printed 4); block 2; data row 12; model SegmentNT-3kb (NTv1 human; 2.5B); column lncRNA
Configuration: SegmentNT-3kb (NTv1 human; 2.5B)Protocol: SegmentNT human genome annotation lncRNA MCC: lncRNA: per-nucleotide annotation (MCC)
Dataset subset: Human genome lncRNA test chromosomes 20 and 21 (SegmentNT human genome annotation split)
0.04 (± 0.004) mcc
dimensionless · higher

Uncertainty: type: standard_deviation; value: 0.004

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · source checked
Methods, coverage and source

SegmentNT-3kb (NTv1 human; 2.5B) on SegmentNT human genome annotation lncRNA MCC: lncRNA: per-nucleotide annotation (MCC)

Test chromosomes 20 and 21; validation chromosome 22; training remaining chromosomes. Test chunks with genes homologous to train/validation genes excluded using Ensembl BioMart accessed 2024-05-08; homologous distal regulatory elements not excluded. Ten test-set samplings with sliding windows beginning at different genomic starting positions. Mean plus reported standard deviation; not ten training seeds. Best validation checkpoint by average MCC across 14 elements. Per-nucleotide predictions pooled across test sequences separately for each genomic element. MCC and auPRC are separate metrics, not cross-element or cross-table pooled comparisons.

Aggregation: Not reported

SegmentNT supplementary information: complete Tables 2 and 3 · Supplementary Table 2; PDF page 4 (printed 3); block 2; data row 12; model SegmentNT-3kb (NTv1 human; 2.5B); column lncRNA
Configuration: SegmentNT-3kb (NTv1 human; 2.5B)Protocol: SegmentNT human genome annotation polyA signal auPRC: polyA signal: per-nucleotide annotation (auPRC)
Dataset subset: Human genome polyA signal test chromosomes 20 and 21 (SegmentNT human genome annotation split)
0.19 (± 0.009) auprc
dimensionless · higher

Uncertainty: type: standard_deviation; value: 0.009

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · source checked
Methods, coverage and source

SegmentNT-3kb (NTv1 human; 2.5B) on SegmentNT human genome annotation polyA signal auPRC: polyA signal: per-nucleotide annotation (auPRC)

Test chromosomes 20 and 21; validation chromosome 22; training remaining chromosomes. Test chunks with genes homologous to train/validation genes excluded using Ensembl BioMart accessed 2024-05-08; homologous distal regulatory elements not excluded. Ten test-set samplings with sliding windows beginning at different genomic starting positions. Mean plus reported standard deviation; not ten training seeds. Best validation checkpoint by average MCC across 14 elements. Per-nucleotide predictions pooled across test sequences separately for each genomic element. MCC and auPRC are separate metrics, not cross-element or cross-table pooled comparisons.

Aggregation: Not reported

SegmentNT supplementary information: complete Tables 2 and 3 · Supplementary Table 3; PDF page 5 (printed 4); block 2; data row 12; model SegmentNT-3kb (NTv1 human; 2.5B); column polyA signal
Configuration: SegmentNT-3kb (NTv1 human; 2.5B)Protocol: SegmentNT human genome annotation polyA signal MCC: polyA signal: per-nucleotide annotation (MCC)
Dataset subset: Human genome polyA signal test chromosomes 20 and 21 (SegmentNT human genome annotation split)
0.28 (± 0.006) mcc
dimensionless · higher

Uncertainty: type: standard_deviation; value: 0.006

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · source checked
Methods, coverage and source

SegmentNT-3kb (NTv1 human; 2.5B) on SegmentNT human genome annotation polyA signal MCC: polyA signal: per-nucleotide annotation (MCC)

Test chromosomes 20 and 21; validation chromosome 22; training remaining chromosomes. Test chunks with genes homologous to train/validation genes excluded using Ensembl BioMart accessed 2024-05-08; homologous distal regulatory elements not excluded. Ten test-set samplings with sliding windows beginning at different genomic starting positions. Mean plus reported standard deviation; not ten training seeds. Best validation checkpoint by average MCC across 14 elements. Per-nucleotide predictions pooled across test sequences separately for each genomic element. MCC and auPRC are separate metrics, not cross-element or cross-table pooled comparisons.

Aggregation: Not reported

SegmentNT supplementary information: complete Tables 2 and 3 · Supplementary Table 2; PDF page 4 (printed 3); block 2; data row 12; model SegmentNT-3kb (NTv1 human; 2.5B); column polyA signal
Configuration: SegmentNT-3kb (NTv1 human; 2.5B)Protocol: SegmentNT human genome annotation promoter tissue-invariant auPRC: promoter tissue-invariant: per-nucleotide annotation (auPRC)
Dataset subset: Human genome promoter tissue-invariant test chromosomes 20 and 21 (SegmentNT human genome annotation split)
0.55 (± 0.007) auprc
dimensionless · higher

Uncertainty: type: standard_deviation; value: 0.007

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · source checked
Methods, coverage and source

SegmentNT-3kb (NTv1 human; 2.5B) on SegmentNT human genome annotation promoter tissue-invariant auPRC: promoter tissue-invariant: per-nucleotide annotation (auPRC)

Test chromosomes 20 and 21; validation chromosome 22; training remaining chromosomes. Test chunks with genes homologous to train/validation genes excluded using Ensembl BioMart accessed 2024-05-08; homologous distal regulatory elements not excluded. Ten test-set samplings with sliding windows beginning at different genomic starting positions. Mean plus reported standard deviation; not ten training seeds. Best validation checkpoint by average MCC across 14 elements. Per-nucleotide predictions pooled across test sequences separately for each genomic element. MCC and auPRC are separate metrics, not cross-element or cross-table pooled comparisons.

Aggregation: Not reported

SegmentNT supplementary information: complete Tables 2 and 3 · Supplementary Table 3; PDF page 5 (printed 4); block 2; data row 12; model SegmentNT-3kb (NTv1 human; 2.5B); column promoter tissue-invariant
Configuration: SegmentNT-3kb (NTv1 human; 2.5B)Protocol: SegmentNT human genome annotation promoter tissue-invariant MCC: promoter tissue-invariant: per-nucleotide annotation (MCC)
Dataset subset: Human genome promoter tissue-invariant test chromosomes 20 and 21 (SegmentNT human genome annotation split)
0.56 (± 0.007) mcc
dimensionless · higher

Uncertainty: type: standard_deviation; value: 0.007

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · source checked
Methods, coverage and source

SegmentNT-3kb (NTv1 human; 2.5B) on SegmentNT human genome annotation promoter tissue-invariant MCC: promoter tissue-invariant: per-nucleotide annotation (MCC)

Test chromosomes 20 and 21; validation chromosome 22; training remaining chromosomes. Test chunks with genes homologous to train/validation genes excluded using Ensembl BioMart accessed 2024-05-08; homologous distal regulatory elements not excluded. Ten test-set samplings with sliding windows beginning at different genomic starting positions. Mean plus reported standard deviation; not ten training seeds. Best validation checkpoint by average MCC across 14 elements. Per-nucleotide predictions pooled across test sequences separately for each genomic element. MCC and auPRC are separate metrics, not cross-element or cross-table pooled comparisons.

Aggregation: Not reported

SegmentNT supplementary information: complete Tables 2 and 3 · Supplementary Table 2; PDF page 4 (printed 3); block 2; data row 12; model SegmentNT-3kb (NTv1 human; 2.5B); column promoter tissue-invariant
Configuration: SegmentNT-3kb (NTv1 human; 2.5B)Protocol: SegmentNT human genome annotation promoter tissue-specific auPRC: promoter tissue-specific: per-nucleotide annotation (auPRC)
Dataset subset: Human genome promoter tissue-specific test chromosomes 20 and 21 (SegmentNT human genome annotation split)
0.17 (± 0.006) auprc
dimensionless · higher

Uncertainty: type: standard_deviation; value: 0.006

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · source checked
Methods, coverage and source

SegmentNT-3kb (NTv1 human; 2.5B) on SegmentNT human genome annotation promoter tissue-specific auPRC: promoter tissue-specific: per-nucleotide annotation (auPRC)

Test chromosomes 20 and 21; validation chromosome 22; training remaining chromosomes. Test chunks with genes homologous to train/validation genes excluded using Ensembl BioMart accessed 2024-05-08; homologous distal regulatory elements not excluded. Ten test-set samplings with sliding windows beginning at different genomic starting positions. Mean plus reported standard deviation; not ten training seeds. Best validation checkpoint by average MCC across 14 elements. Per-nucleotide predictions pooled across test sequences separately for each genomic element. MCC and auPRC are separate metrics, not cross-element or cross-table pooled comparisons.

Aggregation: Not reported

SegmentNT supplementary information: complete Tables 2 and 3 · Supplementary Table 3; PDF page 5 (printed 4); block 2; data row 12; model SegmentNT-3kb (NTv1 human; 2.5B); column promoter tissue-specific
Configuration: SegmentNT-3kb (NTv1 human; 2.5B)Protocol: SegmentNT human genome annotation promoter tissue-specific MCC: promoter tissue-specific: per-nucleotide annotation (MCC)
Dataset subset: Human genome promoter tissue-specific test chromosomes 20 and 21 (SegmentNT human genome annotation split)
0.22 (± 0.010) mcc
dimensionless · higher

Uncertainty: type: standard_deviation; value: 0.010

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · source checked
Methods, coverage and source

SegmentNT-3kb (NTv1 human; 2.5B) on SegmentNT human genome annotation promoter tissue-specific MCC: promoter tissue-specific: per-nucleotide annotation (MCC)

Test chromosomes 20 and 21; validation chromosome 22; training remaining chromosomes. Test chunks with genes homologous to train/validation genes excluded using Ensembl BioMart accessed 2024-05-08; homologous distal regulatory elements not excluded. Ten test-set samplings with sliding windows beginning at different genomic starting positions. Mean plus reported standard deviation; not ten training seeds. Best validation checkpoint by average MCC across 14 elements. Per-nucleotide predictions pooled across test sequences separately for each genomic element. MCC and auPRC are separate metrics, not cross-element or cross-table pooled comparisons.

Aggregation: Not reported

SegmentNT supplementary information: complete Tables 2 and 3 · Supplementary Table 2; PDF page 4 (printed 3); block 2; data row 12; model SegmentNT-3kb (NTv1 human; 2.5B); column promoter tissue-specific
Configuration: SegmentNT-3kb (NTv1 human; 2.5B)Protocol: SegmentNT human genome annotation protein coding gene auPRC: protein coding gene: per-nucleotide annotation (auPRC)
Dataset subset: Human genome protein coding gene test chromosomes 20 and 21 (SegmentNT human genome annotation split)
0.69 (± 0.001) auprc
dimensionless · higher

Uncertainty: type: standard_deviation; value: 0.001

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · source checked
Methods, coverage and source

SegmentNT-3kb (NTv1 human; 2.5B) on SegmentNT human genome annotation protein coding gene auPRC: protein coding gene: per-nucleotide annotation (auPRC)

Test chromosomes 20 and 21; validation chromosome 22; training remaining chromosomes. Test chunks with genes homologous to train/validation genes excluded using Ensembl BioMart accessed 2024-05-08; homologous distal regulatory elements not excluded. Ten test-set samplings with sliding windows beginning at different genomic starting positions. Mean plus reported standard deviation; not ten training seeds. Best validation checkpoint by average MCC across 14 elements. Per-nucleotide predictions pooled across test sequences separately for each genomic element. MCC and auPRC are separate metrics, not cross-element or cross-table pooled comparisons.

Aggregation: Not reported

SegmentNT supplementary information: complete Tables 2 and 3 · Supplementary Table 3; PDF page 5 (printed 4); block 2; data row 12; model SegmentNT-3kb (NTv1 human; 2.5B); column protein coding gene
Configuration: SegmentNT-3kb (NTv1 human; 2.5B)Protocol: SegmentNT human genome annotation protein coding gene MCC: protein coding gene: per-nucleotide annotation (MCC)
Dataset subset: Human genome protein coding gene test chromosomes 20 and 21 (SegmentNT human genome annotation split)
0.41 (± 0.002) mcc
dimensionless · higher

Uncertainty: type: standard_deviation; value: 0.002

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · source checked
Methods, coverage and source

SegmentNT-3kb (NTv1 human; 2.5B) on SegmentNT human genome annotation protein coding gene MCC: protein coding gene: per-nucleotide annotation (MCC)

Test chromosomes 20 and 21; validation chromosome 22; training remaining chromosomes. Test chunks with genes homologous to train/validation genes excluded using Ensembl BioMart accessed 2024-05-08; homologous distal regulatory elements not excluded. Ten test-set samplings with sliding windows beginning at different genomic starting positions. Mean plus reported standard deviation; not ten training seeds. Best validation checkpoint by average MCC across 14 elements. Per-nucleotide predictions pooled across test sequences separately for each genomic element. MCC and auPRC are separate metrics, not cross-element or cross-table pooled comparisons.

Aggregation: Not reported

SegmentNT supplementary information: complete Tables 2 and 3 · Supplementary Table 2; PDF page 4 (printed 3); block 2; data row 12; model SegmentNT-3kb (NTv1 human; 2.5B); column protein coding gene
Configuration: SegmentNT-3kb (NTv1 human; 2.5B)Protocol: SegmentNT human genome annotation splice acceptor auPRC: splice acceptor: per-nucleotide annotation (auPRC)
Dataset subset: Human genome splice acceptor test chromosomes 20 and 21 (SegmentNT human genome annotation split)
0.57 (± 0.002) auprc
dimensionless · higher

Uncertainty: type: standard_deviation; value: 0.002

Coverage: Not reported scored / Not reported eligible

Author-reported evaluation · source checked
Methods, coverage and source

SegmentNT-3kb (NTv1 human; 2.5B) on SegmentNT human genome annotation splice acceptor auPRC: splice acceptor: per-nucleotide annotation (auPRC)

Test chromosomes 20 and 21; validation chromosome 22; training remaining chromosomes. Test chunks with genes homologous to train/validation genes excluded using Ensembl BioMart accessed 2024-05-08; homologous distal regulatory elements not excluded. Ten test-set samplings with sliding windows beginning at different genomic starting positions. Mean plus reported standard deviation; not ten training seeds. Best validation checkpoint by average MCC across 14 elements. Per-nucleotide predictions pooled across test sequences separately for each genomic element. MCC and auPRC are separate metrics, not cross-element or cross-table pooled comparisons.

Aggregation: Not reported

SegmentNT supplementary information: complete Tables 2 and 3 · Supplementary Table 3; PDF page 5 (printed 4); block 2; data row 12; model SegmentNT-3kb (NTv1 human; 2.5B); column splice acceptor

Source checking is not independent reproduction. Release 2026-09-23-2b89723c6dd9.

Use this model

How it works, versions and access

Related profile: SegmentNT. This page retains the exact record and its evaluation context.

Underlying model: Nucleotide Transformer. Results on this page belong to this configuration and its evaluated settings.

This configuration

Evaluated configuration as printed in Supplementary Tables 2 and 3; source reports training and checkpoint selection in Sec16–17.

record
SegmentNT-3kb (NTv1 human; 2.5B)
configuration
Not reported
entity type
Configuration

How it works

How it works

SegmentNT labels genomic elements at individual nucleotide positions using a pretrained DNA backbone. Nucleotide Transformer backbone with a one-dimensional U-Net segmentation head; YaRN rescales positions for longer inputs. The documented inputs are DNA sequences without N bases, tokenized into 6-mers under the documented length constraints. The output consists of per-base probabilities for 14 genomic-element classes.

Sources (7)instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; InstaDeepAI/segment_nt: README.md; InstaDeepAI/segment_nt: config.json; segmentnt: Journal full-text XML · SegmentNT paper Methods: SegmentNT architecture, Model ablations, Human genomic elements and Multispecies training; official model-card Training data and licence metadata
Versions and reproducibility

segment_nt and segment_nt_multi_species; SegmentEnformer and SegmentBorzoi are distinct pipelines. Trained on 30kb; inference up to 50kb requires the documented rescaling. Input token count also has divisibility constraints.

Sources (7)instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; InstaDeepAI/segment_nt: README.md; InstaDeepAI/segment_nt: config.json; segmentnt: Journal full-text XML · SegmentNT paper Methods: SegmentNT architecture, Model ablations, Human genomic elements and Multispecies training; official model-card Training data and licence metadata
Strengths, limitations and unresolved questions

Strengths and limitations

Strengths supported by sources

Limitations and conditions

Profile review details

Inspected pinned official documentation, relevant implementation files and named primary-paper sections. Claims are limited to those artifacts. Remaining field extraction and identity conflicts are explicit; no new performance claims, model runs or human review are implied.

Stable record: discovery-model-segmentnt

Specifications

Inputs, training, access and other details

Explanatory profile: source reviewed · Automated source review, 2026-09-16. Review applies to the cited claims; unresolved fields are listed below. Numerical results retain their own review status.

Inputs, outputs and configuration
PropertyDescription and evidence
Model typeDNA encoder with nucleotide-level segmentation head
Sources (7)instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; InstaDeepAI/segment_nt: README.md; InstaDeepAI/segment_nt: config.json; segmentnt: Journal full-text XML · SegmentNT paper Methods: SegmentNT architecture, Model ablations, Human genomic elements and Multispecies training; official model-card Training data and licence metadata
ArchitectureNucleotide Transformer backbone with a one-dimensional U-Net segmentation head; YaRN rescales positions for longer inputs.
Sources (7)instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; InstaDeepAI/segment_nt: README.md; InstaDeepAI/segment_nt: config.json; segmentnt: Journal full-text XML · SegmentNT paper Methods: SegmentNT architecture, Model ablations, Human genomic elements and Multispecies training; official model-card Training data and licence metadata
InputsDNA sequences without N bases, tokenized into 6-mers under the documented length constraints.
Sources (7)instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; InstaDeepAI/segment_nt: README.md; InstaDeepAI/segment_nt: config.json; segmentnt: Journal full-text XML · SegmentNT paper Methods: SegmentNT architecture, Model ablations, Human genomic elements and Multispecies training; official model-card Training data and licence metadata
OutputsPer-base probabilities for 14 genomic-element classes.
Sources (7)instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; InstaDeepAI/segment_nt: README.md; InstaDeepAI/segment_nt: config.json; segmentnt: Journal full-text XML · SegmentNT paper Methods: SegmentNT architecture, Model ablations, Human genomic elements and Multispecies training; official model-card Training data and licence metadata
Parameters563M for NT-v2 500M plus the 63M segmentation head; alternative encoder ablations are different complete pipelines.
Sources (7)instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; InstaDeepAI/segment_nt: README.md; InstaDeepAI/segment_nt: config.json; segmentnt: Journal full-text XML · SegmentNT paper Methods: SegmentNT architecture, Model ablations, Human genomic elements and Multispecies training; official model-card Training data and licence metadata
Known versionssegment_nt and segment_nt_multi_species; SegmentEnformer and SegmentBorzoi are distinct pipelines.
Sources (7)instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; InstaDeepAI/segment_nt: README.md; InstaDeepAI/segment_nt: config.json; segmentnt: Journal full-text XML · SegmentNT paper Methods: SegmentNT architecture, Model ablations, Human genomic elements and Multispecies training; official model-card Training data and licence metadata
Training dataHuman 14-element segmentation labels from GENCODE V44 and ENCODE SCREEN/DHS annotations. Human chromosomes 20 and 21 are held out for testing and 22 for validation. A separate multispecies model adds mouse, chicken, fly, zebrafish and worm.
Sources (7)instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; InstaDeepAI/segment_nt: README.md; InstaDeepAI/segment_nt: config.json; segmentnt: Journal full-text XML · SegmentNT paper Methods: SegmentNT architecture, Model ablations, Human genomic elements and Multispecies training; official model-card Training data and licence metadata
Training cutoffGENCODE V44 and the named ENCODE SCREEN/DHS resources define label provenance. The inspected paper does not give one latest-experiment date shared by every resource.
Sources (7)instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; InstaDeepAI/segment_nt: README.md; InstaDeepAI/segment_nt: config.json; segmentnt: Journal full-text XML · SegmentNT paper Methods: SegmentNT architecture, Model ablations, Human genomic elements and Multispecies training; official model-card Training data and licence metadata
Context limitsTrained on 30kb; inference up to 50kb requires the documented rescaling. Input token count also has divisibility constraints.
Sources (7)instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; InstaDeepAI/segment_nt: README.md; InstaDeepAI/segment_nt: config.json; segmentnt: Journal full-text XML · SegmentNT paper Methods: SegmentNT architecture, Model ablations, Human genomic elements and Multispecies training; official model-card Training data and licence metadata
Weights licenceCC-BY-NC-SA-4.0 declared by the official InstaDeepAI/segment_nt model card.
Sources (7)instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; InstaDeepAI/segment_nt: README.md; InstaDeepAI/segment_nt: config.json; segmentnt: Journal full-text XML · SegmentNT paper Methods: SegmentNT architecture, Model ablations, Human genomic elements and Multispecies training; official model-card Training data and licence metadata
AccessOfficial project documentation and implementation: https://github.com/instadeepai/nucleotide-transformer
Sources (7)instadeepai/nucleotide-transformer: README.md; instadeepai/nucleotide-transformer: docs/nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/agro_nucleotide_transformer.md; instadeepai/nucleotide-transformer: docs/segment_nt.md; InstaDeepAI/segment_nt: README.md; InstaDeepAI/segment_nt: config.json; segmentnt: Journal full-text XML · SegmentNT paper Methods: SegmentNT architecture, Model ablations, Human genomic elements and Multispecies training; official model-card Training data and licence metadata
Code licenceCC-BY-NC-SA-4.0
Sourcesinstadeepai/nucleotide-transformer: LICENSE.md · LICENSE.md: licence text

Evidence

Source checking verifies the cited claim or transcription. It does not establish independent reproduction.

Evidence table

Inspect claims, sources and review details

Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.

One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.

2 evidence rows matching the loaded filters

Claims, original sources and review scope · Release 2026-09-23-2b89723c6dd9
Property and statementOriginal source and locationReview and provenance
Relationship: family
discovery-model-segmentnt
Individual claims
SegmentNT supplementary information: complete Tables 2 and 3

Original source ↗

Supplementary Tables 2 and 3 (PDF pages 4 and 5); primary article Methods / Model training and evaluation (Sec16), Model ablations and baselines (Sec17); row SegmentNT-3kb (NTv1 human; 2.5B)

Version: Published supplementary information to s41592-025-02881-2; SHA-256 pinned snapshot
Retrieved: 2026-09-23T11:23:43.985410+00:00

source checked

automated source review · 2026-09-23

Audit details

Source-backed evaluated identity only; no independent reproduction.

Field: links:family:discovery-model-segmentnt

Claim: segmentnt-supplement-2025-method-segmentnt-3kb-ntv1-human-2-5b-discovery-model-segmentnt-identity-claim

Source artifact SHA-256: c4ad23a62ab161a464fe789e5dd167cf2bfe9f43a91a9c3fb1261b99a5594b0f

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Relationship: uses model
discovery-model-nucleotide-transformer
Individual claims
SegmentNT supplementary information: complete Tables 2 and 3

Original source ↗

Supplementary Tables 2 and 3 (PDF pages 4 and 5); primary article Methods / Model training and evaluation (Sec16), Model ablations and baselines (Sec17); row SegmentNT-3kb (NTv1 human; 2.5B)

Version: Published supplementary information to s41592-025-02881-2; SHA-256 pinned snapshot
Retrieved: 2026-09-23T11:23:43.985410+00:00

source checked

automated source review · 2026-09-23

Audit details

Source-backed evaluated identity only; no independent reproduction.

Field: links:uses_model:discovery-model-nucleotide-transformer

Claim: segmentnt-supplement-2025-method-segmentnt-3kb-ntv1-human-2-5b-discovery-model-nucleotide-transformer-identity-claim

Source artifact SHA-256: c4ad23a62ab161a464fe789e5dd167cf2bfe9f43a91a9c3fb1261b99a5594b0f

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Sources and history

View linked audit checks and correction history

Release 2026-09-23-2b89723c6dd9 · Record review: source checked

1 source records and release historyDownload this release
Technical metadata and extraction receipts

Stable ID: segmentnt-supplement-2025-method-segmentnt-3kb-ntv1-human-2-5b

areas
genomics
source locator
Supplementary Tables 2 and 3 (PDF pages 4 and 5); primary article Methods / Model training and evaluation (Sec16), Model ablations and baselines (Sec17); row SegmentNT-3kb (NTv1 human; 2.5B)
missing metadata
checkpoint revision: unreported; parameters: unextracted
Related records

Suggest a correction