Model type
DNA encoder with nucleotide-level segmentation head
SegmentNT labels genomic elements at individual nucleotide positions using a pretrained DNA backbone.
Conceptual summary of the documented data flow; optional inputs and configured downstream stages must be reported for a reproducible evaluation.
DNA encoder with nucleotide-level segmentation head
DNA sequences without N bases, tokenized into 6-mers under the documented length constraints.
Per-base probabilities for 14 genomic-element classes.
Official project documentation and implementation: https://github.com/instadeepai/nucleotide-transformer
Source reviewed · Automated source review, 2026-09-16. All specifications and missing details
14 evaluations · 28 metric rows. Different protocols are not a single leaderboard.
Applied filters: All linked evaluations
| Tested configuration | Protocol and dataset | Finding | Evidence and details |
|---|---|---|---|
| Configuration: SegmentNT-30kb | Protocol: SegmentNT human genome annotation 3UTR auPRC: 3UTR: per-nucleotide annotation (auPRC) Dataset subset: Human genome 3UTR test chromosomes 20 and 21 (SegmentNT human genome annotation split) | 0.70 (± 0.006) auprc dimensionless · higher Uncertainty: type: standard_deviation; value: 0.006 Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · source checkedMethods, coverage and sourceTest chromosomes 20 and 21; validation chromosome 22; training remaining chromosomes. Test chunks with genes homologous to train/validation genes excluded using Ensembl BioMart accessed 2024-05-08; homologous distal regulatory elements not excluded. Ten test-set samplings with sliding windows beginning at different genomic starting positions. Mean plus reported standard deviation; not ten training seeds. Best validation checkpoint by average MCC across 14 elements. Per-nucleotide predictions pooled across test sequences separately for each genomic element. MCC and auPRC are separate metrics, not cross-element or cross-table pooled comparisons. Aggregation: Not reported SegmentNT supplementary information: complete Tables 2 and 3 · Supplementary Table 3; PDF page 5 (printed 4); block 1; data row 17; model SegmentNT-30kb; column 3UTR |
| Configuration: SegmentNT-30kb | Protocol: SegmentNT human genome annotation 3UTR MCC: 3UTR: per-nucleotide annotation (MCC) Dataset subset: Human genome 3UTR test chromosomes 20 and 21 (SegmentNT human genome annotation split) | 0.70 (± 0.007) mcc dimensionless · higher Uncertainty: type: standard_deviation; value: 0.007 Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · source checkedMethods, coverage and sourceSegmentNT-30kb on SegmentNT human genome annotation 3UTR MCC: 3UTR: per-nucleotide annotation (MCC) Test chromosomes 20 and 21; validation chromosome 22; training remaining chromosomes. Test chunks with genes homologous to train/validation genes excluded using Ensembl BioMart accessed 2024-05-08; homologous distal regulatory elements not excluded. Ten test-set samplings with sliding windows beginning at different genomic starting positions. Mean plus reported standard deviation; not ten training seeds. Best validation checkpoint by average MCC across 14 elements. Per-nucleotide predictions pooled across test sequences separately for each genomic element. MCC and auPRC are separate metrics, not cross-element or cross-table pooled comparisons. Aggregation: Not reported SegmentNT supplementary information: complete Tables 2 and 3 · Supplementary Table 2; PDF page 4 (printed 3); block 1; data row 17; model SegmentNT-30kb; column 3UTR |
| Configuration: SegmentNT-30kb | Protocol: SegmentNT human genome annotation 5UTR auPRC: 5UTR: per-nucleotide annotation (auPRC) Dataset subset: Human genome 5UTR test chromosomes 20 and 21 (SegmentNT human genome annotation split) | 0.44 (± 0.004) auprc dimensionless · higher Uncertainty: type: standard_deviation; value: 0.004 Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · source checkedMethods, coverage and sourceTest chromosomes 20 and 21; validation chromosome 22; training remaining chromosomes. Test chunks with genes homologous to train/validation genes excluded using Ensembl BioMart accessed 2024-05-08; homologous distal regulatory elements not excluded. Ten test-set samplings with sliding windows beginning at different genomic starting positions. Mean plus reported standard deviation; not ten training seeds. Best validation checkpoint by average MCC across 14 elements. Per-nucleotide predictions pooled across test sequences separately for each genomic element. MCC and auPRC are separate metrics, not cross-element or cross-table pooled comparisons. Aggregation: Not reported SegmentNT supplementary information: complete Tables 2 and 3 · Supplementary Table 3; PDF page 5 (printed 4); block 1; data row 17; model SegmentNT-30kb; column 5UTR |
| Configuration: SegmentNT-30kb | Protocol: SegmentNT human genome annotation 5UTR MCC: 5UTR: per-nucleotide annotation (MCC) Dataset subset: Human genome 5UTR test chromosomes 20 and 21 (SegmentNT human genome annotation split) | 0.50 (± 0.003) mcc dimensionless · higher Uncertainty: type: standard_deviation; value: 0.003 Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · source checkedMethods, coverage and sourceSegmentNT-30kb on SegmentNT human genome annotation 5UTR MCC: 5UTR: per-nucleotide annotation (MCC) Test chromosomes 20 and 21; validation chromosome 22; training remaining chromosomes. Test chunks with genes homologous to train/validation genes excluded using Ensembl BioMart accessed 2024-05-08; homologous distal regulatory elements not excluded. Ten test-set samplings with sliding windows beginning at different genomic starting positions. Mean plus reported standard deviation; not ten training seeds. Best validation checkpoint by average MCC across 14 elements. Per-nucleotide predictions pooled across test sequences separately for each genomic element. MCC and auPRC are separate metrics, not cross-element or cross-table pooled comparisons. Aggregation: Not reported SegmentNT supplementary information: complete Tables 2 and 3 · Supplementary Table 2; PDF page 4 (printed 3); block 1; data row 17; model SegmentNT-30kb; column 5UTR |
| Configuration: SegmentNT-30kb | Protocol: SegmentNT human genome annotation CTCF-bound auPRC: CTCF-bound: per-nucleotide annotation (auPRC) Dataset subset: Human genome CTCF-bound test chromosomes 20 and 21 (SegmentNT human genome annotation split) | 0.04 (± 0.001) auprc dimensionless · higher Uncertainty: type: standard_deviation; value: 0.001 Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · source checkedMethods, coverage and sourceTest chromosomes 20 and 21; validation chromosome 22; training remaining chromosomes. Test chunks with genes homologous to train/validation genes excluded using Ensembl BioMart accessed 2024-05-08; homologous distal regulatory elements not excluded. Ten test-set samplings with sliding windows beginning at different genomic starting positions. Mean plus reported standard deviation; not ten training seeds. Best validation checkpoint by average MCC across 14 elements. Per-nucleotide predictions pooled across test sequences separately for each genomic element. MCC and auPRC are separate metrics, not cross-element or cross-table pooled comparisons. Aggregation: Not reported SegmentNT supplementary information: complete Tables 2 and 3 · Supplementary Table 3; PDF page 5 (printed 4); block 1; data row 17; model SegmentNT-30kb; column CTCF-bound |
| Configuration: SegmentNT-30kb | Protocol: SegmentNT human genome annotation CTCF-bound MCC: CTCF-bound: per-nucleotide annotation (MCC) Dataset subset: Human genome CTCF-bound test chromosomes 20 and 21 (SegmentNT human genome annotation split) | 0.08 (± 0.002) mcc dimensionless · higher Uncertainty: type: standard_deviation; value: 0.002 Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · source checkedMethods, coverage and sourceTest chromosomes 20 and 21; validation chromosome 22; training remaining chromosomes. Test chunks with genes homologous to train/validation genes excluded using Ensembl BioMart accessed 2024-05-08; homologous distal regulatory elements not excluded. Ten test-set samplings with sliding windows beginning at different genomic starting positions. Mean plus reported standard deviation; not ten training seeds. Best validation checkpoint by average MCC across 14 elements. Per-nucleotide predictions pooled across test sequences separately for each genomic element. MCC and auPRC are separate metrics, not cross-element or cross-table pooled comparisons. Aggregation: Not reported SegmentNT supplementary information: complete Tables 2 and 3 · Supplementary Table 2; PDF page 4 (printed 3); block 1; data row 17; model SegmentNT-30kb; column CTCF-bound |
| Configuration: SegmentNT-30kb | Protocol: SegmentNT human genome annotation enhancer tissue-invariant auPRC: enhancer tissue-invariant: per-nucleotide annotation (auPRC) Dataset subset: Human genome enhancer tissue-invariant test chromosomes 20 and 21 (SegmentNT human genome annotation split) | 0.10 (± 0.004) auprc dimensionless · higher Uncertainty: type: standard_deviation; value: 0.004 Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · source checkedMethods, coverage and sourceTest chromosomes 20 and 21; validation chromosome 22; training remaining chromosomes. Test chunks with genes homologous to train/validation genes excluded using Ensembl BioMart accessed 2024-05-08; homologous distal regulatory elements not excluded. Ten test-set samplings with sliding windows beginning at different genomic starting positions. Mean plus reported standard deviation; not ten training seeds. Best validation checkpoint by average MCC across 14 elements. Per-nucleotide predictions pooled across test sequences separately for each genomic element. MCC and auPRC are separate metrics, not cross-element or cross-table pooled comparisons. Aggregation: Not reported SegmentNT supplementary information: complete Tables 2 and 3 · Supplementary Table 3; PDF page 5 (printed 4); block 1; data row 17; model SegmentNT-30kb; column enhancer tissue-invariant |
| Configuration: SegmentNT-30kb | Protocol: SegmentNT human genome annotation enhancer tissue-invariant MCC: enhancer tissue-invariant: per-nucleotide annotation (MCC) Dataset subset: Human genome enhancer tissue-invariant test chromosomes 20 and 21 (SegmentNT human genome annotation split) | 0.19 (± 0.005) mcc dimensionless · higher Uncertainty: type: standard_deviation; value: 0.005 Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · source checkedMethods, coverage and sourceTest chromosomes 20 and 21; validation chromosome 22; training remaining chromosomes. Test chunks with genes homologous to train/validation genes excluded using Ensembl BioMart accessed 2024-05-08; homologous distal regulatory elements not excluded. Ten test-set samplings with sliding windows beginning at different genomic starting positions. Mean plus reported standard deviation; not ten training seeds. Best validation checkpoint by average MCC across 14 elements. Per-nucleotide predictions pooled across test sequences separately for each genomic element. MCC and auPRC are separate metrics, not cross-element or cross-table pooled comparisons. Aggregation: Not reported SegmentNT supplementary information: complete Tables 2 and 3 · Supplementary Table 2; PDF page 4 (printed 3); block 1; data row 17; model SegmentNT-30kb; column enhancer tissue-invariant |
| Configuration: SegmentNT-30kb | Protocol: SegmentNT human genome annotation enhancer tissue-specific auPRC: enhancer tissue-specific: per-nucleotide annotation (auPRC) Dataset subset: Human genome enhancer tissue-specific test chromosomes 20 and 21 (SegmentNT human genome annotation split) | 0.29 (± 0.002) auprc dimensionless · higher Uncertainty: type: standard_deviation; value: 0.002 Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · source checkedMethods, coverage and sourceTest chromosomes 20 and 21; validation chromosome 22; training remaining chromosomes. Test chunks with genes homologous to train/validation genes excluded using Ensembl BioMart accessed 2024-05-08; homologous distal regulatory elements not excluded. Ten test-set samplings with sliding windows beginning at different genomic starting positions. Mean plus reported standard deviation; not ten training seeds. Best validation checkpoint by average MCC across 14 elements. Per-nucleotide predictions pooled across test sequences separately for each genomic element. MCC and auPRC are separate metrics, not cross-element or cross-table pooled comparisons. Aggregation: Not reported SegmentNT supplementary information: complete Tables 2 and 3 · Supplementary Table 3; PDF page 5 (printed 4); block 1; data row 17; model SegmentNT-30kb; column enhancer tissue-specific |
| Configuration: SegmentNT-30kb | Protocol: SegmentNT human genome annotation enhancer tissue-specific MCC: enhancer tissue-specific: per-nucleotide annotation (MCC) Dataset subset: Human genome enhancer tissue-specific test chromosomes 20 and 21 (SegmentNT human genome annotation split) | 0.27 (± 0.002) mcc dimensionless · higher Uncertainty: type: standard_deviation; value: 0.002 Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · source checkedMethods, coverage and sourceTest chromosomes 20 and 21; validation chromosome 22; training remaining chromosomes. Test chunks with genes homologous to train/validation genes excluded using Ensembl BioMart accessed 2024-05-08; homologous distal regulatory elements not excluded. Ten test-set samplings with sliding windows beginning at different genomic starting positions. Mean plus reported standard deviation; not ten training seeds. Best validation checkpoint by average MCC across 14 elements. Per-nucleotide predictions pooled across test sequences separately for each genomic element. MCC and auPRC are separate metrics, not cross-element or cross-table pooled comparisons. Aggregation: Not reported SegmentNT supplementary information: complete Tables 2 and 3 · Supplementary Table 2; PDF page 4 (printed 3); block 1; data row 17; model SegmentNT-30kb; column enhancer tissue-specific |
| Configuration: SegmentNT-30kb | Protocol: SegmentNT human genome annotation exon auPRC: exon: per-nucleotide annotation (auPRC) Dataset subset: Human genome exon test chromosomes 20 and 21 (SegmentNT human genome annotation split) | 0.58 (± 0.004) auprc dimensionless · higher Uncertainty: type: standard_deviation; value: 0.004 Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · source checkedMethods, coverage and sourceTest chromosomes 20 and 21; validation chromosome 22; training remaining chromosomes. Test chunks with genes homologous to train/validation genes excluded using Ensembl BioMart accessed 2024-05-08; homologous distal regulatory elements not excluded. Ten test-set samplings with sliding windows beginning at different genomic starting positions. Mean plus reported standard deviation; not ten training seeds. Best validation checkpoint by average MCC across 14 elements. Per-nucleotide predictions pooled across test sequences separately for each genomic element. MCC and auPRC are separate metrics, not cross-element or cross-table pooled comparisons. Aggregation: Not reported SegmentNT supplementary information: complete Tables 2 and 3 · Supplementary Table 3; PDF page 5 (printed 4); block 1; data row 17; model SegmentNT-30kb; column exon |
| Configuration: SegmentNT-30kb | Protocol: SegmentNT human genome annotation exon MCC: exon: per-nucleotide annotation (MCC) Dataset subset: Human genome exon test chromosomes 20 and 21 (SegmentNT human genome annotation split) | 0.62 (± 0.003) mcc dimensionless · higher Uncertainty: type: standard_deviation; value: 0.003 Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · source checkedMethods, coverage and sourceSegmentNT-30kb on SegmentNT human genome annotation exon MCC: exon: per-nucleotide annotation (MCC) Test chromosomes 20 and 21; validation chromosome 22; training remaining chromosomes. Test chunks with genes homologous to train/validation genes excluded using Ensembl BioMart accessed 2024-05-08; homologous distal regulatory elements not excluded. Ten test-set samplings with sliding windows beginning at different genomic starting positions. Mean plus reported standard deviation; not ten training seeds. Best validation checkpoint by average MCC across 14 elements. Per-nucleotide predictions pooled across test sequences separately for each genomic element. MCC and auPRC are separate metrics, not cross-element or cross-table pooled comparisons. Aggregation: Not reported SegmentNT supplementary information: complete Tables 2 and 3 · Supplementary Table 2; PDF page 4 (printed 3); block 1; data row 17; model SegmentNT-30kb; column exon |
| Configuration: SegmentNT-30kb | Protocol: SegmentNT human genome annotation intron auPRC: intron: per-nucleotide annotation (auPRC) Dataset subset: Human genome intron test chromosomes 20 and 21 (SegmentNT human genome annotation split) | 0.84 (± 0.002) auprc dimensionless · higher Uncertainty: type: standard_deviation; value: 0.002 Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · source checkedMethods, coverage and sourceTest chromosomes 20 and 21; validation chromosome 22; training remaining chromosomes. Test chunks with genes homologous to train/validation genes excluded using Ensembl BioMart accessed 2024-05-08; homologous distal regulatory elements not excluded. Ten test-set samplings with sliding windows beginning at different genomic starting positions. Mean plus reported standard deviation; not ten training seeds. Best validation checkpoint by average MCC across 14 elements. Per-nucleotide predictions pooled across test sequences separately for each genomic element. MCC and auPRC are separate metrics, not cross-element or cross-table pooled comparisons. Aggregation: Not reported SegmentNT supplementary information: complete Tables 2 and 3 · Supplementary Table 3; PDF page 5 (printed 4); block 1; data row 17; model SegmentNT-30kb; column intron |
| Configuration: SegmentNT-30kb | Protocol: SegmentNT human genome annotation intron MCC: intron: per-nucleotide annotation (MCC) Dataset subset: Human genome intron test chromosomes 20 and 21 (SegmentNT human genome annotation split) | 0.52 (± 0.005) mcc dimensionless · higher Uncertainty: type: standard_deviation; value: 0.005 Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · source checkedMethods, coverage and sourceTest chromosomes 20 and 21; validation chromosome 22; training remaining chromosomes. Test chunks with genes homologous to train/validation genes excluded using Ensembl BioMart accessed 2024-05-08; homologous distal regulatory elements not excluded. Ten test-set samplings with sliding windows beginning at different genomic starting positions. Mean plus reported standard deviation; not ten training seeds. Best validation checkpoint by average MCC across 14 elements. Per-nucleotide predictions pooled across test sequences separately for each genomic element. MCC and auPRC are separate metrics, not cross-element or cross-table pooled comparisons. Aggregation: Not reported SegmentNT supplementary information: complete Tables 2 and 3 · Supplementary Table 2; PDF page 4 (printed 3); block 1; data row 17; model SegmentNT-30kb; column intron |
| Configuration: SegmentNT-30kb | Protocol: SegmentNT human genome annotation lncRNA auPRC: lncRNA: per-nucleotide annotation (auPRC) Dataset subset: Human genome lncRNA test chromosomes 20 and 21 (SegmentNT human genome annotation split) | 0.25 (± 0.003) auprc dimensionless · higher Uncertainty: type: standard_deviation; value: 0.003 Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · source checkedMethods, coverage and sourceTest chromosomes 20 and 21; validation chromosome 22; training remaining chromosomes. Test chunks with genes homologous to train/validation genes excluded using Ensembl BioMart accessed 2024-05-08; homologous distal regulatory elements not excluded. Ten test-set samplings with sliding windows beginning at different genomic starting positions. Mean plus reported standard deviation; not ten training seeds. Best validation checkpoint by average MCC across 14 elements. Per-nucleotide predictions pooled across test sequences separately for each genomic element. MCC and auPRC are separate metrics, not cross-element or cross-table pooled comparisons. Aggregation: Not reported SegmentNT supplementary information: complete Tables 2 and 3 · Supplementary Table 3; PDF page 5 (printed 4); block 2; data row 17; model SegmentNT-30kb; column lncRNA |
| Configuration: SegmentNT-30kb | Protocol: SegmentNT human genome annotation lncRNA MCC: lncRNA: per-nucleotide annotation (MCC) Dataset subset: Human genome lncRNA test chromosomes 20 and 21 (SegmentNT human genome annotation split) | 0.06 (± 0.014) mcc dimensionless · higher Uncertainty: type: standard_deviation; value: 0.014 Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · source checkedMethods, coverage and sourceTest chromosomes 20 and 21; validation chromosome 22; training remaining chromosomes. Test chunks with genes homologous to train/validation genes excluded using Ensembl BioMart accessed 2024-05-08; homologous distal regulatory elements not excluded. Ten test-set samplings with sliding windows beginning at different genomic starting positions. Mean plus reported standard deviation; not ten training seeds. Best validation checkpoint by average MCC across 14 elements. Per-nucleotide predictions pooled across test sequences separately for each genomic element. MCC and auPRC are separate metrics, not cross-element or cross-table pooled comparisons. Aggregation: Not reported SegmentNT supplementary information: complete Tables 2 and 3 · Supplementary Table 2; PDF page 4 (printed 3); block 2; data row 17; model SegmentNT-30kb; column lncRNA |
| Configuration: SegmentNT-30kb | Protocol: SegmentNT human genome annotation polyA signal auPRC: polyA signal: per-nucleotide annotation (auPRC) Dataset subset: Human genome polyA signal test chromosomes 20 and 21 (SegmentNT human genome annotation split) | 0.31 (± 0.005) auprc dimensionless · higher Uncertainty: type: standard_deviation; value: 0.005 Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · source checkedMethods, coverage and sourceTest chromosomes 20 and 21; validation chromosome 22; training remaining chromosomes. Test chunks with genes homologous to train/validation genes excluded using Ensembl BioMart accessed 2024-05-08; homologous distal regulatory elements not excluded. Ten test-set samplings with sliding windows beginning at different genomic starting positions. Mean plus reported standard deviation; not ten training seeds. Best validation checkpoint by average MCC across 14 elements. Per-nucleotide predictions pooled across test sequences separately for each genomic element. MCC and auPRC are separate metrics, not cross-element or cross-table pooled comparisons. Aggregation: Not reported SegmentNT supplementary information: complete Tables 2 and 3 · Supplementary Table 3; PDF page 5 (printed 4); block 2; data row 17; model SegmentNT-30kb; column polyA signal |
| Configuration: SegmentNT-30kb | Protocol: SegmentNT human genome annotation polyA signal MCC: polyA signal: per-nucleotide annotation (MCC) Dataset subset: Human genome polyA signal test chromosomes 20 and 21 (SegmentNT human genome annotation split) | 0.40 (± 0.004) mcc dimensionless · higher Uncertainty: type: standard_deviation; value: 0.004 Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · source checkedMethods, coverage and sourceTest chromosomes 20 and 21; validation chromosome 22; training remaining chromosomes. Test chunks with genes homologous to train/validation genes excluded using Ensembl BioMart accessed 2024-05-08; homologous distal regulatory elements not excluded. Ten test-set samplings with sliding windows beginning at different genomic starting positions. Mean plus reported standard deviation; not ten training seeds. Best validation checkpoint by average MCC across 14 elements. Per-nucleotide predictions pooled across test sequences separately for each genomic element. MCC and auPRC are separate metrics, not cross-element or cross-table pooled comparisons. Aggregation: Not reported SegmentNT supplementary information: complete Tables 2 and 3 · Supplementary Table 2; PDF page 4 (printed 3); block 2; data row 17; model SegmentNT-30kb; column polyA signal |
| Configuration: SegmentNT-30kb | Protocol: SegmentNT human genome annotation promoter tissue-invariant auPRC: promoter tissue-invariant: per-nucleotide annotation (auPRC) Dataset subset: Human genome promoter tissue-invariant test chromosomes 20 and 21 (SegmentNT human genome annotation split) | 0.55 (± 0.012) auprc dimensionless · higher Uncertainty: type: standard_deviation; value: 0.012 Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · source checkedMethods, coverage and sourceTest chromosomes 20 and 21; validation chromosome 22; training remaining chromosomes. Test chunks with genes homologous to train/validation genes excluded using Ensembl BioMart accessed 2024-05-08; homologous distal regulatory elements not excluded. Ten test-set samplings with sliding windows beginning at different genomic starting positions. Mean plus reported standard deviation; not ten training seeds. Best validation checkpoint by average MCC across 14 elements. Per-nucleotide predictions pooled across test sequences separately for each genomic element. MCC and auPRC are separate metrics, not cross-element or cross-table pooled comparisons. Aggregation: Not reported SegmentNT supplementary information: complete Tables 2 and 3 · Supplementary Table 3; PDF page 5 (printed 4); block 2; data row 17; model SegmentNT-30kb; column promoter tissue-invariant |
| Configuration: SegmentNT-30kb | Protocol: SegmentNT human genome annotation promoter tissue-invariant MCC: promoter tissue-invariant: per-nucleotide annotation (MCC) Dataset subset: Human genome promoter tissue-invariant test chromosomes 20 and 21 (SegmentNT human genome annotation split) | 0.58 (± 0.007) mcc dimensionless · higher Uncertainty: type: standard_deviation; value: 0.007 Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · source checkedMethods, coverage and sourceTest chromosomes 20 and 21; validation chromosome 22; training remaining chromosomes. Test chunks with genes homologous to train/validation genes excluded using Ensembl BioMart accessed 2024-05-08; homologous distal regulatory elements not excluded. Ten test-set samplings with sliding windows beginning at different genomic starting positions. Mean plus reported standard deviation; not ten training seeds. Best validation checkpoint by average MCC across 14 elements. Per-nucleotide predictions pooled across test sequences separately for each genomic element. MCC and auPRC are separate metrics, not cross-element or cross-table pooled comparisons. Aggregation: Not reported SegmentNT supplementary information: complete Tables 2 and 3 · Supplementary Table 2; PDF page 4 (printed 3); block 2; data row 17; model SegmentNT-30kb; column promoter tissue-invariant |
| Configuration: SegmentNT-30kb | Protocol: SegmentNT human genome annotation promoter tissue-specific auPRC: promoter tissue-specific: per-nucleotide annotation (auPRC) Dataset subset: Human genome promoter tissue-specific test chromosomes 20 and 21 (SegmentNT human genome annotation split) | 0.15 (± 0.004) auprc dimensionless · higher Uncertainty: type: standard_deviation; value: 0.004 Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · source checkedMethods, coverage and sourceTest chromosomes 20 and 21; validation chromosome 22; training remaining chromosomes. Test chunks with genes homologous to train/validation genes excluded using Ensembl BioMart accessed 2024-05-08; homologous distal regulatory elements not excluded. Ten test-set samplings with sliding windows beginning at different genomic starting positions. Mean plus reported standard deviation; not ten training seeds. Best validation checkpoint by average MCC across 14 elements. Per-nucleotide predictions pooled across test sequences separately for each genomic element. MCC and auPRC are separate metrics, not cross-element or cross-table pooled comparisons. Aggregation: Not reported SegmentNT supplementary information: complete Tables 2 and 3 · Supplementary Table 3; PDF page 5 (printed 4); block 2; data row 17; model SegmentNT-30kb; column promoter tissue-specific |
| Configuration: SegmentNT-30kb | Protocol: SegmentNT human genome annotation promoter tissue-specific MCC: promoter tissue-specific: per-nucleotide annotation (MCC) Dataset subset: Human genome promoter tissue-specific test chromosomes 20 and 21 (SegmentNT human genome annotation split) | 0.23 (± 0.006) mcc dimensionless · higher Uncertainty: type: standard_deviation; value: 0.006 Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · source checkedMethods, coverage and sourceTest chromosomes 20 and 21; validation chromosome 22; training remaining chromosomes. Test chunks with genes homologous to train/validation genes excluded using Ensembl BioMart accessed 2024-05-08; homologous distal regulatory elements not excluded. Ten test-set samplings with sliding windows beginning at different genomic starting positions. Mean plus reported standard deviation; not ten training seeds. Best validation checkpoint by average MCC across 14 elements. Per-nucleotide predictions pooled across test sequences separately for each genomic element. MCC and auPRC are separate metrics, not cross-element or cross-table pooled comparisons. Aggregation: Not reported SegmentNT supplementary information: complete Tables 2 and 3 · Supplementary Table 2; PDF page 4 (printed 3); block 2; data row 17; model SegmentNT-30kb; column promoter tissue-specific |
| Configuration: SegmentNT-30kb | Protocol: SegmentNT human genome annotation protein coding gene auPRC: protein coding gene: per-nucleotide annotation (auPRC) Dataset subset: Human genome protein coding gene test chromosomes 20 and 21 (SegmentNT human genome annotation split) | 0.89 (± 0.001) auprc dimensionless · higher Uncertainty: type: standard_deviation; value: 0.001 Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · source checkedMethods, coverage and sourceTest chromosomes 20 and 21; validation chromosome 22; training remaining chromosomes. Test chunks with genes homologous to train/validation genes excluded using Ensembl BioMart accessed 2024-05-08; homologous distal regulatory elements not excluded. Ten test-set samplings with sliding windows beginning at different genomic starting positions. Mean plus reported standard deviation; not ten training seeds. Best validation checkpoint by average MCC across 14 elements. Per-nucleotide predictions pooled across test sequences separately for each genomic element. MCC and auPRC are separate metrics, not cross-element or cross-table pooled comparisons. Aggregation: Not reported SegmentNT supplementary information: complete Tables 2 and 3 · Supplementary Table 3; PDF page 5 (printed 4); block 2; data row 17; model SegmentNT-30kb; column protein coding gene |
| Configuration: SegmentNT-30kb | Protocol: SegmentNT human genome annotation protein coding gene MCC: protein coding gene: per-nucleotide annotation (MCC) Dataset subset: Human genome protein coding gene test chromosomes 20 and 21 (SegmentNT human genome annotation split) | 0.70 (± 0.003) mcc dimensionless · higher Uncertainty: type: standard_deviation; value: 0.003 Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · source checkedMethods, coverage and sourceTest chromosomes 20 and 21; validation chromosome 22; training remaining chromosomes. Test chunks with genes homologous to train/validation genes excluded using Ensembl BioMart accessed 2024-05-08; homologous distal regulatory elements not excluded. Ten test-set samplings with sliding windows beginning at different genomic starting positions. Mean plus reported standard deviation; not ten training seeds. Best validation checkpoint by average MCC across 14 elements. Per-nucleotide predictions pooled across test sequences separately for each genomic element. MCC and auPRC are separate metrics, not cross-element or cross-table pooled comparisons. Aggregation: Not reported SegmentNT supplementary information: complete Tables 2 and 3 · Supplementary Table 2; PDF page 4 (printed 3); block 2; data row 17; model SegmentNT-30kb; column protein coding gene |
| Configuration: SegmentNT-30kb | Protocol: SegmentNT human genome annotation splice acceptor auPRC: splice acceptor: per-nucleotide annotation (auPRC) Dataset subset: Human genome splice acceptor test chromosomes 20 and 21 (SegmentNT human genome annotation split) | 0.65 (± 0.002) auprc dimensionless · higher Uncertainty: type: standard_deviation; value: 0.002 Coverage: Not reported scored / Not reported eligible | Author-reported evaluation · source checkedMethods, coverage and sourceTest chromosomes 20 and 21; validation chromosome 22; training remaining chromosomes. Test chunks with genes homologous to train/validation genes excluded using Ensembl BioMart accessed 2024-05-08; homologous distal regulatory elements not excluded. Ten test-set samplings with sliding windows beginning at different genomic starting positions. Mean plus reported standard deviation; not ten training seeds. Best validation checkpoint by average MCC across 14 elements. Per-nucleotide predictions pooled across test sequences separately for each genomic element. MCC and auPRC are separate metrics, not cross-element or cross-table pooled comparisons. Aggregation: Not reported SegmentNT supplementary information: complete Tables 2 and 3 · Supplementary Table 3; PDF page 5 (printed 4); block 2; data row 17; model SegmentNT-30kb; column splice acceptor |
Source checking is not independent reproduction. Release 2026-09-23-2b89723c6dd9.
Related profile: SegmentNT. This page retains the exact record and its evaluation context.
Underlying model: Nucleotide Transformer. Results on this page belong to this configuration and its evaluated settings.
Evaluated configuration as printed in Supplementary Tables 2 and 3; source reports training and checkpoint selection in Sec16–17.
SegmentNT labels genomic elements at individual nucleotide positions using a pretrained DNA backbone. Nucleotide Transformer backbone with a one-dimensional U-Net segmentation head; YaRN rescales positions for longer inputs. The documented inputs are DNA sequences without N bases, tokenized into 6-mers under the documented length constraints. The output consists of per-base probabilities for 14 genomic-element classes.
segment_nt and segment_nt_multi_species; SegmentEnformer and SegmentBorzoi are distinct pipelines. Trained on 30kb; inference up to 50kb requires the documented rescaling. Input token count also has divisibility constraints.
Inspected pinned official documentation, relevant implementation files and named primary-paper sections. Claims are limited to those artifacts. Remaining field extraction and identity conflicts are explicit; no new performance claims, model runs or human review are implied.
Stable record: discovery-model-segmentntExplanatory profile: source reviewed · Automated source review, 2026-09-16. Review applies to the cited claims; unresolved fields are listed below. Numerical results retain their own review status.
Source checking verifies the cited claim or transcription. It does not establish independent reproduction.
Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.
One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.
2 evidence rows matching the loaded filters
| Property and statement | Original source and location | Review and provenance |
|---|---|---|
| Relationship: family discovery-model-segmentnt Individual claims | SegmentNT supplementary information: complete Tables 2 and 3 Supplementary Tables 2 and 3 (PDF pages 4 and 5); primary article Methods / Model training and evaluation (Sec16), Model ablations and baselines (Sec17); row SegmentNT-30kb Version: Published supplementary information to s41592-025-02881-2; SHA-256 pinned snapshot | source checked automated source review · 2026-09-23 Audit detailsSource-backed evaluated identity only; no independent reproduction. Field: Claim: segmentnt-supplement-2025-method-segmentnt-30kb-discovery-model-segmentnt-identity-claim Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Relationship: uses model discovery-model-nucleotide-transformer Individual claims | SegmentNT supplementary information: complete Tables 2 and 3 Supplementary Tables 2 and 3 (PDF pages 4 and 5); primary article Methods / Model training and evaluation (Sec16), Model ablations and baselines (Sec17); row SegmentNT-30kb Version: Published supplementary information to s41592-025-02881-2; SHA-256 pinned snapshot | source checked automated source review · 2026-09-23 Audit detailsSource-backed evaluated identity only; no independent reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
View linked audit checks and correction history
Release 2026-09-23-2b89723c6dd9 · Record review: source checked
Stable ID: segmentnt-supplement-2025-method-segmentnt-30kb