Model type
Transformer representation pipeline; this record is the paper-specific evaluated configuration.
ProkBERT-mini learns microbial DNA representations with local-context-aware tokenisation.
Conceptual input–method–output guide. Check the procedure text and linked evaluation for fitted components, additional inputs and exact settings.
Transformer representation pipeline; this record is the paper-specific evaluated configuration.
Microbial DNA sequences
DNA embeddings and task-specific genomic classifications
limited source coverage · Automated source review, 2026-09-16. All specifications and missing details
Release 2026-09-17-d277315f7d76 · 1 evaluation · 4 metric rows. Different protocols are not a single leaderboard.
| Metric and finding | Coverage and uncertainty | Evidence |
|---|---|---|
| ProkBERT-mini: E. coli sigma70 promoter prediction Configuration: ProkBERT-miniProtocol: E. coli sigma70 independent promoter test (E. coli sigma70 promoter prediction)Dataset: E. coli sigma70 promoter dataset Test-only evaluation on Cassiano and Silva-Rocha 2020 data; methods have different training histories. 865 high-evidence RegulonDB 10.5 promoters and 1,000 nucleotide-distribution-matched negative sequences. Promoter exact matches removed from model training. Author-reported evaluation · Evaluation metadata: needs review | ||
| 0.87 Accuracy Unit: unitless · Direction: higher | Uncertainty: not reported in legacy extract Scored: Not reported · Eligible: Not reported | source checkedProkBERT family: genomic language models for microbiome applications; ProkBERT family: genomic language models for microbiome applications · Table 3, ProkBERT-mini row, Accuracy column Source checking is not independent reproduction. |
| 0.90 Sensitivity Unit: fraction · Direction: higher | Uncertainty: unreported Scored: Not reported · Eligible: Not reported | source checkedProkBERT family: genomic language models for microbiome applications · Table 3, row ProkBERT-mini, column Sensitivity; XML row2 column4 Source checking is not independent reproduction. |
| 0.85 Specificity Unit: fraction · Direction: higher | Uncertainty: unreported Scored: Not reported · Eligible: Not reported | source checkedProkBERT family: genomic language models for microbiome applications · Table 3, row ProkBERT-mini, column Specificity; XML row2 column5 Source checking is not independent reproduction. |
| 0.74 MCC Unit: dimensionless · Direction: higher | Uncertainty: unreported Scored: Not reported · Eligible: Not reported | source checkedProkBERT family: genomic language models for microbiome applications · Table 3, row ProkBERT-mini, column MCC; XML row2 column3 Source checking is not independent reproduction. |
A transformer encoder uses local-context-aware 6-mer tokenisation for the ProkBERT-mini variant, followed by self-supervised pretraining and task-specific fine-tuning.
The linked evaluation record identifies ProkBERT-mini: E. coli sigma70 promoter prediction. Its dataset, split, adaptation and evidence origin remain attached to the reported results.
Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.
Stable record: reported-model-4438513d9cd42cExplanatory profile: limited source coverage · Automated source review, 2026-09-16. Review applies to the cited claims; unresolved fields are listed below. Numerical results retain their own review status.
| Property | Description and evidence |
|---|---|
| Model type | Transformer representation pipeline; this record is the paper-specific evaluated configuration.SourcesProkBERT family: genomic language models for microbiome applications · 2 Materials and methods/2.1 Sequence data/2.1.1 Sequence segmentation and tokenization (paragraph 6); 2 Materials and methods (paragraph 1) |
| Architecture / procedure | A transformer encoder uses local-context-aware 6-mer tokenisation for the ProkBERT-mini variant, followed by self-supervised pretraining and task-specific fine-tuning.SourcesProkBERT family: genomic language models for microbiome applications · 2 Materials and methods/2.1 Sequence data/2.1.1 Sequence segmentation and tokenization (paragraph 6); 2 Materials and methods (paragraph 1) |
| Biological inputs | Microbial DNA sequencesSourcesProkBERT family: genomic language models for microbiome applications · 1 Introduction (paragraph 7); 4 Conclusion (paragraph 10) |
| Outputs | DNA embeddings and task-specific genomic classificationsSourcesProkBERT family: genomic language models for microbiome applications · 2 Materials and methods/2.5 Applied metrics (paragraph 1); 3 Results and discussion/3.1 ProkBERT's learned representations capture genomic structure and phylogeny (paragraph 4) |
| Parameters | Approximately 20 million parameters.SourcesProkBERT family: genomic language models for microbiome applications · 4 Conclusion (paragraph 8); 2 Materials and methods/2.2 Pretraining and learning sequence representations/2.2.2 Training process/2.2.2.2 Training phases and configuration (paragraph 1) |
| Known versions / configuration | ProkBERT-mini is the comparison-table label; that label does not specify an immutable weight revision. · Not reported in inspected sourcesSourcesProkBERT family: genomic language models for microbiome applications · Model identification in the comparison table and corresponding Methods; immutable checkpoint revision is not supplied by the table label. |
| Training data / fitting | Unlabelled microbial genome sequences followed by supervised task data, as specified in the paper.SourcesProkBERT family: genomic language models for microbiome applications · 2 Materials and methods/2.4 Application II: phage sequence analysis/2.4.2 Model training for phage sequence analysis (paragraph 1); 2 Materials and methods/2.2 Pretraining and learning sequence representations/2.2.4 Analysis of encoder outputs (paragraph 6) |
| Context limits | Approximately 1 kb for ProkBERT-mini; the distinct mini-long variant supports approximately 2 kb.SourcesProkBERT family: genomic language models for microbiome applications · 2 Materials and methods/2.1 Sequence data/2.1.1 Sequence segmentation and tokenization (paragraph 6); Table T3 (paragraph 1) |
| Access | Official study implementation and usage documentation: https://github.com/nbrg-ppcu/prokbert/blob/8670ae92b816cff158a0b85647a8dea122e251eb/README.md. This pinned documentation revision is not automatically the evaluated weight revision.Sourcesnbrg-ppcu/prokbert README.md · README.md; installation, model download and usage instructions |
| Code licence | MIT (study repository code at the cited revision; this does not establish every dependency or historical checkpoint licence).Sourcesnbrg-ppcu/prokbert LICENSE · LICENSE; complete licence text |
| Weights licence | The inspected model-access documentation does not explicitly identify terms for this exact evaluated checkpoint or fitted head; repository code terms are shown separately. · Not reported in inspected sourcesSourcesnbrg-ppcu/prokbert README.md · README.md; checkpoint/access documentation and licence scope |
Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.
One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.
19 evidence rows matching the loaded filters
| Property and statement | Original source and location | Review and provenance |
|---|---|---|
| Diagram caption Conceptual input–method–output guide. Check the procedure text and linked evaluation for fitted components, additional inputs and exact settings. Individual claims | ProkBERT family: genomic language models for microbiome applications 2 Materials and methods/2.1 Sequence data/2.1.1 Sequence segmentation and tokenization (paragraph 6); 2 Materials and methods (paragraph 1) Version: PMC10810988.1 | source checked automated source review · 2026-09-16 Audit detailsPrimary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Diagram steps ["Microbial DNA sequences","ProkBERT-mini","DNA embeddings and task-specific genomic classifications"] Individual claims | ProkBERT family: genomic language models for microbiome applications 2 Materials and methods/2.1 Sequence data/2.1.1 Sequence segmentation and tokenization (paragraph 6); 2 Materials and methods (paragraph 1) Version: PMC10810988.1 | source checked automated source review · 2026-09-16 Audit detailsPrimary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Diagram title Evaluated procedure (conceptual) Individual claims | ProkBERT family: genomic language models for microbiome applications 2 Materials and methods/2.1 Sequence data/2.1.1 Sequence segmentation and tokenization (paragraph 6); 2 Materials and methods (paragraph 1) Version: PMC10810988.1 | source checked automated source review · 2026-09-16 Audit detailsPrimary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Model type Transformer representation pipeline; this record is the paper-specific evaluated configuration. Individual claims | ProkBERT family: genomic language models for microbiome applications 2 Materials and methods/2.1 Sequence data/2.1.1 Sequence segmentation and tokenization (paragraph 6); 2 Materials and methods (paragraph 1) Version: PMC10810988.1 | source checked automated source review · 2026-09-16 Audit detailsPrimary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Architecture / procedure A transformer encoder uses local-context-aware 6-mer tokenisation for the ProkBERT-mini variant, followed by self-supervised pretraining and task-specific fine-tuning. Individual claims | ProkBERT family: genomic language models for microbiome applications 2 Materials and methods/2.1 Sequence data/2.1.1 Sequence segmentation and tokenization (paragraph 6); 2 Materials and methods (paragraph 1) Version: PMC10810988.1 | source checked automated source review · 2026-09-16 Audit detailsPrimary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Weights licence The inspected model-access documentation does not explicitly identify terms for this exact evaluated checkpoint or fitted head; repository code terms are shown separately. Individual claims | nbrg-ppcu/prokbert README.md README.md; checkpoint/access documentation and licence scope Version: 8670ae92b816cff158a0b85647a8dea122e251eb | unreported automated source review · 2026-09-16 Audit detailsPrimary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Biological inputs Microbial DNA sequences Individual claims | ProkBERT family: genomic language models for microbiome applications 1 Introduction (paragraph 7); 4 Conclusion (paragraph 10) Version: PMC10810988.1 | source checked automated source review · 2026-09-16 Audit detailsPrimary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Outputs DNA embeddings and task-specific genomic classifications Individual claims | ProkBERT family: genomic language models for microbiome applications 2 Materials and methods/2.5 Applied metrics (paragraph 1); 3 Results and discussion/3.1 ProkBERT's learned representations capture genomic structure and phylogeny (paragraph 4) Version: PMC10810988.1 | source checked automated source review · 2026-09-16 Audit detailsPrimary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Parameters Approximately 20 million parameters. Individual claims | ProkBERT family: genomic language models for microbiome applications 4 Conclusion (paragraph 8); 2 Materials and methods/2.2 Pretraining and learning sequence representations/2.2.2 Training process/2.2.2.2 Training phases and configuration (paragraph 1) Version: PMC10810988.1 | source checked automated source review · 2026-09-16 Audit detailsPrimary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
| Known versions / configuration ProkBERT-mini is the comparison-table label; that label does not specify an immutable weight revision. Individual claims | ProkBERT family: genomic language models for microbiome applications Model identification in the comparison table and corresponding Methods; immutable checkpoint revision is not supplied by the table label. Version: PMC10810988.1 | unreported automated source review · 2026-09-16 Audit detailsPrimary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction. Field: Source artifact SHA-256: Hash scope: Hash scope not separately documented; inspect source record |
Release 2026-09-17-d277315f7d76 · Record review: needs review
Stable ID: reported-model-4438513d9cd42c