rewire.it
Configuration

ProkBERT-mini

ProkBERT-mini learns microbial DNA representations with local-context-aware tokenisation.

SourcesProkBERT family: genomic language models for microbiome applications · 2 Materials and methods/2.1 Sequence data/2.1.1 Sequence segmentation and tokenization (paragraph 6); 4 Conclusion (paragraph 5)

1 evaluation · 4 metric rows

How it worksEvaluated procedure (conceptual)
Evaluated procedure (conceptual)1. Microbial DNA sequences. Then: 2. ProkBERT-mini. Then: 3. DNA embeddings and task-specific genomic classificationsEvaluated procedure (conceptual)1. Microbial DNA sequences. Then: 2. ProkBERT-mini. Then: 3. DNA embeddings and task-specific genomic classificationsEvaluated procedure (conceptual)1. Microbial DNA sequences. Then: 2. ProkBERT-mini. Then: 3. DNA embeddings and task-specific genomic classifications

Conceptual input–method–output guide. Check the procedure text and linked evaluation for fitted components, additional inputs and exact settings.

SourcesProkBERT family: genomic language models for microbiome applications · 2 Materials and methods/2.1 Sequence data/2.1.1 Sequence segmentation and tokenization (paragraph 6); 2 Materials and methods (paragraph 1)

At a glance

limited source coverage · Automated source review, 2026-09-16. All specifications and missing details

Evaluations and results

Release 2026-09-17-d277315f7d76 · 1 evaluation · 4 metric rows. Different protocols are not a single leaderboard.

Results grouped by the exact reported evaluation
Metric and findingCoverage and uncertaintyEvidence
ProkBERT-mini: E. coli sigma70 promoter prediction

Test-only evaluation on Cassiano and Silva-Rocha 2020 data; methods have different training histories. 865 high-evidence RegulonDB 10.5 promoters and 1,000 nucleotide-distribution-matched negative sequences. Promoter exact matches removed from model training.

Author-reported evaluation · Evaluation metadata: needs review

0.87 Accuracy

Unit: unitless · Direction: higher

Uncertainty: not reported in legacy extract

Scored: Not reported · Eligible: Not reported

source checkedProkBERT family: genomic language models for microbiome applications; ProkBERT family: genomic language models for microbiome applications · Table 3, ProkBERT-mini row, Accuracy column

Source checking is not independent reproduction.

0.90 Sensitivity

Unit: fraction · Direction: higher

Uncertainty: unreported

Scored: Not reported · Eligible: Not reported

source checkedProkBERT family: genomic language models for microbiome applications · Table 3, row ProkBERT-mini, column Sensitivity; XML row2 column4

Source checking is not independent reproduction.

0.85 Specificity

Unit: fraction · Direction: higher

Uncertainty: unreported

Scored: Not reported · Eligible: Not reported

source checkedProkBERT family: genomic language models for microbiome applications · Table 3, row ProkBERT-mini, column Specificity; XML row2 column5

Source checking is not independent reproduction.

0.74 MCC

Unit: dimensionless · Direction: higher

Uncertainty: unreported

Scored: Not reported · Eligible: Not reported

source checkedProkBERT family: genomic language models for microbiome applications · Table 3, row ProkBERT-mini, column MCC; XML row2 column3

Source checking is not independent reproduction.

How it works

How the evaluated method works

A transformer encoder uses local-context-aware 6-mer tokenisation for the ProkBERT-mini variant, followed by self-supervised pretraining and task-specific fine-tuning.

SourcesProkBERT family: genomic language models for microbiome applications · 2 Materials and methods/2.1 Sequence data/2.1.1 Sequence segmentation and tokenization (paragraph 6); 2 Materials and methods (paragraph 1)
What was evaluated

The linked evaluation record identifies ProkBERT-mini: E. coli sigma70 promoter prediction. Its dataset, split, adaptation and evidence origin remain attached to the reported results.

SourcesProkBERT family: genomic language models for microbiome applications · The named evaluation’s methods and comparison table; exact preserved evaluation IDs: evaluation-lit-033

Strengths and limitations

Strengths and considerations

  • Targets microbial sequence distributions and evaluates promoter-related downstream tasks.
    SourcesProkBERT family: genomic language models for microbiome applications · 2 Materials and methods/2.3 Application I: bacterial promoter prediction/2.3.1 Dataset overview/2.3.1.2 Dataset construction for multispecies train, test and validation sets (paragraph 6); 3 Results and discussion/3.4 ProkBERT swiftly and accurately identifies phage sequences, even in challenging settings (paragraph 2)

Limitations and conditions

  • The mini checkpoint and tokenizer settings define the evaluated system; results cannot be transferred to every ProkBERT variant.
    SourcesProkBERT family: genomic language models for microbiome applications · 2 Materials and methods/2.1 Sequence data/2.1.1 Sequence segmentation and tokenization (paragraph 6); 3 Results and discussion/3.3 ProkBERT performs accurately and robustly in promoter sequence recognition (paragraph 8)
Profile review details

Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.

Stable record: reported-model-4438513d9cd42c

Specifications

Inputs, training, access and other details

Explanatory profile: limited source coverage · Automated source review, 2026-09-16. Review applies to the cited claims; unresolved fields are listed below. Numerical results retain their own review status.

Inputs, outputs and configuration
PropertyDescription and evidence
Model typeTransformer representation pipeline; this record is the paper-specific evaluated configuration.
SourcesProkBERT family: genomic language models for microbiome applications · 2 Materials and methods/2.1 Sequence data/2.1.1 Sequence segmentation and tokenization (paragraph 6); 2 Materials and methods (paragraph 1)
Architecture / procedureA transformer encoder uses local-context-aware 6-mer tokenisation for the ProkBERT-mini variant, followed by self-supervised pretraining and task-specific fine-tuning.
SourcesProkBERT family: genomic language models for microbiome applications · 2 Materials and methods/2.1 Sequence data/2.1.1 Sequence segmentation and tokenization (paragraph 6); 2 Materials and methods (paragraph 1)
Biological inputsMicrobial DNA sequences
SourcesProkBERT family: genomic language models for microbiome applications · 1 Introduction (paragraph 7); 4 Conclusion (paragraph 10)
OutputsDNA embeddings and task-specific genomic classifications
SourcesProkBERT family: genomic language models for microbiome applications · 2 Materials and methods/2.5 Applied metrics (paragraph 1); 3 Results and discussion/3.1 ProkBERT's learned representations capture genomic structure and phylogeny (paragraph 4)
ParametersApproximately 20 million parameters.
SourcesProkBERT family: genomic language models for microbiome applications · 4 Conclusion (paragraph 8); 2 Materials and methods/2.2 Pretraining and learning sequence representations/2.2.2 Training process/2.2.2.2 Training phases and configuration (paragraph 1)
Known versions / configurationProkBERT-mini is the comparison-table label; that label does not specify an immutable weight revision. · Not reported in inspected sources
SourcesProkBERT family: genomic language models for microbiome applications · Model identification in the comparison table and corresponding Methods; immutable checkpoint revision is not supplied by the table label.
Training data / fittingUnlabelled microbial genome sequences followed by supervised task data, as specified in the paper.
SourcesProkBERT family: genomic language models for microbiome applications · 2 Materials and methods/2.4 Application II: phage sequence analysis/2.4.2 Model training for phage sequence analysis (paragraph 1); 2 Materials and methods/2.2 Pretraining and learning sequence representations/2.2.4 Analysis of encoder outputs (paragraph 6)
Context limitsApproximately 1 kb for ProkBERT-mini; the distinct mini-long variant supports approximately 2 kb.
SourcesProkBERT family: genomic language models for microbiome applications · 2 Materials and methods/2.1 Sequence data/2.1.1 Sequence segmentation and tokenization (paragraph 6); Table T3 (paragraph 1)
AccessOfficial study implementation and usage documentation: https://github.com/nbrg-ppcu/prokbert/blob/8670ae92b816cff158a0b85647a8dea122e251eb/README.md. This pinned documentation revision is not automatically the evaluated weight revision.
Sourcesnbrg-ppcu/prokbert README.md · README.md; installation, model download and usage instructions
Code licenceMIT (study repository code at the cited revision; this does not establish every dependency or historical checkpoint licence).
Sourcesnbrg-ppcu/prokbert LICENSE · LICENSE; complete licence text
Weights licenceThe inspected model-access documentation does not explicitly identify terms for this exact evaluated checkpoint or fitted head; repository code terms are shown separately. · Not reported in inspected sources
Sourcesnbrg-ppcu/prokbert README.md · README.md; checkpoint/access documentation and licence scope

Evidence table

Inspect claims, sources and review details

Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.

One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.

19 evidence rows matching the loaded filters

Claims, original sources and review scope · Release 2026-09-17-d277315f7d76
Property and statementOriginal source and locationReview and provenance
Diagram caption

Conceptual input–method–output guide. Check the procedure text and linked evaluation for fitted components, additional inputs and exact settings.

Individual claims
ProkBERT family: genomic language models for microbiome applications

Original source ↗

2 Materials and methods/2.1 Sequence data/2.1.1 Sequence segmentation and tokenization (paragraph 6); 2 Materials and methods (paragraph 1)

Version: PMC10810988.1
Retrieved: 2026-09-16T10:33:36.197Z

source checked

automated source review · 2026-09-16

Audit details

Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.

Field: attributes.profile.diagram.caption

Source artifact SHA-256: 8610e2a54aa877c8dc565a9cdb6e82099f284c5e0907a52cab18d994ea732436

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Diagram steps

["Microbial DNA sequences","ProkBERT-mini","DNA embeddings and task-specific genomic classifications"]

Individual claims
ProkBERT family: genomic language models for microbiome applications

Original source ↗

2 Materials and methods/2.1 Sequence data/2.1.1 Sequence segmentation and tokenization (paragraph 6); 2 Materials and methods (paragraph 1)

Version: PMC10810988.1
Retrieved: 2026-09-16T10:33:36.197Z

source checked

automated source review · 2026-09-16

Audit details

Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.

Field: attributes.profile.diagram.steps

Source artifact SHA-256: 8610e2a54aa877c8dc565a9cdb6e82099f284c5e0907a52cab18d994ea732436

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Diagram title

Evaluated procedure (conceptual)

Individual claims
ProkBERT family: genomic language models for microbiome applications

Original source ↗

2 Materials and methods/2.1 Sequence data/2.1.1 Sequence segmentation and tokenization (paragraph 6); 2 Materials and methods (paragraph 1)

Version: PMC10810988.1
Retrieved: 2026-09-16T10:33:36.197Z

source checked

automated source review · 2026-09-16

Audit details

Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.

Field: attributes.profile.diagram.title

Source artifact SHA-256: 8610e2a54aa877c8dc565a9cdb6e82099f284c5e0907a52cab18d994ea732436

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Model type

Transformer representation pipeline; this record is the paper-specific evaluated configuration.

Individual claims
ProkBERT family: genomic language models for microbiome applications

Original source ↗

2 Materials and methods/2.1 Sequence data/2.1.1 Sequence segmentation and tokenization (paragraph 6); 2 Materials and methods (paragraph 1)

Version: PMC10810988.1
Retrieved: 2026-09-16T10:33:36.197Z

source checked

automated source review · 2026-09-16

Audit details

Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.

Field: attributes.profile.facts.0.value

Source artifact SHA-256: 8610e2a54aa877c8dc565a9cdb6e82099f284c5e0907a52cab18d994ea732436

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Architecture / procedure

A transformer encoder uses local-context-aware 6-mer tokenisation for the ProkBERT-mini variant, followed by self-supervised pretraining and task-specific fine-tuning.

Individual claims
ProkBERT family: genomic language models for microbiome applications

Original source ↗

2 Materials and methods/2.1 Sequence data/2.1.1 Sequence segmentation and tokenization (paragraph 6); 2 Materials and methods (paragraph 1)

Version: PMC10810988.1
Retrieved: 2026-09-16T10:33:36.197Z

source checked

automated source review · 2026-09-16

Audit details

Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.

Field: attributes.profile.facts.1.value

Source artifact SHA-256: 8610e2a54aa877c8dc565a9cdb6e82099f284c5e0907a52cab18d994ea732436

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Weights licence

The inspected model-access documentation does not explicitly identify terms for this exact evaluated checkpoint or fitted head; repository code terms are shown separately.

Individual claims
nbrg-ppcu/prokbert README.md

Original source ↗

README.md; checkpoint/access documentation and licence scope

Version: 8670ae92b816cff158a0b85647a8dea122e251eb
Retrieved: 2026-09-16T19:54:21.078483+00:00

unreported

automated source review · 2026-09-16

Audit details

Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.

Field: attributes.profile.facts.10.value

Source artifact SHA-256: 29de39c6ad006ce704ab14240cfd97af93da411fb63ec89ebe499c2646928cfc

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Biological inputs

Microbial DNA sequences

Individual claims
ProkBERT family: genomic language models for microbiome applications

Original source ↗

1 Introduction (paragraph 7); 4 Conclusion (paragraph 10)

Version: PMC10810988.1
Retrieved: 2026-09-16T10:33:36.197Z

source checked

automated source review · 2026-09-16

Audit details

Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.

Field: attributes.profile.facts.2.value

Source artifact SHA-256: 8610e2a54aa877c8dc565a9cdb6e82099f284c5e0907a52cab18d994ea732436

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Outputs

DNA embeddings and task-specific genomic classifications

Individual claims
ProkBERT family: genomic language models for microbiome applications

Original source ↗

2 Materials and methods/2.5 Applied metrics (paragraph 1); 3 Results and discussion/3.1 ProkBERT's learned representations capture genomic structure and phylogeny (paragraph 4)

Version: PMC10810988.1
Retrieved: 2026-09-16T10:33:36.197Z

source checked

automated source review · 2026-09-16

Audit details

Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.

Field: attributes.profile.facts.3.value

Source artifact SHA-256: 8610e2a54aa877c8dc565a9cdb6e82099f284c5e0907a52cab18d994ea732436

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Parameters

Approximately 20 million parameters.

Individual claims
ProkBERT family: genomic language models for microbiome applications

Original source ↗

4 Conclusion (paragraph 8); 2 Materials and methods/2.2 Pretraining and learning sequence representations/2.2.2 Training process/2.2.2.2 Training phases and configuration (paragraph 1)

Version: PMC10810988.1
Retrieved: 2026-09-16T10:33:36.197Z

source checked

automated source review · 2026-09-16

Audit details

Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.

Field: attributes.profile.facts.4.value

Source artifact SHA-256: 8610e2a54aa877c8dc565a9cdb6e82099f284c5e0907a52cab18d994ea732436

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Known versions / configuration

ProkBERT-mini is the comparison-table label; that label does not specify an immutable weight revision.

Individual claims
ProkBERT family: genomic language models for microbiome applications

Original source ↗

Model identification in the comparison table and corresponding Methods; immutable checkpoint revision is not supplied by the table label.

Version: PMC10810988.1
Retrieved: 2026-09-16T10:33:36.197Z

unreported

automated source review · 2026-09-16

Audit details

Primary full text and the available official implementation/model documentation were inspected. Explanatory claims are source-backed; unresolved exact-configuration metadata is labelled explicitly. This is automated review, not a human review or independent benchmark reproduction.

Field: attributes.profile.facts.5.value

Source artifact SHA-256: 8610e2a54aa877c8dc565a9cdb6e82099f284c5e0907a52cab18d994ea732436

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Sources and history

Release 2026-09-17-d277315f7d76 · Record review: needs review

3 source records and release historyDownload this release
Technical metadata and extraction receipts

Stable ID: reported-model-4438513d9cd42c

areas
microbes-communities
entity level
method
version
Not reported
reported name
ProkBERT-mini
historical missing metadata
version: not_reported_in_legacy_extract; checkpoint revision: not_reported_in_legacy_extract; training data: not_reported_in_legacy_extract; licence: not_reported_in_legacy_extract
metadata review scope
historical_missing_metadata preserves the original discovery state. Current descriptive evidence and missingness are recorded in profile.facts; numerical-result review is separate.
legacy kinds
model
entity classification
review date: 2026-09-17; rationale: This source-scoped entry preserves the method/configuration actually named in an evaluation. It is neither a global family identity nor proof of an immutable checkpoint; the linked evaluation retains adaptation, fitting and scoring details.; source ids: prokbert-2024; source locator: 2 Materials and methods/2.1 Sequence data/2.1.1 Sequence segmentation and tokenization (paragraph 6); 2 Materials and methods (paragraph 1) | 2 Materials and methods/2.1 Sequence data/2.1.1 Sequence segmentation and tokenization (paragraph 6); 4 Conclusion (paragraph 5); ambiguities: Configuration means the source-labelled evaluated identity. It does not establish missing checkpoint hashes, default settings or equivalence to same-named records in other papers.
Related records

Suggest a correction