rewire.it
Protocol

Gene-MTEB genomic embedding benchmark Human-Virus-2: Human-Virus-2 genomic embedding evaluation

Human-Virus-2 genomic embedding evaluation. Scored with accuracy on Human-Virus-2 Gene-MTEB data used in METAGENE-1 Table 3. Frozen model last-hidden-state embeddings, mean pooled. Logistic regression for eight classification tasks and mini-batch k-means for eight clustering tasks. Source paper Section 5.3. Current official gene-mteb implementation corroborates task identities but its commit is not proven identical to the paper evaluator. Exact dataset release, split hash, scored counts and evaluator commit are unreported in the paper. Do not substitute current mutable HF main for the original split.

5 evaluations · 5 metric rows

Overview

Human-Virus-2 genomic embedding evaluation. Scored with accuracy on Human-Virus-2 Gene-MTEB data used in METAGENE-1 Table 3. Frozen model last-hidden-state embeddings, mean pooled. Logistic regression for eight classification tasks and mini-batch k-means for eight clustering tasks. Source paper Section 5.3. Current official gene-mteb implementation corroborates task identities but its commit is not proven identical to the paper evaluator. Exact dataset release, split hash, scored counts and evaluator commit are unreported in the paper. Do not substitute current mutable HF main for the original split.

Consult the linked sources for architecture or protocol details. Missing evidence is not evidence of a missing capability.

Results

Each comparison retains its reviewed evaluation scope, dataset and metric. Results are shown without a pooled ranking.

Gene-MTEB genomic embedding benchmark Human-Virus-2: Human-Virus-2 genomic embedding evaluation

accuracy (fraction) · Higher values are better.

Every method Gene-MTEB genomic embedding benchmark reports on Human-Virus-2 genomic embedding evaluation, scored with accuracy on Human-Virus-2 Gene-MTEB data used in METAGENE-1 Table 3.

Gene-MTEB genomic embedding benchmark Human-Virus-2: Human-Virus-2 genomic embedding evaluation · Human-Virus-2 Gene-MTEB data used in METAGENE-1 Table 3 (Gene-MTEB genomic embedding benchmark split)

Evidence origin: Author-reported evaluation. Numerical source review does not establish independent reproduction.

METAGENE-1 v1: Gene-MTEB complete component tasks · Table 3 (S5.T3), row Human-Virus-2, column DNABERT-2; Section 5.3 paragraphs S5.SS3.p2.1 and S5.SS3.p3.1

All 16 component tasks and five methods are included. Twenty-five printed aggregate cells are retained in a companion receipt, excluded from task-level result counts to avoid counting derived averages as extra experiments.

All comparison limitations (5)
  • All 16 component tasks and five methods are included. Twenty-five printed aggregate cells are retained in a companion receipt, excluded from task-level result counts to avoid counting derived averages as extra experiments.
  • Classification accuracy and clustering V-measure must not be combined into a universal performance metric.
  • Paper source does not pin exact evaluator/data revisions or per-task denominators. Current code metadata uses mutable dataset revision main; its development metadata is not asserted to be the evaluated paper version.
  • Paper calls embeddings zero-shot; this does not mean classification is label-free: logistic-regression probes use training labels.
  • No uncertainty intervals or repeats are printed.

Automated source review: 2026-09-23.

No unavailable values; missing scores remain labelled and are never plotted as zero.

Showing 5 of 5 matching rows.

Dots show point estimates. Whiskers show only explicitly defined uncertainty (standard deviation, standard error or a labelled interval); their definitions remain in Table. Unresolved uncertainty is not plotted. Differences do not establish statistical significance.

Methods and evaluation design

Procedure, tasks and evaluated configurations

Evaluation design

Benchmarks bring together tasks and protocols. A task describes the biological question; a protocol defines a particular test.

Benchmarks

These source-backed links do not make different protocols or scores interchangeable.

Recorded evaluations

Each evaluation records what was tested and under which conditions.

Baseline coverage

Reference methods help show what a model adds beyond simple controls. We track a null control and a conventional method for each protocol.

0 of 2 active baseline roles have published Rewire measurements in this release. Measurements on a selected protocol do not establish coverage of an entire suite.

No execution recipe linked to this protocol. Recipe availability does not establish a completed evaluation.

Author-reported evaluations
5

Literature evidence is not a Rewire measurement. Executed but unpublished runs and private review status are not included.

Null control

Proposed control: requires review

Unintegrated representation using protocol preprocessing

Protocol-specific applicability, permitted inputs, access, split, evaluator and execution requirements need review before implementation or execution.

This is a suggested selection rule, not a validated method or a measured score.

Conventional reference

Proposed control: requires review

Protocol-specific conventional integration method

Protocol-specific applicability, permitted inputs, access, split, evaluator and execution requirements need review before implementation or execution.

This is a suggested selection rule, not a validated method or a measured score.

Protocol coverage CSV · Model evaluation matrix · Source table · Release and checksums

Coverage is derived from release 2026-09-23-2b89723c6dd9. Source citations describe the original records; they do not validate an unreviewed baseline proposal. No results have been generated by this audit.

Run instructions

No runnable recipe has been reviewed for this protocol. Dataset access, model requirements, licences and compute requirements must be checked against its sources before execution.

Strengths, limitations and unresolved questions

Evidence

Source checking verifies the cited claim or transcription. It does not establish independent reproduction.

Evidence table

Inspect claims, sources and review details

Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.

One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.

1 evidence row matching the loaded filters

Claims, original sources and review scope · Release 2026-09-23-2b89723c6dd9
Property and statementOriginal source and locationReview and provenance
Relationship: part of
discovery-benchmark-gene-mteb
Individual claims
METAGENE-1 v1: Gene-MTEB complete component tasks

Original source ↗

Table 3 (S5.T3), row Human-Virus-2, column DNABERT-2; Section 5.3 paragraphs S5.SS3.p2.1 and S5.SS3.p3.1

Version: arXiv:2501.02045v1
Retrieved: 2026-09-23T11:22:00.101850+00:00

source checked

automated source review · 2026-09-23

Audit details

Primary-source transcription with no human sign-off and no independent reproduction.

Field: links:part_of:discovery-benchmark-gene-mteb

Claim: metagene-gene-mteb-association-human-virus-2

Source artifact SHA-256: a426c77d137138cfc33410a1e54a1cce521007bdc9e3176f4774d0105424a7fe

Hash scope: Hash scope not separately documented; inspect source record

Inspected artifact

Sources and history

View linked audit checks and correction history

Release 2026-09-23-2b89723c6dd9 · Record review: source checked

1 source records and release historyDownload this release
Technical metadata and extraction receipts

Stable ID: metagene-gene-mteb-task-human-virus-2

areas
dna-genomes
tasks
Human-Virus-2 genomic embedding evaluation
metric
accuracy
metric direction
higher
dataset
Human-Virus-2 Gene-MTEB data used in METAGENE-1 Table 3
protocol
Frozen model last-hidden-state embeddings, mean pooled. Logistic regression for eight classification tasks and mini-batch k-means for eight clustering tasks. Source paper Section 5.3. Current official gene-mteb implementation corroborates task identities but its commit is not proven identical to the paper evaluator. Exact dataset release, split hash, scored counts and evaluator commit are unreported in the paper. Do not substitute current mutable HF main for the original split.
source locator
Table 3 (S5.T3), row Human-Virus-2, column DNABERT-2; Section 5.3 paragraphs S5.SS3.p2.1 and S5.SS3.p3.1
comparison panels
id: metagene-gene-mteb-panel-human-virus-2; title: Gene-MTEB genomic embedding benchmark Human-Virus-2: Human-Virus-2 genomic embedding evaluation; protocol id: metagene-gene-mteb-task-human-virus-2; dataset id: metagene-gene-mteb-dataset-human-virus-2-gene-mteb-data-used-in-metagene-1-table-3; metric: accuracy; unit: fraction; direction: higher; result ids: metagene-gene-mteb-result-dnabert-2-human-virus-2-accuracy; metagene-gene-mteb-result-dnabert-s-human-virus-2-accuracy; metagene-gene-mteb-result-nt-2-5b-multi-human-virus-2-accuracy; metagene-gene-mteb-result-nt-2-5b-1000g-human-virus-2-accuracy; metagene-gene-mteb-result-metagene-1-human-virus-2-accuracy; source ids: coverage-source-metagene-paper-v1; source locator: Table 3 (S5.T3), row Human-Virus-2, column DNABERT-2; Section 5.3 paragraphs S5.SS3.p2.1 and S5.SS3.p3.1; context: Every method Gene-MTEB genomic embedding benchmark reports on Human-Virus-2 genomic embedding evaluation, scored with accuracy on Human-Virus-2 Gene-MTEB data used in METAGENE-1 Table 3.; caveats: All 16 component tasks and five methods are included. Twenty-five printed aggregate cells are retained in a companion receipt, excluded from task-level result counts to avoid counting derived averages as extra experiments.; Classification accuracy and clustering V-measure must not be combined into a universal performance metric.; Paper source does not pin exact evaluator/data revisions or per-task denominators. Current code metadata uses mutable dataset revision main; its development metadata is not asserted to be the evaluated paper version.; Paper calls embeddings zero-shot; this does not mean classification is label-free: logistic-regression probes use training labels.; No uncertainty intervals or repeats are printed.; review: method: automated_source_review; date: 2026-09-23
entity level
protocol
Related records

Suggest a correction