GenomeOcean-4B on GENEB MLP-PROBE: Average macro-MCC across the 13 representative tasks, MLP probe
GENEB evaluation of GenomeOcean-4B on Average macro-MCC across the 13 representative tasks, MLP probe, scored with Macro-MCC.
Methods and reproduction
GENEB evaluation of GenomeOcean-4B on Average macro-MCC across the 13 representative tasks, MLP probe, scored with Macro-MCC.
- task
- GENEB MLP-PROBE: Average macro-MCC across the 13 representative tasks, MLP probe
- configuration
- GenomeOcean-4B
- dataset subset
- GENEB representative task subset (GENEB split)
- Split
- Not reported
- Adaptation
- Not reported
- Scoring implementation
- Macro-MCC
No execution recipe has been verified for this exact configuration and evaluation. A benchmark's general instructions may use different inputs, splits or model settings.
Reproducing this published result requires matching its model configuration, data, split and scorer. Source checking or a successful smoke test does not establish score reproduction.
Evaluation procedure
Frozen embeddings scored with a probe over the thirteen representative tasks, one from each functional category (Table 7).
- Configuration
- GenomeOcean-4B
- Task
- GENEB MLP-PROBE: Average macro-MCC across the 13 representative tasks, MLP probe
- Dataset subset
- GENEB representative task subset (GENEB split)
- origin
- Author-reported evaluation
- configuration
- Not reported
- protocol id
- geneb-task-mlp-probe
- metric implementation
- Macro-MCC
Metadata review: source checked. Unreported conditions prevent automatic comparisons.
Evaluation results
Release 2026-09-17-134cd1815de8 · 1 evaluation · 1 metric row. Different protocols are not a single leaderboard. Where several source tables report the same metric, the published comparisons above offer a pooled view that names what it does not hold constant.
| Metric and finding | Coverage and uncertainty | Evidence |
|---|---|---|
| GenomeOcean-4B on GENEB MLP-PROBE: Average macro-MCC across the 13 representative tasks, MLP probe Configuration: GenomeOcean-4BTask: GENEB MLP-PROBE: Average macro-MCC across the 13 representative tasks, MLP probeDataset subset: GENEB representative task subset (GENEB split) Frozen embeddings scored with a probe over the thirteen representative tasks, one from each functional category (Table 7). Author-reported evaluation · Evaluation metadata: source checked | ||
| 0.552 macro_mcc Unit: correlation · Direction: higher | Uncertainty: Not reported Scored: Not reported · Eligible: Not reported | source checkedGENEB: Why Genomic Models Are Hard to Compare · Table 8, row(GenomeOcean-4B), column(MLP MCC) Source checking is not independent reproduction. |
Evidence table
Inspect claims, sources and review details
Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.
One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.
10 evidence rows matching the loaded filters
| Property and statement | Original source and location | Review and provenance |
|---|---|---|
| attributes.comparison.metric_implementation Macro-MCC Context-only references | GENEB: Why Genomic Models Are Hard to Compare Table 8, row(GenomeOcean-4B), column(MLP MCC) Version: 2606.04525v1 | not individually reviewed No individual claim review recorded author reported Audit detailsField: Source artifact SHA-256: Hash scope: Exact retrieved primary paper artifact bytes. |
| attributes.comparison.protocol_id geneb-task-mlp-probe Context-only references | GENEB: Why Genomic Models Are Hard to Compare Table 8, row(GenomeOcean-4B), column(MLP MCC) Version: 2606.04525v1 | not individually reviewed No individual claim review recorded author reported Audit detailsField: Source artifact SHA-256: Hash scope: Exact retrieved primary paper artifact bytes. |
| attributes.origin author_reported Context-only references | GENEB: Why Genomic Models Are Hard to Compare Table 8, row(GenomeOcean-4B), column(MLP MCC) Version: 2606.04525v1 | not individually reviewed No individual claim review recorded author reported Audit detailsField: Source artifact SHA-256: Hash scope: Exact retrieved primary paper artifact bytes. |
| attributes.protocol Frozen embeddings scored with a probe over the thirteen representative tasks, one from each functional category (Table 7). Context-only references | GENEB: Why Genomic Models Are Hard to Compare Table 8, row(GenomeOcean-4B), column(MLP MCC) Version: 2606.04525v1 | not individually reviewed No individual claim review recorded author reported Audit detailsField: Source artifact SHA-256: Hash scope: Exact retrieved primary paper artifact bytes. |
| attributes.source_locator Table 8, row(GenomeOcean-4B), column(MLP MCC) Context-only references | GENEB: Why Genomic Models Are Hard to Compare Table 8, row(GenomeOcean-4B), column(MLP MCC) Version: 2606.04525v1 | not individually reviewed No individual claim review recorded author reported Audit detailsField: Source artifact SHA-256: Hash scope: Exact retrieved primary paper artifact bytes. |
| description GENEB evaluation of GenomeOcean-4B on Average macro-MCC across the 13 representative tasks, MLP probe, scored with Macro-MCC. Context-only references | GENEB: Why Genomic Models Are Hard to Compare Table 8, row(GenomeOcean-4B), column(MLP MCC) Version: 2606.04525v1 | not individually reviewed No individual claim review recorded author reported Audit detailsField: Source artifact SHA-256: Hash scope: Exact retrieved primary paper artifact bytes. |
| Relationship: benchmark geneb-task-mlp-probe Context-only references | GENEB: Why Genomic Models Are Hard to Compare Table 8, row(GenomeOcean-4B), column(MLP MCC) Version: 2606.04525v1 | not individually reviewed No individual claim review recorded author reported Audit detailsField: Source artifact SHA-256: Hash scope: Exact retrieved primary paper artifact bytes. |
| Relationship: dataset geneb-dataset-geneb-representative-task-subset Context-only references | GENEB: Why Genomic Models Are Hard to Compare Table 8, row(GenomeOcean-4B), column(MLP MCC) Version: 2606.04525v1 | not individually reviewed No individual claim review recorded author reported Audit detailsField: Source artifact SHA-256: Hash scope: Exact retrieved primary paper artifact bytes. |
| Relationship: model geneb-method-genomeocean-4b Context-only references | GENEB: Why Genomic Models Are Hard to Compare Table 8, row(GenomeOcean-4B), column(MLP MCC) Version: 2606.04525v1 | not individually reviewed No individual claim review recorded author reported Audit detailsField: Source artifact SHA-256: Hash scope: Exact retrieved primary paper artifact bytes. |
| name GenomeOcean-4B on GENEB MLP-PROBE: Average macro-MCC across the 13 representative tasks, MLP probe Context-only references | GENEB: Why Genomic Models Are Hard to Compare Table 8, row(GenomeOcean-4B), column(MLP MCC) Version: 2606.04525v1 | not individually reviewed No individual claim review recorded author reported Audit detailsField: Source artifact SHA-256: Hash scope: Exact retrieved primary paper artifact bytes. |
Sources and history
Release 2026-09-17-134cd1815de8 · Record review: source checked
1 source records and release history
- GENEB: Why Genomic Models Are Hard to Compare · Original source · 2606.04525v1
Technical metadata and extraction receipts
Stable ID: geneb-evaluation-genomeocean-4b-mlp-probe
- areas
- dna-genomes
- tasks
- Average macro-MCC across the 13 representative tasks, MLP probe
- origin
- author_reported
- protocol
- Frozen embeddings scored with a probe over the thirteen representative tasks, one from each functional category (Table 7).
- comparison
- protocol id: geneb-task-mlp-probe; metric implementation: Macro-MCC
- missing metadata
- checkpoint revision: unreported; seeds: unreported; budget: unreported; split manifest: unextracted
- source locator
- Table 8, row(GenomeOcean-4B), column(MLP MCC)