GENEB representative task subset (GENEB split)
The split of GENEB representative task subset that GENEB evaluated on. The upstream dataset release is not catalogued here, so no claim is made that this matches its original splits.
Subset and evaluation context
This record describes a particular subset or cohort used in an evaluation. Its results do not describe the full dataset.
Evaluation results
Release 2026-09-17-134cd1815de8 · 22 evaluations · 22 metric rows. Different protocols are not a single leaderboard. Where several source tables report the same metric, the published comparisons above offer a pooled view that names what it does not hold constant.
| Metric and finding | Coverage and uncertainty | Evidence |
|---|---|---|
| GENA-LM-Large-T2T on GENEB LINEAR-PROBE: Average macro-MCC across the 13 representative tasks, linear probe Configuration: GENA-LM-Large-T2TTask: GENEB LINEAR-PROBE: Average macro-MCC across the 13 representative tasks, linear probeDataset subset: GENEB representative task subset (GENEB split) Frozen embeddings scored with a probe over the thirteen representative tasks, one from each functional category (Table 7). Author-reported evaluation · Evaluation metadata: source checked | ||
| 0.530 macro_mcc Unit: correlation · Direction: higher | Uncertainty: Not reported Scored: Not reported · Eligible: Not reported | source checkedGENEB: Why Genomic Models Are Hard to Compare · Table 8, row(GENA-LM-Large-T2T), column(Linear MCC) Source checking is not independent reproduction. |
| GENA-LM-Large-T2T on GENEB MLP-PROBE: Average macro-MCC across the 13 representative tasks, MLP probe Configuration: GENA-LM-Large-T2TTask: GENEB MLP-PROBE: Average macro-MCC across the 13 representative tasks, MLP probeDataset subset: GENEB representative task subset (GENEB split) Frozen embeddings scored with a probe over the thirteen representative tasks, one from each functional category (Table 7). Author-reported evaluation · Evaluation metadata: source checked | ||
| 0.535 macro_mcc Unit: correlation · Direction: higher | Uncertainty: Not reported Scored: Not reported · Eligible: Not reported | source checkedGENEB: Why Genomic Models Are Hard to Compare · Table 8, row(GENA-LM-Large-T2T), column(MLP MCC) Source checking is not independent reproduction. |
| GENERator-Eukaryote-1.2B on GENEB LINEAR-PROBE: Average macro-MCC across the 13 representative tasks, linear probe Configuration: GENERator-Eukaryote-1.2BTask: GENEB LINEAR-PROBE: Average macro-MCC across the 13 representative tasks, linear probeDataset subset: GENEB representative task subset (GENEB split) Frozen embeddings scored with a probe over the thirteen representative tasks, one from each functional category (Table 7). Author-reported evaluation · Evaluation metadata: source checked | ||
| 0.579 macro_mcc Unit: correlation · Direction: higher | Uncertainty: Not reported Scored: Not reported · Eligible: Not reported | source checkedGENEB: Why Genomic Models Are Hard to Compare · Table 8, row(GENERator-Eukaryote-1.2B), column(Linear MCC) Source checking is not independent reproduction. |
| GENERator-Eukaryote-1.2B on GENEB MLP-PROBE: Average macro-MCC across the 13 representative tasks, MLP probe Configuration: GENERator-Eukaryote-1.2BTask: GENEB MLP-PROBE: Average macro-MCC across the 13 representative tasks, MLP probeDataset subset: GENEB representative task subset (GENEB split) Frozen embeddings scored with a probe over the thirteen representative tasks, one from each functional category (Table 7). Author-reported evaluation · Evaluation metadata: source checked | ||
| 0.595 macro_mcc Unit: correlation · Direction: higher | Uncertainty: Not reported Scored: Not reported · Eligible: Not reported | source checkedGENEB: Why Genomic Models Are Hard to Compare · Table 8, row(GENERator-Eukaryote-1.2B), column(MLP MCC) Source checking is not independent reproduction. |
| GENERator-Eukaryote-3B on GENEB LINEAR-PROBE: Average macro-MCC across the 13 representative tasks, linear probe Configuration: GENERator-Eukaryote-3BTask: GENEB LINEAR-PROBE: Average macro-MCC across the 13 representative tasks, linear probeDataset subset: GENEB representative task subset (GENEB split) Frozen embeddings scored with a probe over the thirteen representative tasks, one from each functional category (Table 7). Author-reported evaluation · Evaluation metadata: source checked | ||
| 0.605 macro_mcc Unit: correlation · Direction: higher | Uncertainty: Not reported Scored: Not reported · Eligible: Not reported | source checkedGENEB: Why Genomic Models Are Hard to Compare · Table 8, row(GENERator-Eukaryote-3B), column(Linear MCC) Source checking is not independent reproduction. |
| GENERator-Eukaryote-3B on GENEB MLP-PROBE: Average macro-MCC across the 13 representative tasks, MLP probe Configuration: GENERator-Eukaryote-3BTask: GENEB MLP-PROBE: Average macro-MCC across the 13 representative tasks, MLP probeDataset subset: GENEB representative task subset (GENEB split) Frozen embeddings scored with a probe over the thirteen representative tasks, one from each functional category (Table 7). Author-reported evaluation · Evaluation metadata: source checked | ||
| 0.609 macro_mcc Unit: correlation · Direction: higher | Uncertainty: Not reported Scored: Not reported · Eligible: Not reported | source checkedGENEB: Why Genomic Models Are Hard to Compare · Table 8, row(GENERator-Eukaryote-3B), column(MLP MCC) Source checking is not independent reproduction. |
| GenomeOcean-4B on GENEB LINEAR-PROBE: Average macro-MCC across the 13 representative tasks, linear probe Configuration: GenomeOcean-4BTask: GENEB LINEAR-PROBE: Average macro-MCC across the 13 representative tasks, linear probeDataset subset: GENEB representative task subset (GENEB split) Frozen embeddings scored with a probe over the thirteen representative tasks, one from each functional category (Table 7). Author-reported evaluation · Evaluation metadata: source checked | ||
| 0.552 macro_mcc Unit: correlation · Direction: higher | Uncertainty: Not reported Scored: Not reported · Eligible: Not reported | source checkedGENEB: Why Genomic Models Are Hard to Compare · Table 8, row(GenomeOcean-4B), column(Linear MCC) Source checking is not independent reproduction. |
| GenomeOcean-4B on GENEB MLP-PROBE: Average macro-MCC across the 13 representative tasks, MLP probe Configuration: GenomeOcean-4BTask: GENEB MLP-PROBE: Average macro-MCC across the 13 representative tasks, MLP probeDataset subset: GENEB representative task subset (GENEB split) Frozen embeddings scored with a probe over the thirteen representative tasks, one from each functional category (Table 7). Author-reported evaluation · Evaluation metadata: source checked | ||
| 0.552 macro_mcc Unit: correlation · Direction: higher | Uncertainty: Not reported Scored: Not reported · Eligible: Not reported | source checkedGENEB: Why Genomic Models Are Hard to Compare · Table 8, row(GenomeOcean-4B), column(MLP MCC) Source checking is not independent reproduction. |
| GenomeOcean-500M on GENEB LINEAR-PROBE: Average macro-MCC across the 13 representative tasks, linear probe Configuration: GenomeOcean-500MTask: GENEB LINEAR-PROBE: Average macro-MCC across the 13 representative tasks, linear probeDataset subset: GENEB representative task subset (GENEB split) Frozen embeddings scored with a probe over the thirteen representative tasks, one from each functional category (Table 7). Author-reported evaluation · Evaluation metadata: source checked | ||
| 0.536 macro_mcc Unit: correlation · Direction: higher | Uncertainty: Not reported Scored: Not reported · Eligible: Not reported | source checkedGENEB: Why Genomic Models Are Hard to Compare · Table 8, row(GenomeOcean-500M), column(Linear MCC) Source checking is not independent reproduction. |
| GenomeOcean-500M on GENEB MLP-PROBE: Average macro-MCC across the 13 representative tasks, MLP probe Configuration: GenomeOcean-500MTask: GENEB MLP-PROBE: Average macro-MCC across the 13 representative tasks, MLP probeDataset subset: GENEB representative task subset (GENEB split) Frozen embeddings scored with a probe over the thirteen representative tasks, one from each functional category (Table 7). Author-reported evaluation · Evaluation metadata: source checked | ||
| 0.535 macro_mcc Unit: correlation · Direction: higher | Uncertainty: Not reported Scored: Not reported · Eligible: Not reported | source checkedGENEB: Why Genomic Models Are Hard to Compare · Table 8, row(GenomeOcean-500M), column(MLP MCC) Source checking is not independent reproduction. |
| GROVER on GENEB LINEAR-PROBE: Average macro-MCC across the 13 representative tasks, linear probe Configuration: GROVERTask: GENEB LINEAR-PROBE: Average macro-MCC across the 13 representative tasks, linear probeDataset subset: GENEB representative task subset (GENEB split) Frozen embeddings scored with a probe over the thirteen representative tasks, one from each functional category (Table 7). Author-reported evaluation · Evaluation metadata: source checked | ||
| 0.466 macro_mcc Unit: correlation · Direction: higher | Uncertainty: Not reported Scored: Not reported · Eligible: Not reported | source checkedGENEB: Why Genomic Models Are Hard to Compare · Table 8, row(GROVER), column(Linear MCC) Source checking is not independent reproduction. |
| GROVER on GENEB MLP-PROBE: Average macro-MCC across the 13 representative tasks, MLP probe Configuration: GROVERTask: GENEB MLP-PROBE: Average macro-MCC across the 13 representative tasks, MLP probeDataset subset: GENEB representative task subset (GENEB split) Frozen embeddings scored with a probe over the thirteen representative tasks, one from each functional category (Table 7). Author-reported evaluation · Evaluation metadata: source checked | ||
| 0.477 macro_mcc Unit: correlation · Direction: higher | Uncertainty: Not reported Scored: Not reported · Eligible: Not reported | source checkedGENEB: Why Genomic Models Are Hard to Compare · Table 8, row(GROVER), column(MLP MCC) Source checking is not independent reproduction. |
| HyenaDNA-Large-1M on GENEB LINEAR-PROBE: Average macro-MCC across the 13 representative tasks, linear probe Configuration: HyenaDNA-Large-1MTask: GENEB LINEAR-PROBE: Average macro-MCC across the 13 representative tasks, linear probeDataset subset: GENEB representative task subset (GENEB split) Frozen embeddings scored with a probe over the thirteen representative tasks, one from each functional category (Table 7). Author-reported evaluation · Evaluation metadata: source checked | ||
| 0.427 macro_mcc Unit: correlation · Direction: higher | Uncertainty: Not reported Scored: Not reported · Eligible: Not reported | source checkedGENEB: Why Genomic Models Are Hard to Compare · Table 8, row(HyenaDNA-Large-1M), column(Linear MCC) Source checking is not independent reproduction. |
| HyenaDNA-Large-1M on GENEB MLP-PROBE: Average macro-MCC across the 13 representative tasks, MLP probe Configuration: HyenaDNA-Large-1MTask: GENEB MLP-PROBE: Average macro-MCC across the 13 representative tasks, MLP probeDataset subset: GENEB representative task subset (GENEB split) Frozen embeddings scored with a probe over the thirteen representative tasks, one from each functional category (Table 7). Author-reported evaluation · Evaluation metadata: source checked | ||
| 0.479 macro_mcc Unit: correlation · Direction: higher | Uncertainty: Not reported Scored: Not reported · Eligible: Not reported | source checkedGENEB: Why Genomic Models Are Hard to Compare · Table 8, row(HyenaDNA-Large-1M), column(MLP MCC) Source checking is not independent reproduction. |
| LucaOne on GENEB LINEAR-PROBE: Average macro-MCC across the 13 representative tasks, linear probe Configuration: LucaOneTask: GENEB LINEAR-PROBE: Average macro-MCC across the 13 representative tasks, linear probeDataset subset: GENEB representative task subset (GENEB split) Frozen embeddings scored with a probe over the thirteen representative tasks, one from each functional category (Table 7). Author-reported evaluation · Evaluation metadata: source checked | ||
| 0.573 macro_mcc Unit: correlation · Direction: higher | Uncertainty: Not reported Scored: Not reported · Eligible: Not reported | source checkedGENEB: Why Genomic Models Are Hard to Compare · Table 8, row(LucaOne), column(Linear MCC) Source checking is not independent reproduction. |
| LucaOne on GENEB MLP-PROBE: Average macro-MCC across the 13 representative tasks, MLP probe Configuration: LucaOneTask: GENEB MLP-PROBE: Average macro-MCC across the 13 representative tasks, MLP probeDataset subset: GENEB representative task subset (GENEB split) Frozen embeddings scored with a probe over the thirteen representative tasks, one from each functional category (Table 7). Author-reported evaluation · Evaluation metadata: source checked | ||
| 0.600 macro_mcc Unit: correlation · Direction: higher | Uncertainty: Not reported Scored: Not reported · Eligible: Not reported | source checkedGENEB: Why Genomic Models Are Hard to Compare · Table 8, row(LucaOne), column(MLP MCC) Source checking is not independent reproduction. |
| MutBERT on GENEB LINEAR-PROBE: Average macro-MCC across the 13 representative tasks, linear probe Configuration: MutBERTTask: GENEB LINEAR-PROBE: Average macro-MCC across the 13 representative tasks, linear probeDataset subset: GENEB representative task subset (GENEB split) Frozen embeddings scored with a probe over the thirteen representative tasks, one from each functional category (Table 7). Author-reported evaluation · Evaluation metadata: source checked | ||
| 0.516 macro_mcc Unit: correlation · Direction: higher | Uncertainty: Not reported Scored: Not reported · Eligible: Not reported | source checkedGENEB: Why Genomic Models Are Hard to Compare · Table 8, row(MutBERT), column(Linear MCC) Source checking is not independent reproduction. |
| MutBERT on GENEB MLP-PROBE: Average macro-MCC across the 13 representative tasks, MLP probe Configuration: MutBERTTask: GENEB MLP-PROBE: Average macro-MCC across the 13 representative tasks, MLP probeDataset subset: GENEB representative task subset (GENEB split) Frozen embeddings scored with a probe over the thirteen representative tasks, one from each functional category (Table 7). Author-reported evaluation · Evaluation metadata: source checked | ||
| 0.517 macro_mcc Unit: correlation · Direction: higher | Uncertainty: Not reported Scored: Not reported · Eligible: Not reported | source checkedGENEB: Why Genomic Models Are Hard to Compare · Table 8, row(MutBERT), column(MLP MCC) Source checking is not independent reproduction. |
| NT-v2-50M-MS on GENEB LINEAR-PROBE: Average macro-MCC across the 13 representative tasks, linear probe Configuration: NT-v2-50M-MSTask: GENEB LINEAR-PROBE: Average macro-MCC across the 13 representative tasks, linear probeDataset subset: GENEB representative task subset (GENEB split) Frozen embeddings scored with a probe over the thirteen representative tasks, one from each functional category (Table 7). Author-reported evaluation · Evaluation metadata: source checked | ||
| 0.511 macro_mcc Unit: correlation · Direction: higher | Uncertainty: Not reported Scored: Not reported · Eligible: Not reported | source checkedGENEB: Why Genomic Models Are Hard to Compare · Table 8, row(NT-v2-50M-MS), column(Linear MCC) Source checking is not independent reproduction. |
| NT-v2-50M-MS on GENEB MLP-PROBE: Average macro-MCC across the 13 representative tasks, MLP probe Configuration: NT-v2-50M-MSTask: GENEB MLP-PROBE: Average macro-MCC across the 13 representative tasks, MLP probeDataset subset: GENEB representative task subset (GENEB split) Frozen embeddings scored with a probe over the thirteen representative tasks, one from each functional category (Table 7). Author-reported evaluation · Evaluation metadata: source checked | ||
| 0.521 macro_mcc Unit: correlation · Direction: higher | Uncertainty: Not reported Scored: Not reported · Eligible: Not reported | source checkedGENEB: Why Genomic Models Are Hard to Compare · Table 8, row(NT-v2-50M-MS), column(MLP MCC) Source checking is not independent reproduction. |
| Omni-DNA-1B on GENEB LINEAR-PROBE: Average macro-MCC across the 13 representative tasks, linear probe Configuration: Omni-DNA-1BTask: GENEB LINEAR-PROBE: Average macro-MCC across the 13 representative tasks, linear probeDataset subset: GENEB representative task subset (GENEB split) Frozen embeddings scored with a probe over the thirteen representative tasks, one from each functional category (Table 7). Author-reported evaluation · Evaluation metadata: source checked | ||
| 0.550 macro_mcc Unit: correlation · Direction: higher | Uncertainty: Not reported Scored: Not reported · Eligible: Not reported | source checkedGENEB: Why Genomic Models Are Hard to Compare · Table 8, row(Omni-DNA-1B), column(Linear MCC) Source checking is not independent reproduction. |
| Omni-DNA-1B on GENEB MLP-PROBE: Average macro-MCC across the 13 representative tasks, MLP probe Configuration: Omni-DNA-1BTask: GENEB MLP-PROBE: Average macro-MCC across the 13 representative tasks, MLP probeDataset subset: GENEB representative task subset (GENEB split) Frozen embeddings scored with a probe over the thirteen representative tasks, one from each functional category (Table 7). Author-reported evaluation · Evaluation metadata: source checked | ||
| 0.542 macro_mcc Unit: correlation · Direction: higher | Uncertainty: Not reported Scored: Not reported · Eligible: Not reported | source checkedGENEB: Why Genomic Models Are Hard to Compare · Table 8, row(Omni-DNA-1B), column(MLP MCC) Source checking is not independent reproduction. |
Evidence table
Inspect claims, sources and review details
Trace each statement to its source and review. A context-only reference supports the record generally; it does not verify an individual field. Source checking does not reproduce an experiment.
One row per statement and cited source. Multiple citations are not independent evaluations. Shared locators are labelled explicitly.
2 evidence rows matching the loaded filters
| Property and statement | Original source and location | Review and provenance |
|---|---|---|
| description The split of GENEB representative task subset that GENEB evaluated on. The upstream dataset release is not catalogued here, so no claim is made that this matches its original splits. Context-only references | GENEB: Why Genomic Models Are Hard to Compare No field-specific location recorded Version: 2606.04525v1 | not individually reviewed No individual claim review recorded Audit detailsField: Source artifact SHA-256: Hash scope: Exact retrieved primary paper artifact bytes. |
| name GENEB representative task subset (GENEB split) Context-only references | GENEB: Why Genomic Models Are Hard to Compare No field-specific location recorded Version: 2606.04525v1 | not individually reviewed No individual claim review recorded Audit detailsField: Source artifact SHA-256: Hash scope: Exact retrieved primary paper artifact bytes. |
Sources and history
Release 2026-09-17-134cd1815de8 · Record review: source checked
1 source records and release history
- GENEB: Why Genomic Models Are Hard to Compare · Original source · 2606.04525v1
Technical metadata and extraction receipts
Stable ID: geneb-dataset-geneb-representative-task-subset
- areas
- dna-genomes
- missing metadata
- version: unreported; url: unextracted
Related records
- dataset: GENA-LM-Large-T2T on GENEB LINEAR-PROBE: Average macro-MCC across the 13 representative tasks, linear probe
- dataset: GENA-LM-Large-T2T on GENEB MLP-PROBE: Average macro-MCC across the 13 representative tasks, MLP probe
- dataset: GENERator-Eukaryote-1.2B on GENEB LINEAR-PROBE: Average macro-MCC across the 13 representative tasks, linear probe
- dataset: GENERator-Eukaryote-1.2B on GENEB MLP-PROBE: Average macro-MCC across the 13 representative tasks, MLP probe
- dataset: GENERator-Eukaryote-3B on GENEB LINEAR-PROBE: Average macro-MCC across the 13 representative tasks, linear probe
- dataset: GENERator-Eukaryote-3B on GENEB MLP-PROBE: Average macro-MCC across the 13 representative tasks, MLP probe
- dataset: GenomeOcean-4B on GENEB LINEAR-PROBE: Average macro-MCC across the 13 representative tasks, linear probe
- dataset: GenomeOcean-4B on GENEB MLP-PROBE: Average macro-MCC across the 13 representative tasks, MLP probe
- dataset: GenomeOcean-500M on GENEB LINEAR-PROBE: Average macro-MCC across the 13 representative tasks, linear probe
- dataset: GenomeOcean-500M on GENEB MLP-PROBE: Average macro-MCC across the 13 representative tasks, MLP probe
- dataset: GROVER on GENEB LINEAR-PROBE: Average macro-MCC across the 13 representative tasks, linear probe
- dataset: GROVER on GENEB MLP-PROBE: Average macro-MCC across the 13 representative tasks, MLP probe
- dataset: HyenaDNA-Large-1M on GENEB LINEAR-PROBE: Average macro-MCC across the 13 representative tasks, linear probe
- dataset: HyenaDNA-Large-1M on GENEB MLP-PROBE: Average macro-MCC across the 13 representative tasks, MLP probe
- dataset: LucaOne on GENEB LINEAR-PROBE: Average macro-MCC across the 13 representative tasks, linear probe
- dataset: LucaOne on GENEB MLP-PROBE: Average macro-MCC across the 13 representative tasks, MLP probe
- dataset: MutBERT on GENEB LINEAR-PROBE: Average macro-MCC across the 13 representative tasks, linear probe
- dataset: MutBERT on GENEB MLP-PROBE: Average macro-MCC across the 13 representative tasks, MLP probe
- dataset: NT-v2-50M-MS on GENEB LINEAR-PROBE: Average macro-MCC across the 13 representative tasks, linear probe
- dataset: NT-v2-50M-MS on GENEB MLP-PROBE: Average macro-MCC across the 13 representative tasks, MLP probe
- dataset: Omni-DNA-1B on GENEB LINEAR-PROBE: Average macro-MCC across the 13 representative tasks, linear probe
- dataset: Omni-DNA-1B on GENEB MLP-PROBE: Average macro-MCC across the 13 representative tasks, MLP probe