A pre-season model tested on a mostly novel 2024 hybrid cohort
The primary result is an independently timed 2024 holdout of a frozen corrected GBLUP pipeline. Predictions used only genotype and historical field-trial information; 90.2% of the 1,063 test hybrids had no eligible 2014–2023 phenotype record, making the holdout substantially harder than the earlier 2023 diagnostic.
Can historical genotype and field-trial data rank maize hybrids in a future season?
Multi-environment maize trials are sparse across years: hybrids, testers, locations, and genotyping platforms change, while most hybrid-by-environment combinations are never observed. The breeding question is therefore prospective—whether data available before a new season can rank candidates within each future environment.
Scientific question
Which leakage-controlled genomic approach best predicts within-environment grain-yield rankings, how well does a frozen pre-season GBLUP transfer to the 2024 season, and how much does performance change for truly novel hybrids?
| Organism | Zea mays maize hybrids |
|---|---|
| Source | G2F 2024 Maize Genotype by Environment Prediction Competition Data |
| Historical phenotype | 2014–2023 grain-yield field trials |
| Genotype panel | 5,899 hybrids × 2,425 SNP markers |
| Frozen test | 10,057 requested hybrid × environment rows in 23 2024 environments |
| Primary endpoint | Equal-weight macro within-environment Spearman rank correlation |
| Prediction timing | Pre-season; no within-season environmental covariates |
What was asked of Pipette?
Audit the G2F experimental design; compare genomic, G×E, machine-learning, and eligible sequence-model approaches under expanding-window forward-year validation; lock the selected pre-season pipeline; then predict and evaluate 2024 without tuning after outcome access.
The endpoint is yield ranking within each environment, not prediction of absolute yield across environments. Reported correlations are predictive ability against observed phenotypes; they are not accuracy against unobserved true breeding values.
Forward-year model search, a frozen GBLUP specification, and a pre-outcome prediction lock
Model-family differences were small beside year-to-year variation
Across six forward-year development folds, the leading models clustered near macro Spearman ρ = 0.20. Random forest with marker and environmental-covariate features reached 0.2007; corrected GBLUP reached 0.1986; RR-BLUP and BayesB were similarly close. The practical pre-season choice was corrected GBLUP because it used genotype and historical trial data without growing-season environmental covariates.

The corrected genomic signal exceeded a within-environment permutation null
The initial GBLUP relationship matrix used a non-standard dosage scale. Correcting the scale changed kernel magnitude but preserved rankings: development macro Spearman shifted from 0.1985 to 0.1986. In 100 exact-model within-environment permutations, no null score reached the observed corrected score; the plus-one empirical p value was 0.0099.

The post-season RF-plus-environmental-covariate model had the highest development mean by 0.0021, but it requires growing-season summaries and was not evaluated as the locked pre-season 2024 model. The case study does not collapse those timing scenarios into one claim.
The locked pre-season GBLUP produced a positive but modest 2024 ranking signal
ρ = 0.1426 across 22 evaluable 2024 environments; all 22 environment-level correlations were positive.
r = 0.2239 across the same evaluable environments.
Environment-level Spearman ranged from 0.0124 to 0.2964, with SD 0.0795.
9,486 of 10,057 prediction rows had observed values; SCH1_2024 lacked outcomes for 571 rows.

Novel hybrids drove the performance drop
Only 104 of 1,063 test hybrids had eligible historical records; 959, or 90.2%, were novel. Seen hybrids reached macro Spearman 0.3184, while novel hybrids reached 0.0888. The overall 0.1426 score therefore describes a test cohort dominated by genetically new candidates rather than a repeat-hybrid scenario.

The positive rank correlations support limited predictive transfer, especially for previously observed hybrids. They do not show causal marker effects, stable accuracy in later years, or sufficient reliability to replace multi-environment field testing.
Top-10% selection was near random; top-20% selection was only modestly better
| Metric | Top 10% | Top 20% |
|---|---|---|
| Macro precision / recall | 0.0979 | 0.2230 |
| Random baseline | 0.10 | 0.20 |
| Macro realized gain | +0.0964 Mg/ha | +0.2116 Mg/ha |
| Gain range | −0.3474 to +0.6835 Mg/ha | −0.1916 to +0.6078 Mg/ha |
Selecting the predicted top 10% recovered essentially the random fraction of observed top performers. The mean yield of the predicted top 10% exceeded the environment mean by 0.096 Mg/ha, but gain was positive in only 13 of 22 environments.

DNABERT-2 sequence features did not add supported predictive value
A later analysis embedded 2,047-nt B73 v5 windows centered on 2,383 eligible biallelic SNPs. Across the frozen 2017–2022 development folds, GBLUP reached macro Spearman 0.2026, DNABERT-only reached 0.1629, and a GBLUP-plus-DNABERT multi-kernel model reached 0.2035.

The multi-kernel improvement over GBLUP was practically negligible.
The observed increment was indistinguishable from the exact-model null distribution.
The multi-kernel fit assigned 96.3% of modeled genetic variance to the genomic relationship matrix and 3.7% to the DNABERT kernel.
For novel hybrids, combined-model Spearman was 0.1271 versus 0.1251 for GBLUP.
The finding applies to this sparse 2,425-SNP panel, the 2,047-nt allele-window representation, and the tested kernel and random-forest integrations. It does not establish that sequence foundation models can never help with denser or haplotype-level maize data.
Novel germplasm, sparse markers, and one test season constrain the breeding claim
- The 2024 cohort contains 959 novel hybrids, or 90.2% of test hybrids. Their macro Spearman of 0.0888 is close to zero and substantially below the seen-hybrid result.
- The shared genotype representation contains only 2,425 SNPs; 48 markers exceeded 50% missingness in a 200-hybrid 2024 audit sample and were column-mean imputed under the frozen protocol.
- Genotyping platforms vary by cohort: GBS, WGS, and exome data contributed across years. The shared panel reduces but does not eliminate platform-batch confounding.
- Tester panels changed across year cohorts, and many hybrid-by-environment groups had low replication, adding uncertainty to BLUE targets.
- SCH1_2024 had 571 requested prediction rows but no observed values, so evaluation covers 22 of 23 requested environments.
- Only one 2024 location was novel. The observed 0.0867 correlation at ONH3 cannot support a general claim about transfer to new locations.
- The locked holdout covers one season. Year-to-year development accuracy was highly variable, so another pre-committed future-year validation is needed.
- Selection gain depends on the chosen fraction and environment; top-10% precision was at the random baseline and nine environments had negative realized gain.
- DNABERT-2 conclusions are limited to a sparse marker-derived representation; the exposed 2023 and 2024 outcomes were diagnostic only for that later comparison.
The strongest evidence is the pre-outcome lock and one-time 2024 evaluation of corrected pre-season GBLUP. It supports a modest positive ranking signal overall, useful accuracy for seen hybrids, weak transfer to novel hybrids, and no dependable top-10% selection advantage.
Frozen contracts, row-level predictions, model registries, and evaluation tables
The files below are recorded in the 86-page session inventory. This website publishes the complete session PDF and selected report-derived figures; the row-level scientific artifacts remain documented for provenance.
Frozen model contractCorrected GBLUP specification, eligibility rules, input hashes, and outcome-access policy2024 prediction lock manifestSHA-256, timestamp, schema, and model provenance for the read-only prediction fileLocked 2024 predictions10,057 GEBVs with stable row identifiers and seen-versus-novel labels2024 cohort auditHybrid and location novelty, genotype coverage, missingness, and prediction eligibilityPer-environment evaluationSpearman, Pearson, top-fraction selection, and realized-gain metricsForward-year model registryAttempted model families, fold metrics, hyperparameters, failures, and eligibility decisionsPermutation controlsCorrected GBLUP null scores and DNABERT incremental-value null resultsDNABERT-2 manifestsMarker audit, 2,047-nt windows, embedding provenance, kernel comparisons, and variance components