Back to Case Studies

Case study · Plant breeding and genomic prediction

Maize Genomic Prediction: Future-Season Yield Ranking

This case study uses historical Genomes to Fields genotype and field-trial data to predict maize hybrid grain-yield rankings in a future season. A corrected pre-season GBLUP pipeline was frozen on 2014–2023 data, and 10,057 predictions were hashed and locked before 2024 observed yields were opened.

Zea maysG2F 2014–20242,425 SNPs10,057 locked predictionsGBLUP

Original 86-page session report · PDF · 5.2 MB

At a glance

A pre-season model tested on a mostly novel 2024 hybrid cohort

The primary result is an independently timed 2024 holdout of a frozen corrected GBLUP pipeline. Predictions used only genotype and historical field-trial information; 90.2% of the 1,063 test hybrids had no eligible 2014–2023 phenotype record, making the holdout substantially harder than the earlier 2023 diagnostic.

2024 predictions10,057 rows
Novel hybrids959 of 1,063
Macro Spearmanρ = 0.143
Prediction lockBefore outcome access
01 / Scientific context

Can historical genotype and field-trial data rank maize hybrids in a future season?

Multi-environment maize trials are sparse across years: hybrids, testers, locations, and genotyping platforms change, while most hybrid-by-environment combinations are never observed. The breeding question is therefore prospective—whether data available before a new season can rank candidates within each future environment.

Scientific question

Which leakage-controlled genomic approach best predicts within-environment grain-yield rankings, how well does a frozen pre-season GBLUP transfer to the 2024 season, and how much does performance change for truly novel hybrids?
OrganismZea mays maize hybrids
SourceG2F 2024 Maize Genotype by Environment Prediction Competition Data
Historical phenotype2014–2023 grain-yield field trials
Genotype panel5,899 hybrids × 2,425 SNP markers
Frozen test10,057 requested hybrid × environment rows in 23 2024 environments
Primary endpointEqual-weight macro within-environment Spearman rank correlation
Prediction timingPre-season; no within-season environmental covariates

What was asked of Pipette?

Audit the G2F experimental design; compare genomic, G×E, machine-learning, and eligible sequence-model approaches under expanding-window forward-year validation; lock the selected pre-season pipeline; then predict and evaluate 2024 without tuning after outcome access.
Prediction target

The endpoint is yield ranking within each environment, not prediction of absolute yield across environments. Reported correlations are predictive ability against observed phenotypes; they are not accuracy against unobserved true breeding values.

02 / Analysis workflow

Forward-year model search, a frozen GBLUP specification, and a pre-outcome prediction lock

Audit the trial systemThe session reconciled phenotypes, hybrid IDs, 2,425-marker genotypes, environment-years, weather, soil, management, replication, missingness, and platform changes.
Construct within-environment targetsHybrid BLUEs were estimated with a replicate random effect when at least two replicates were available; plot means were used otherwise. Targets were centered within environment.
Search on forward yearsSix expanding-window development folds held out 2017 through 2022 in turn. Baselines, RR-BLUP, GBLUP, BayesB, random forest, boosted trees, ElasticNet, G×E models, and ensembles were compared.
Correct and freeze GBLUPThe uploaded {0, 0.5, 1} dosage was transformed to {−1, 0, 1} before rrBLUP::A.mat. The corrected pre-season GBLUP specification was refit on all eligible 2014–2023 records without 2024 outcomes.
Generate and lock predictionsAll 10,057 requested 2024 GEBVs were written, made read-only, and recorded in a manifest with input provenance and SHA-256 6a48cdeba5443a60c67c4c13689c1827ccd8a3640c89df7082d2d65cb4f619f8.
Open outcomes onceOnly after the prediction hash was verified were 2024 observed yields extracted and joined for per-environment ranking, selection, novelty, and realized-gain evaluation.
Leakage boundaryNo 2024 outcome was used to tune, reselect, reweight, or modify the model. Later DNABERT-2 comparisons treat both 2023 and 2024 as exposed diagnostic years; only the 2017–2022 development folds support model-comparison conclusions.
04 / 2024 temporal holdout

The locked pre-season GBLUP produced a positive but modest 2024 ranking signal

Macro Spearman

ρ = 0.1426 across 22 evaluable 2024 environments; all 22 environment-level correlations were positive.

Macro Pearson

r = 0.2239 across the same evaluable environments.

Accuracy range

Environment-level Spearman ranged from 0.0124 to 0.2964, with SD 0.0795.

Outcome attrition

9,486 of 10,057 prediction rows had observed values; SCH1_2024 lacked outcomes for 571 rows.

Bar chart of corrected GBLUP yield-ranking accuracy across 22 evaluable 2024 G2F environments
Figure 3. Per-environment 2024 ranking accuracy. Within-environment Spearman was positive in every evaluable environment but varied from 0.012 to 0.296; the macro mean was 0.143. The single novel location, ONH3, is highlighted.

Novel hybrids drove the performance drop

Only 104 of 1,063 test hybrids had eligible historical records; 959, or 90.2%, were novel. Seen hybrids reached macro Spearman 0.3184, while novel hybrids reached 0.0888. The overall 0.1426 score therefore describes a test cohort dominated by genetically new candidates rather than a repeat-hybrid scenario.

Boxplots showing higher within-environment Spearman correlations for seen maize hybrids than for novel hybrids in 2024
Figure 4. Prediction accuracy by hybrid novelty. Seen-hybrid macro Spearman was 0.318; novel-hybrid macro Spearman was 0.089. This descriptive gap is central to the validated prediction scenario.
Interpretation boundary

The positive rank correlations support limited predictive transfer, especially for previously observed hybrids. They do not show causal marker effects, stable accuracy in later years, or sufficient reliability to replace multi-environment field testing.

05 / Breeding-selection utility

Top-10% selection was near random; top-20% selection was only modestly better

MetricTop 10%Top 20%
Macro precision / recall0.09790.2230
Random baseline0.100.20
Macro realized gain+0.0964 Mg/ha+0.2116 Mg/ha
Gain range−0.3474 to +0.6835 Mg/ha−0.1916 to +0.6078 Mg/ha

Selecting the predicted top 10% recovered essentially the random fraction of observed top performers. The mean yield of the predicted top 10% exceeded the environment mean by 0.096 Mg/ha, but gain was positive in only 13 of 22 environments.

Per-environment realized maize yield gain from selecting the top 10 percent by genomic estimated breeding value
Figure 5. Realized yield gain from top-10% GEBV selection. Gains were positive in 13 environments and negative in nine. The macro gain of 0.096 Mg/ha is exploratory and highly environment-dependent.
Operational limitThe top-20% gain is more encouraging than the top-10% result, but a single holdout year with broad between-environment variation is not enough to claim dependable selection utility. Replication in another locked season is required.
06 / Foundation-model comparison

DNABERT-2 sequence features did not add supported predictive value

A later analysis embedded 2,047-nt B73 v5 windows centered on 2,383 eligible biallelic SNPs. Across the frozen 2017–2022 development folds, GBLUP reached macro Spearman 0.2026, DNABERT-only reached 0.1629, and a GBLUP-plus-DNABERT multi-kernel model reached 0.2035.

Forward-year development Spearman correlations for GBLUP, DNABERT-only, and combined GBLUP plus DNABERT maize models
Figure 6. DNABERT-2 comparison across forward-year folds. DNABERT-only underperformed GBLUP, while the combined model differed from GBLUP by only +0.0009 macro Spearman.
+0.0009 Spearman

The multi-kernel improvement over GBLUP was practically negligible.

Permutation p = 0.4217

The observed increment was indistinguishable from the exact-model null distribution.

3.7% sequence variance

The multi-kernel fit assigned 96.3% of modeled genetic variance to the genomic relationship matrix and 3.7% to the DNABERT kernel.

No novel-hybrid rescue

For novel hybrids, combined-model Spearman was 0.1271 versus 0.1251 for GBLUP.

Scope of the null result

The finding applies to this sparse 2,425-SNP panel, the 2,047-nt allele-window representation, and the tested kernel and random-forest integrations. It does not establish that sequence foundation models can never help with denser or haplotype-level maize data.

07 / Limitations

Novel germplasm, sparse markers, and one test season constrain the breeding claim

  • The 2024 cohort contains 959 novel hybrids, or 90.2% of test hybrids. Their macro Spearman of 0.0888 is close to zero and substantially below the seen-hybrid result.
  • The shared genotype representation contains only 2,425 SNPs; 48 markers exceeded 50% missingness in a 200-hybrid 2024 audit sample and were column-mean imputed under the frozen protocol.
  • Genotyping platforms vary by cohort: GBS, WGS, and exome data contributed across years. The shared panel reduces but does not eliminate platform-batch confounding.
  • Tester panels changed across year cohorts, and many hybrid-by-environment groups had low replication, adding uncertainty to BLUE targets.
  • SCH1_2024 had 571 requested prediction rows but no observed values, so evaluation covers 22 of 23 requested environments.
  • Only one 2024 location was novel. The observed 0.0867 correlation at ONH3 cannot support a general claim about transfer to new locations.
  • The locked holdout covers one season. Year-to-year development accuracy was highly variable, so another pre-committed future-year validation is needed.
  • Selection gain depends on the chosen fraction and environment; top-10% precision was at the random baseline and nine environments had negative realized gain.
  • DNABERT-2 conclusions are limited to a sparse marker-derived representation; the exposed 2023 and 2024 outcomes were diagnostic only for that later comparison.
Evidence status

The strongest evidence is the pre-outcome lock and one-time 2024 evaluation of corrected pre-season GBLUP. It supports a modest positive ranking signal overall, useful accuracy for seen hybrids, weak transfer to novel hybrids, and no dependable top-10% selection advantage.

08 / Reproducibility and outputs

Frozen contracts, row-level predictions, model registries, and evaluation tables

The files below are recorded in the 86-page session inventory. This website publishes the complete session PDF and selected report-derived figures; the row-level scientific artifacts remain documented for provenance.

Frozen model contractCorrected GBLUP specification, eligibility rules, input hashes, and outcome-access policy
2024 prediction lock manifestSHA-256, timestamp, schema, and model provenance for the read-only prediction file
Locked 2024 predictions10,057 GEBVs with stable row identifiers and seen-versus-novel labels
2024 cohort auditHybrid and location novelty, genotype coverage, missingness, and prediction eligibility
Per-environment evaluationSpearman, Pearson, top-fraction selection, and realized-gain metrics
Forward-year model registryAttempted model families, fold metrics, hyperparameters, failures, and eligibility decisions
Permutation controlsCorrected GBLUP null scores and DNABERT incremental-value null results
DNABERT-2 manifestsMarker audit, 2,047-nt windows, embedding provenance, kernel comparisons, and variance components
Original Pipette session report86-page PDF · 5.2 MB · generated September 24, 2026. The HTML article is the primary case study; the PDF is the preserved session record.

Run an auditable genomic-prediction study

Bring genotype, phenotype, environment, and pedigree data, then define the future breeding decision and prediction-time information boundary.

Start in Pipette