Back to Case Studies

Case study · Plant breeding and environmental prediction

Soybean Genomic Prediction: When Environmental Covariates Help

This SoyNAM case study compares genomic, genotype-by-environment, and genotype-by-environmental-covariate models for soybean grain yield. Environmental covariates improved prediction for untested or partially observed genotypes in known environments, but degraded prediction when the environment itself was withheld.

Glycine max1,379 genotypes4 environments4,420 SNPsDryad dataset

Original 16-page session report · PDF · 1.2 MB

At a glance

Environmental data helped inside known environments but failed to transfer to a withheld environment

Filtered daily environmental covariates produced the highest macro Pearson predictive ability in CV1 and CV2. For CV0, where each environment was withheld in turn, genomic-only prediction was best and every environmental-covariate model reduced accuracy.

SoyNAM lines1,379 genotypes
Observed environments4 in 2012
Best CV1 / CV2WFILT · 0.539 / 0.550
Best CV0Genomic only · 0.333
01 / Scientific context

When does environmental information improve soybean genomic prediction?

Soybean grain yield depends on both genotype and environment. Environmental kernels can help models share information among trials with similar weather patterns, but the value of that information depends on whether the target environment is already represented in training.

Scientific question

Do daily, filtered, stage-summarized, or season-average environmental covariates improve soybean yield prediction beyond genomic-only and genotype-by-environment models—and do those gains persist when predicting a novel environment?
OrganismGlycine max soybean
Public datasetDryad DOI 10.5061/dryad.2fqz6133v
PhenotypeGrain yield in kg/ha; 5,516 observations
Genotypes1,379 SoyNAM lines
EnvironmentsIA, IL, IN, and NE in 2012
Markers4,611 SNPs supplied; 4,420 retained after QC
Primary endpointMacro within-environment Pearson predictive ability

Three breeding scenarios

ScenarioHeld outDecision represented
CV1GenotypesUntested genotypes in observed environments
CV2Observations within genotypeTested genotypes with observations withheld
CV0EnvironmentTested genotypes in a novel environment
Environmental-covariate boundary

Environmental covariates are predictive inputs here, not causal effects. CV0 uses weather data for the withheld environment to project its environmental kernel; it tests whether an EC-to-phenotype relationship transfers, not whether weather data are available.

02 / Analysis workflow

Matched validation splits for genomic, G×E, environmental-kernel, and consensus models

Reconcile the SoyNAM filesGenotypes, yield observations, environment IDs, 4,611 markers, and four environmental-covariate representations were joined and audited.
Prepare genomic inputsSNPs were quality controlled to 4,420 markers. The remaining 4.87% marker missingness was mean-imputed before genomic-kernel construction.
Build four EC representationsModels used 1,306 daily covariates (W), seven season averages (WCONV), 1,127 R²-filtered daily covariates (WFILT), or 70 stage-summarized covariates (WSTG).
Fit matched model familiesBayesian RKHS models compared genomic-only, marker-based G×E, and EC-plus-genotype-by-EC kernels. RR-BLUP and equal-weight consensus variants extended the comparison.
Freeze scenario-specific splitsCV1 and CV2 used five folds with fixed seeds; CV0 used four leave-one-environment-out folds. Comparable models were scored on identical held-out rows.
Measure prediction and selectionThe primary metric was macro within-environment Pearson r. Per-environment r and recovery of the observed top 10% of genotypes were secondary endpoints.
Model-fitting precisionBGLR used 4,000 iterations with 1,000 burn-in iterations and eigendecomposition truncated at 99.9% variance. The published analysis used longer chains and ten CV repetitions, so this case study emphasizes direction and ranking over exact numerical reproduction.
03 / Observed-environment prediction

Filtered daily environmental covariates improved CV1 and CV2 by about 23%

CV1: 0.539

WFILT improved macro Pearson by +0.102, or 23.3%, over the genomic-only baseline of 0.437.

CV2: 0.550

WFILT improved macro Pearson by +0.101, or 22.6%, over the genomic-only baseline of 0.448.

Small gain over G×E

WFILT exceeded marker-based G×E by only +0.007 in CV1 and +0.008 in CV2.

Daily detail mattered

Season-average WCONV reached 0.487 in both scenarios, below daily, filtered, and stage-summarized representations.

Macro within-environment Pearson accuracy for soybean genomic and environmental models under CV1, CV2, and CV0
Figure 1. Predictive ability across three breeding scenarios. Filtered daily ECs led CV1 and CV2, while genomic-only and G×E models led CV0. CV0 error bars summarize only four leave-one-environment-out folds and are not confidence intervals.

Performance varied across the four observed environments

Heatmap of per-environment Pearson predictive ability for soybean genomic and environmental models under CV1, CV2, and CV0
Figure 2. Per-environment predictive ability. Environmental models improved all four environments in CV1 and CV2. In CV0, several environmental representations became negatively correlated in IN or NE, while genomic-only and G×E remained positive.

Because EC kernels contain only four unique environment patterns and have rank no greater than three, much of their CV1/CV2 gain may reflect improved environment-level similarity rather than a stable genotype-by-weather interaction. The narrow +0.007 to +0.008 margin over G×E supports that cautious reading.

Change in soybean prediction accuracy relative to the genomic-only baseline for CV1, CV2, and CV0
Figure 3. Change from the genomic-only baseline. EC models added as much as +0.102 Pearson in observed environments, but every EC representation reduced CV0 accuracy.
04 / Novel-environment prediction

Environmental covariates harmed prediction when the environment was withheld

In CV0, genomic-only prediction reached macro Pearson 0.333 and marker-based G×E reached 0.327. WFILT fell to 0.161, daily W to 0.157, season-average WCONV to 0.061, and stage-summarized WSTG to 0.002.

Per-environment soybean prediction accuracy for genomic and environmental models under leave-one-environment-out CV0
Figure 4. CV0 prediction by held-out environment. Genomic-only and G×E models remained positive in all four environments. Several EC models became negatively correlated in IN or NE, showing harmful extrapolation.
G: 0.333

The genomic-only model was the best CV0 model.

WFILT: −51.7%

Filtered daily EC accuracy was 0.172 Pearson lower than genomic-only prediction.

WSTG: 0.002

Stage-summarized environmental prediction was nearly uninformative on average.

Four folds only

Each environment contributes one CV0 fold; the reported SD of 0.142 for genomic-only prediction is not a confidence interval.

Why transfer failed

With one environment withheld, the EC kernel is calibrated on only three environments and has rank no greater than two. The observed failures show that the fitted environmental similarity did not transfer reliably; a larger multi-year environment panel is needed before inferring a general mechanism.

05 / Breeding-selection utility

Environmental models recovered more top lines in known environments but fewer in CV0

ModelCV1 top 10%CV2 top 10%CV0 top 10%
Random10.0%10.0%10.0%
Genomic only25.4%27.4%23.2%
G×E25.9%28.9%22.1%
WFILT29.3%30.1%17.2%
WSTG27.4%29.5%12.0%
Top-10-percent soybean genotype selection accuracy for genomic and environmental models under CV1, CV2, and CV0
Figure 5. Recovery of the observed top 10% of genotypes. WFILT recovered 29–30% in CV1/CV2, while genomic-only prediction led CV0 at 23.2%. The dashed line marks random 10% recovery.
Breeding recommendation from this datasetWFILT or daily W is the strongest choice for lines evaluated in represented target environments. Genomic-only or G×E models are more defensible for a new environment. This recommendation is provisional because the evidence contains only four environments from one year.
06 / Published-study comparison

The direction and model ranking agree with Sagae et al. (2026)

The source study used the same 1,379 genotypes, four environments, and 4,611 supplied SNPs. It reported that environmental covariates improved CV1/CV2, season-average summarization performed worst among EC representations, and genomic or G×E models were strongest in CV0.

ComparisonSagae et al. (2026)This analysis
Genomic baseline, CV1/CV2Approximately 0.420.437 / 0.448
Best EC, CV1/CV2Approximately 0.560.539 / 0.550
Best CV0 modelG×E or genomic, approximately 0.33Genomic only, 0.333
EC rankingAVG < STG < FILT ≈ ALLWCONV < WSTG < W ≈ WFILT
ECs hurt CV0YesYes
Reproduction boundary

The agreement is directional rather than an exact reproduction. This run used one CV repetition, shorter MCMC chains, truncated eigendecomposition, and raw phenotypes; the source study used ten repetitions, longer chains, and its published phenotype-processing pipeline.

07 / Limitations

Four environments cannot establish a transferable weather-response model

  • The analysis contains only four environments from one year. Environmental kernels have rank no greater than three and behave more like contrasts among known trials than continuous weather-response representations.
  • CV0 has four folds—one per environment. Fold-level SD is unstable and must not be interpreted as a confidence interval.
  • CV1 and CV2 used one cross-validation repetition rather than the source study’s ten, limiting precision of absolute performance estimates.
  • BGLR used 4,000 iterations with 1,000 burn-in iterations rather than the published 12,000 and 2,000 settings.
  • Raw yield observations were modeled; the source pipeline may have used environment-adjusted BLUEs.
  • Stage-summarized ECs had 5% missingness and were mean-imputed, potentially contributing to WSTG’s weak CV0 performance.
  • Marker missingness was 4.87% and was mean-imputed before kernel construction.
  • The equal-weight consensus was assembled post hoc without nested training-fold weight selection and did not beat the best individual model in any scenario.
  • CV0 used weather data from the withheld environment. The failure concerns model transfer across environments rather than predictor availability.
  • Between-environment yield correlations ranged from 0.09 to 0.36, indicating substantial environment specificity and a strong need for broader multi-year validation.
Evidence status

The evidence strongly supports the scenario-dependent direction within this four-environment SoyNAM subset: ECs help when target environments are represented and hurt when extrapolating to the fourth environment. It does not establish that the same EC representation will generalize across years, regions, or larger trial networks.

08 / Reproducibility and outputs

Predictions, split registry, model registry, metrics, and literature comparison

The files below are recorded in the session inventory. This website publishes the complete session PDF and all five report-derived figures; the row-level artifacts remain documented for provenance.

metrics_summary.csvMacro and per-environment Pearson metrics for nine model variants across three scenarios
selection_accuracy.csvTop-10% genotype-recovery accuracy by model and scenario
comparison_table.csvCombined baseline, EC, G×E, and consensus metrics
evidence/predictions.csv143,416 row-level prediction records
evidence/registry.jsonComplete 26-entry model registry
splits/split_registry.jsonFrozen fold assignments and split hash
evidence/contract.jsonFrozen prediction and leakage-control contract
literature_comparison.jsonStructured comparison with Sagae et al. (2026)
synthesis.jsonScenario-specific breeding-utility synthesis
Original Pipette session report16-page PDF · 1.2 MB · generated September 24, 2026. The HTML article is the primary case study; the PDF is the preserved session record.

Test environmental covariates in your breeding data

Bring genotype, phenotype, environment, and weather data, then define whether the target decision concerns known trials, new genotypes, or genuinely novel environments.

Start in Pipette