← Back to Blog

Genomic Selection Explained for Plant Breeding

By Pipette.bio Team

Plant breeders routinely face a simple problem with a difficult solution: there are far more potential lines, crosses, and hybrids than can realistically be tested in the field.

A breeding program may generate thousands of candidates, but evaluating every candidate across multiple locations and years is expensive and slow.

Genomic selection provides another way to prioritize those decisions. It is one application of agricultural genomics: genome-wide marker information and historical phenotype data are used to predict which candidates are most likely to have useful genetic merit. Breeders can then decide which material to advance, cross, or test more extensively.

What is genomic selection?

Genomic selection is a breeding strategy in which markers distributed across the genome are used to predict the genetic merit of breeding candidates.

genotypes + phenotypes from tested material → predictive model → genomic predictions for new material

Unlike marker-assisted selection, genomic selection does not usually focus only on a few markers with large known effects. It attempts to use information spread across many markers simultaneously.

This is especially relevant to complex traits such as:

These traits are often influenced by many loci, each contributing a small part of the genetic variation. The foundational genome-wide prediction study by Meuwissen, Hayes, and Goddard showed how dense markers could be modeled jointly to predict total genetic value.

Genomic prediction vs genomic selection

The terms are often used interchangeably, but they describe different stages.

Genomic prediction is the modeling step: genomic and phenotypic data are used to predict genetic merit or future performance.

Genomic selection occurs when those predictions affect breeding decisions.

If a model predicts yield-related merit for 2,000 untested wheat lines, that is genomic prediction. If a breeder advances 200 of those lines using the predictions as part of the decision, that is genomic selection.

The distinction matters. A model with useful predictive performance does not improve a breeding program unless its predictions are used in a selection strategy that creates value.

How genomic selection works

A genomic-selection program usually contains two important groups.

1. The training population

The training population contains individuals with both genotypes and phenotypes. A maize program, for example, may have 3,000 hybrids already genotyped with SNP markers and evaluated for yield across locations and years.

HybridSNP1SNP2SNP3Yield
H0010129.4
H0021018.7
H00321010.2

Real datasets may contain tens or hundreds of thousands of markers. The model learns how genome-wide similarity or marker effects relate to the target phenotype in the training material.

2. The prediction population

The prediction population contains new breeding candidates with genotypes but little or no phenotype information. Genotyping 5,000 new lines may be feasible even when testing all of them across 20 field environments is not.

The model predicts genetic merit for those candidates. The breeder can then focus expensive field testing on candidates that best satisfy the program’s selection objective.

What is a GEBV?

A genomic estimated breeding value, or GEBV, is a prediction of an individual’s genetic merit based on genomic information and a fitted model.

LinePredicted breeding value
L101+1.42
L205+1.18
L077+0.91
L310−0.22

If higher values are favorable for the modeled target, L101 ranks ahead of L310. The GEBV is not necessarily a prediction of raw plot yield. Its interpretation depends on the trait definition, model, genetic effects included, and scale of the training phenotype.

Genotype is not the same as phenotype

A simplified representation of field performance is:

phenotype = overall mean + genetics + environment + genetics × environment + error

A maize hybrid that performs well in one environment may rank differently under heat or drought elsewhere. Breeding datasets therefore often include location, year, management, weather, soil, and trial-design information alongside genotype and phenotype.

Multi-environment genomic-prediction models can represent genotype-by-environment interaction rather than assuming that every genotype has the same relative performance everywhere.

Which genotype is likely to perform well in the target population of environments?

That is often the more useful breeding question.

Genomic selection vs marker-assisted selection

Marker-assisted selection

Marker-assisted selection usually relies on a relatively small number of markers associated with important genes or QTL. If an allele confers a major disease-resistance effect, seedlings can be genotyped and plants carrying the desired allele retained.

This approach is particularly useful when a trait is strongly influenced by one or a few loci.

Genomic selection

Genomic selection uses markers throughout the genome. Instead of asking whether a plant carries a favorable allele at one QTL, it asks what genetic value is predicted from the entire marker profile.

This makes genomic selection attractive for highly polygenic traits, where no single marker explains enough variation to drive selection alone. The two approaches can also coexist: a program may enforce required major alleles while using a genome-wide prediction to rank candidates for complex traits.

What models are used for genomic selection?

A common baseline is genomic best linear unbiased prediction, or GBLUP, and the closely related ridge-regression approach RR-BLUP. These methods model genomic relationships or many marker effects with shrinkage and remain widely used because they are efficient and often competitive.

Other options include:

The most complex model is not automatically the most useful. Training-population design, phenotype quality, relatedness, trait architecture, validation scenario, and operational constraints often matter more than the difference between two algorithms.

Why the training population matters

A yield model trained in one maize breeding population may generalize poorly to genetically distant material. Prediction tends to be stronger when candidates are related to, or otherwise well represented by, the training population.

Training-set size, population structure, relatedness between training and prediction candidates, phenotype precision, environment coverage, and continued updating all affect performance. Empirical work in wheat and maize has repeatedly shown that training-population composition matters, not just its raw size.

A genomic-prediction model is therefore rarely something to train once and use forever.

historical trials → model → selection → new trials → new phenotypes → updated model

That cycle is how genomic selection can accumulate long-term value.

How do you know whether genomic selection works?

Predictions must be tested on material that was not used to fit the model. A basic cross-validation split might train on 80% of lines and test on the remaining 20%.

The correlation between predictions and observed phenotypes is often reported. If that correlation is 0.55, the model contains predictive signal, but whether it is useful depends on the trait, phenotype reliability, selection objective, and cost of mistakes.

Terminology matters here. Correlation with an observed phenotype is often called predictive ability. It is not automatically the accuracy of a true breeding value, because field phenotypes also contain environmental and residual variation.

Random cross-validation can be misleading

If siblings from the same family appear in both training and test sets, random cross-validation can look excellent even when the operational problem is predicting entirely new families in a future year.

More realistic designs may include:

Research in structured Brassica napus populations has shown how random cross-validation can be inflated by family structure. The validation design should reproduce the decision the breeder intends to make.

Genomic selection across locations and years

Field-trial data are multi-environment by nature. A breeder may need to predict how a tested hybrid will perform in a future year, how an untested genotype will perform in known environments, or how candidates will behave in a new target environment.

Models can incorporate:

These are different prediction problems. A model that predicts untested genotypes in observed environments is not automatically validated for untested genotypes in future, unseen environments.

Genomic selection for hybrid breeding

Hybrid crops create a large combinatorial problem. With 500 female and 500 male lines, there are 250,000 possible pairings. Testing every combination is unrealistic.

250,000 possible crosses → predict hybrid performance → rank candidates → field-test a tractable subset

Historical hybrid trials can link parental genotypes, cross identity, field environments, and hybrid performance. Models can then rank untested combinations, provided the parents and target environments are represented well enough to support the prediction.

Large maize studies have demonstrated genome-based prediction of testcross performance, while also showing that prediction can collapse when training and target populations differ too much. Genomic prediction changes where field-testing resources are spent; it does not make the biological coverage problem disappear.

Does genomic selection replace field trials?

No. Genomic selection depends on high-quality phenotyping. Without reliable field data, the model has no useful target to learn.

The objective is to use phenotyping more strategically:

genotype broadly → predict broadly → phenotype selectively → update the model

Well-designed field trials remain necessary for model training, validation, discovering new biology, monitoring changes in the breeding population, and measuring performance in target environments.

What data do you need?

Genotypes

Markers may come from SNP arrays, genotyping-by-sequencing, whole-genome sequencing, or imputed datasets. Marker QC, allele coding, missingness, imputation, and consistent genome coordinates all matter.

Phenotypes

Examples include yield, flowering date, disease score, plant height, quality traits, and stress-response measurements. The phenotype must match the breeding target.

Reliable identities

Genotype, phenotype, pedigree, seed source, and trial records must refer to the same biological material. Misidentified or inconsistently named lines can quietly destroy model quality.

Trial and environmental information

Location, year, replicate, block, treatment, management, weather, soil, and other design variables allow the model or upstream field-trial analysis to separate sources of variation.

BLUPs and genomic selection

Plant breeders often encounter best linear unbiased predictions, or BLUPs, before genomic prediction. Mixed models can account for locations, years, blocks, replicates, spatial field variation, and genetic effects.

multi-environment field trials → mixed model → adjusted genetic target → genomic-prediction model → predictions for new candidates

Depending on the design, a two-stage analysis may use adjusted means, BLUEs, BLUPs, or deregressed values as the genomic-prediction target. These estimates have different amounts of uncertainty. Where the method supports it, their precision or weights should be carried into the second stage rather than treating every adjusted value as equally reliable and error-free.

Why genomic selection can accelerate breeding

The value of genomic selection does not come only from maximizing predictive correlation. It can also come from selecting earlier, evaluating more candidates, or reallocating field capacity.

Genetic gain per unit time depends on selection intensity, accuracy, available genetic variation, and breeding-cycle length. A somewhat imperfect prediction can still be valuable if it permits useful selection earlier or allows a program to screen many more candidates.

Where genomic selection fails

Too little training data

Hundreds of thousands of markers do not compensate for a very small set of reliably phenotyped individuals.

Poor phenotyping

No model can recover biological signal that was never measured reliably.

Training and prediction populations are too different

A model trained on one genetic population, management system, or environment may not transfer to another.

Data leakage

Relatives, repeated observations, preprocessing information, or future records can accidentally cross the validation boundary.

Ignoring environment

Yield and other agronomic traits may be poorly served by a genotype-only model when the operational target depends strongly on environment.

Optimizing prediction instead of selection

The model with the highest correlation is not necessarily the one that creates the most value for the breeding program.

Does using these predictions improve which material we advance?

That is the operational test.

From genomic prediction to predictive breeding

Genomic selection increasingly sits inside a larger predictive-breeding system that combines:

The goal is larger than calculating GEBVs. A mature decision system may help a program determine which lines to advance, which crosses to make, which candidates to test in a target environment, and where an additional field trial would be most informative.

Genomic prediction produces estimates. Genomic selection turns those estimates into decisions. The value comes from the complete loop: representative training data, realistic validation, disciplined selection, new phenotyping, and continuous model updating.

Sources and further reading

  1. Meuwissen THE, Hayes BJ, Goddard ME. Prediction of total genetic value using genome-wide dense marker maps. Genetics (2001).
  2. Crossa J et al. Genomic selection in plant breeding: methods, models, and perspectives. Trends in Plant Science (2017).
  3. Endelman JB. Ridge regression and other kernels for genomic selection with R package rrBLUP. The Plant Genome (2011).
  4. Werner CR et al. How population structure impacts genomic selection accuracy in cross-validation. Frontiers in Plant Science (2020).
  5. Malosetti M et al. Predicting responses in multiple environments: issues in relation to genotype × environment interactions. Crop Science (2016).
  6. Albrecht T et al. Genome-based prediction of testcross values in maize. Theoretical and Applied Genetics (2011).

Genomic selection in Pipette

A genomic-selection workflow involves more than fitting a prediction model. Genotypes must be cleaned, phenotype and trial records reconciled, targets defined, validation designed around the future breeding decision, models compared, and predictions delivered in a form breeders can inspect.

Pipette keeps those datasets, assumptions, validation boundaries, models, and predictions together from preparation through candidate ranking.

The objective is not another cross-validation score. It is evidence that historical breeding data can improve the next decision about what to test, cross, or advance.