Back to Case Studies

Case study · Plant ChIP-seq

Soybean ChIP-seq Analysis: FASTQ to Peaks and Targets

This case study analyzes eight public soybean ChIP-seq and matched-input libraries to characterize genome-wide binding by GLYMA.08G357600 in mid-maturation embryos. The workflow combines read and enrichment QC, replicate-aware peak calling, IDR, peak annotation, motif discovery, and GO enrichment.

Glycine max8 librariesMid-maturation embryoGSE101648MACS3 + IDR

Original 31-page session report · PDF · 2.4 MB

At a glance

Replicate-aware soybean ChIP-seq from raw reads to candidate target processes

The analysis starts with four GLYMA.08G357600 ChIP libraries and four matched inputs from two biological replicates. It reports a broader consensus set for sensitive downstream analysis and a separate IDR set for the most reproducible cross-replicate sites.

Libraries4 ChIP + 4 Input
Design2 biological × 2 technical
Consensus peaks21,067
IDR ≤ 0.057,693 peaks
01 / Scientific context

Where does GLYMA.08G357600 bind in the soybean embryo genome?

GLYMA.08G357600 is a soybean transcription-factor target profiled in mid-maturation embryo tissue. ChIP-seq can locate enriched genomic regions associated with the immunoprecipitated factor, but a peak alone does not prove direct DNA contact or downstream transcriptional regulation.

Scientific question

Which GLYMA.08G357600 binding sites are reproducible across biological replicates, where are those peaks located relative to genes and promoters, which sequence motifs are enriched, and which biological processes are associated with nearby candidate targets?
OrganismGlycine max Williams 82 soybean
Public datasetGEO GSE101648 / SRA SRP112906
TissueMid-maturation embryo
ChIP targetGLYMA.08G357600
LibrariesSRR5849963–SRR5849966 ChIP; SRR5849967–SRR5849970 matched Input
ReplicationTwo biological replicates with two technical replicates each
SequencingIllumina HiSeq 2500, 51-bp single-end
Peak-call referenceNCBI RefSeq GCF_000004515.5 / Glycine_max_v2.1

What was asked of Pipette?

Analyze SRP112906/GSE101648 using matched inputs and biological replicates; assess ChIP-seq QC and concordance; call reproducible GLYMA.08G357600 peaks; annotate their genomic distribution; perform motif enrichment; identify candidate target genes; and interpret associated biological processes.
Reference and annotation warning

The request named Wm82.gnm2.DTC4 and the Wm82.gnm2.ann1.RVB6 gene models, but peak calling and ChIPseeker annotation used NCBI RefSeq GCF_000004515.5 with LOC-style IDs. A post-hoc Ensembl Plants mapping recovered GLYMA IDs for 72.6% of annotated peaks; this only partially resolves the mismatch.

02 / Analysis workflow

FASTQ alignment, enrichment QC, replicate peak calling, annotation, motifs, and GO

Technical replicates were retained for library-level QC and merged within biological replicates for peak calling. ChIP and Input were processed in parallel before replicate-aware signal and peak comparisons.

Align and audit librariesBowtie2 aligned reads; SAMtools filtered at MAPQ ≥ 10 and sorted alignments; Picard marked duplicates without removing them.
Measure ChIP enrichmentphantompeakqualtools estimated NSC, RSC, and fragment length; deepTools produced coverage correlation, PCA, fingerprints, and signal tracks.
Call replicate peaksMACS3 called each biological replicate against matched Input and also called a pooled peak set at q < 0.05.
Define two reproducibility setsbedtools produced the broader pooled-intersection consensus set. IDR 2.0.4.2 was patched for the removed NumPy numpy.int alias, then produced a stricter cross-replicate subset at IDR ≤ 0.05.
Annotate and map genesChIPseeker used a ±2-kb TSS promoter window. Ensembl Plants BioMart mapped NCBI LOC identifiers to GLYMA identifiers where possible.
Test sequence and process enrichmentSTREME discovered motifs in top-peak summit windows, TOMTOM compared motifs with JASPAR2018, FIMO scanned all consensus summits, and clusterProfiler tested GO Biological Process terms.
Peak-set boundaryThe 21,067-peak consensus set is the report’s primary downstream set for sensitivity. The 7,693 IDR peaks are the stricter reproducible subset and should be prioritized for high-confidence validation. Functional analyses were not rerun exclusively on the IDR subset.
03 / Quality and reproducibility

All libraries passed the recorded QC gates and ChIP replicates were highly concordant

Alignment

ChIP libraries aligned at 82.5–85.9%; matched inputs aligned at 98.8–98.9%.

Duplication

ChIP duplication was 5.0–6.8% and Input duplication was 1.8–2.4%, below the report’s 10% warning threshold.

Cross-correlation

NSC ranged from 1.19–1.25 and RSC from 2.23–2.34; every ChIP library received Quality Tag 2/2.

Replicate concordance

ChIP–ChIP Pearson correlations ranged from 0.988–0.992; the IDR rank-correlation parameter was ρ = 0.92.

Alignment and PCR duplication rates for four soybean GLYMA.08G357600 ChIP and four matched Input libraries
Figure 1. Library alignment and duplication metrics. All eight libraries exceed the recorded 80% alignment threshold and remain below the 10% duplication warning threshold.

ChIP and Input clustered as distinct sample classes

The four ChIP libraries form one correlation cluster and the four matched Input libraries form another. High within-condition correlation supports consistent genome-wide coverage, while condition separation supports an enrichment signal rather than interchangeable ChIP and background profiles.

Pearson correlation heatmap clustering four soybean ChIP libraries separately from four matched Input libraries
Figure 2. Genome-wide replicate correlation. The heatmap uses 10-kb coverage bins. ChIP–ChIP correlations were 0.988–0.992 and Input–Input correlations were 0.992–0.995.

The broader consensus and stricter IDR results answer different questions

Peak setCountRole in report
Biological replicate 127,627MACS3 q < 0.05
Biological replicate 224,224MACS3 q < 0.05
Pooled37,623Pooled-library MACS3 result
bedtools reproducible regions19,513BR1/BR2 intersection-derived set
Consensus final21,067Broader primary set used downstream
IDR passing7,693Stricter set at IDR ≤ 0.05

IDR tested 19,515 union peaks and retained 7,693, or 39.4%. The report’s 21,067 consensus peaks come from a different bedtools-and-pooled construction, so 7,693 should not be presented as 39.4% of the consensus set.

04 / Peak annotation

Fifty-nine percent of annotated consensus peaks were promoter-proximal

ChIPseeker assigned genomic features to 21,060 of the 21,067 consensus peaks. Using a ±2-kb promoter window, 12,421 annotated peaks were promoter-proximal: 9,543 within 1 kb of a TSS and 2,878 between 1 and 2 kb.

59.0% promoter

12,421 of 21,060 annotated peaks fell within ±2 kb of a TSS.

34.6% distal

7,285 annotated peaks were classified intergenic or distal.

10,446 target IDs

The post-hoc mapping produced 10,446 unique GLYMA identifiers associated with any peak.

7,445 promoter targets

Promoter-proximal peaks mapped to 7,445 unique GLYMA identifiers before GO coverage filtering.

Bar chart showing promoter, distal, downstream, exonic, and intronic distribution of 21,060 annotated soybean ChIP-seq peaks
Figure 3. Genomic distribution of annotated consensus peaks. Promoter categories total 59.0%; this distribution describes the broader consensus set, not the 7,693-peak IDR subset.
Annotation completeness

15,298 of 21,060 annotated peaks (72.6%) received a GLYMA identifier. The remaining 5,762 peaks lack a direct Entrez-to-GLYMA cross-reference or fall on unplaced scaffolds and are excluded from GLYMA-based enrichment.

05 / Motifs and biological interpretation

A G-box motif and seed-maturation processes emerged from the broader consensus set

STREME reported 16 significant de novo motifs from ±100-bp windows around the top 5,000 peak summits. TOMTOM matched nine motifs to entries in the available JASPAR2018 database, and FIMO then scanned summit sequences from all 21,067 consensus peaks.

FIMO occurrence counts for 16 de novo motifs found in GLYMA.08G357600 ChIP-seq peak summits
Figure 4. De novo motif occurrence counts. Motif 1 contains the CACGTG G-box core, produced 11,310 FIMO hits at p < 10⁻⁴, and matched the HY5 entry at TOMTOM q = 4.4 × 10⁻⁷. Match names are indicative because the preferred plant-specific database was unavailable.

Revised GLYMA-based enrichment returned 47 promoter-target Biological Process terms

The revised background contained GO annotations for 42,883 of 57,147 GLYMA identifiers, or 75.0%. Of 7,445 promoter-target GLYMA identifiers, 4,726 (63.5%) mapped into that background. The top term was protein dephosphorylation (62 genes; adjusted p = 2.49 × 10⁻⁶), with additional terms including circadian rhythm, lipid storage, photosynthesis, ethylene signaling, and gibberellic-acid signaling.

GO Biological Process enrichment dot plot for GLYMA.08G357600 promoter-peak target genes using GLYMA identifiers
Figure 5. Revised GO Biological Process enrichment. The plot shows the top 20 of 47 significant promoter-target terms. Bubble size encodes gene count and color encodes adjusted-p-value strength.
Interpretation boundary

ChIP enrichment supports chromatin association; motif enrichment supports a sequence pattern; and GO over-representation supports a target-set association. None alone proves that GLYMA.08G357600 directly binds the G-box, regulates every nearby gene, or controls lipid storage, ABA signaling, or circadian biology. Those are hypotheses requiring ChIP-qPCR, DNA-binding assays, and perturbation transcriptomics.

06 / Limitations

Reference mapping, motif resources, and experimental scope limit the claims

  • Peak calling and initial annotation used NCBI RefSeq GCF_000004515.5 rather than the requested Wm82.gnm2.DTC4 and RVB6 gene-model resources.
  • Only 72.6% of the 21,060 annotated peaks received GLYMA identifiers; 5,762 peaks are absent from GLYMA-based gene lists and GO analysis.
  • The preferred JASPAR2022 plant database could not be downloaded. TOMTOM used JASPAR2018, which contains some plant entries but is predominantly vertebrate-derived.
  • Only 4,726 of 7,445 promoter-target GLYMA identifiers mapped into the revised GO background, leaving 36.5% outside the foreground analysis.
  • Functional annotation includes electronically inferred terms and can vary in specificity and completeness.
  • The study contains two biological replicates, one tissue stage, and one ChIP target; weak sites, other developmental contexts, and co-binding partners remain unresolved.
  • Downstream annotation, motif scanning, and GO enrichment used the 21,067 consensus set rather than the 7,693-peak IDR subset.
  • Nearby-gene assignment does not establish direct transcriptional regulation, and ChIP-seq does not establish whether binding is direct or mediated through a protein complex.
Evidence status

The run provides strong library-level QC and reproducible ChIP enrichment, a broader consensus peak set, and a stricter IDR subset. Confidence is lower for complete GLYMA target assignment, plant-specific motif identity, and causal biological mechanisms.

07 / Reproducibility and outputs

Peak sets, annotations, enrichment tables, signal tracks, and machine-readable flags

The files below are recorded in the session inventory. The website publishes the complete session PDF and selected report figures; the remaining artifacts are documented for provenance.

peaks/peaks.filtered.bed21,067-peak consensus set used for downstream analysis
peaks/idr_output.txt7,693 peaks passing IDR ≤ 0.05
peaks/annotated_peaks_glyma.csvPeak annotations with post-hoc GLYMA identifiers
results/GO_BP_enrichment_glyma_promoter.csv47 significant promoter-target GO Biological Process terms
results/pearson_correlation_matrix.txtEight-library coverage correlation matrix
results/qc_metrics.jsonMachine-readable alignment, duplication, and enrichment QC
results/analysis_flags.jsonUpdated active and resolved methodology warnings
motif_results/streme_out/streme.txtSixteen de novo motifs in MEME format
motif_results/tomtom_out/tomtom.tsvJASPAR2018 motif comparisons
motif_results/fimo_out/fimo.tsvPer-site motif occurrences across consensus summits
results/chip_vs_input_pooled.bwPooled log₂ ChIP/Input genome-browser signal track
Original Pipette session report31-page PDF · 2.4 MB · generated August 29, 2026. The HTML article is the primary case study; the PDF is the preserved session record.

Run a replicate-aware ChIP-seq analysis

Bring raw reads, matched controls, biological replicate structure, and the reference resources that should govern peak and gene annotation.

Start in Pipette