Replicate-aware soybean ChIP-seq from raw reads to candidate target processes
The analysis starts with four GLYMA.08G357600 ChIP libraries and four matched inputs from two biological replicates. It reports a broader consensus set for sensitive downstream analysis and a separate IDR set for the most reproducible cross-replicate sites.
Where does GLYMA.08G357600 bind in the soybean embryo genome?
GLYMA.08G357600 is a soybean transcription-factor target profiled in mid-maturation embryo tissue. ChIP-seq can locate enriched genomic regions associated with the immunoprecipitated factor, but a peak alone does not prove direct DNA contact or downstream transcriptional regulation.
Scientific question
Which GLYMA.08G357600 binding sites are reproducible across biological replicates, where are those peaks located relative to genes and promoters, which sequence motifs are enriched, and which biological processes are associated with nearby candidate targets?
| Organism | Glycine max Williams 82 soybean |
|---|---|
| Public dataset | GEO GSE101648 / SRA SRP112906 |
| Tissue | Mid-maturation embryo |
| ChIP target | GLYMA.08G357600 |
| Libraries | SRR5849963–SRR5849966 ChIP; SRR5849967–SRR5849970 matched Input |
| Replication | Two biological replicates with two technical replicates each |
| Sequencing | Illumina HiSeq 2500, 51-bp single-end |
| Peak-call reference | NCBI RefSeq GCF_000004515.5 / Glycine_max_v2.1 |
What was asked of Pipette?
Analyze SRP112906/GSE101648 using matched inputs and biological replicates; assess ChIP-seq QC and concordance; call reproducible GLYMA.08G357600 peaks; annotate their genomic distribution; perform motif enrichment; identify candidate target genes; and interpret associated biological processes.
The request named Wm82.gnm2.DTC4 and the Wm82.gnm2.ann1.RVB6 gene models, but peak calling and ChIPseeker annotation used NCBI RefSeq GCF_000004515.5 with LOC-style IDs. A post-hoc Ensembl Plants mapping recovered GLYMA IDs for 72.6% of annotated peaks; this only partially resolves the mismatch.
FASTQ alignment, enrichment QC, replicate peak calling, annotation, motifs, and GO
Technical replicates were retained for library-level QC and merged within biological replicates for peak calling. ChIP and Input were processed in parallel before replicate-aware signal and peak comparisons.
numpy.int alias, then produced a stricter cross-replicate subset at IDR ≤ 0.05.All libraries passed the recorded QC gates and ChIP replicates were highly concordant
ChIP libraries aligned at 82.5–85.9%; matched inputs aligned at 98.8–98.9%.
ChIP duplication was 5.0–6.8% and Input duplication was 1.8–2.4%, below the report’s 10% warning threshold.
NSC ranged from 1.19–1.25 and RSC from 2.23–2.34; every ChIP library received Quality Tag 2/2.
ChIP–ChIP Pearson correlations ranged from 0.988–0.992; the IDR rank-correlation parameter was ρ = 0.92.

ChIP and Input clustered as distinct sample classes
The four ChIP libraries form one correlation cluster and the four matched Input libraries form another. High within-condition correlation supports consistent genome-wide coverage, while condition separation supports an enrichment signal rather than interchangeable ChIP and background profiles.

The broader consensus and stricter IDR results answer different questions
| Peak set | Count | Role in report |
|---|---|---|
| Biological replicate 1 | 27,627 | MACS3 q < 0.05 |
| Biological replicate 2 | 24,224 | MACS3 q < 0.05 |
| Pooled | 37,623 | Pooled-library MACS3 result |
| bedtools reproducible regions | 19,513 | BR1/BR2 intersection-derived set |
| Consensus final | 21,067 | Broader primary set used downstream |
| IDR passing | 7,693 | Stricter set at IDR ≤ 0.05 |
IDR tested 19,515 union peaks and retained 7,693, or 39.4%. The report’s 21,067 consensus peaks come from a different bedtools-and-pooled construction, so 7,693 should not be presented as 39.4% of the consensus set.
Fifty-nine percent of annotated consensus peaks were promoter-proximal
ChIPseeker assigned genomic features to 21,060 of the 21,067 consensus peaks. Using a ±2-kb promoter window, 12,421 annotated peaks were promoter-proximal: 9,543 within 1 kb of a TSS and 2,878 between 1 and 2 kb.
12,421 of 21,060 annotated peaks fell within ±2 kb of a TSS.
7,285 annotated peaks were classified intergenic or distal.
The post-hoc mapping produced 10,446 unique GLYMA identifiers associated with any peak.
Promoter-proximal peaks mapped to 7,445 unique GLYMA identifiers before GO coverage filtering.

15,298 of 21,060 annotated peaks (72.6%) received a GLYMA identifier. The remaining 5,762 peaks lack a direct Entrez-to-GLYMA cross-reference or fall on unplaced scaffolds and are excluded from GLYMA-based enrichment.
A G-box motif and seed-maturation processes emerged from the broader consensus set
STREME reported 16 significant de novo motifs from ±100-bp windows around the top 5,000 peak summits. TOMTOM matched nine motifs to entries in the available JASPAR2018 database, and FIMO then scanned summit sequences from all 21,067 consensus peaks.

Revised GLYMA-based enrichment returned 47 promoter-target Biological Process terms
The revised background contained GO annotations for 42,883 of 57,147 GLYMA identifiers, or 75.0%. Of 7,445 promoter-target GLYMA identifiers, 4,726 (63.5%) mapped into that background. The top term was protein dephosphorylation (62 genes; adjusted p = 2.49 × 10⁻⁶), with additional terms including circadian rhythm, lipid storage, photosynthesis, ethylene signaling, and gibberellic-acid signaling.

ChIP enrichment supports chromatin association; motif enrichment supports a sequence pattern; and GO over-representation supports a target-set association. None alone proves that GLYMA.08G357600 directly binds the G-box, regulates every nearby gene, or controls lipid storage, ABA signaling, or circadian biology. Those are hypotheses requiring ChIP-qPCR, DNA-binding assays, and perturbation transcriptomics.
Reference mapping, motif resources, and experimental scope limit the claims
- Peak calling and initial annotation used NCBI RefSeq GCF_000004515.5 rather than the requested Wm82.gnm2.DTC4 and RVB6 gene-model resources.
- Only 72.6% of the 21,060 annotated peaks received GLYMA identifiers; 5,762 peaks are absent from GLYMA-based gene lists and GO analysis.
- The preferred JASPAR2022 plant database could not be downloaded. TOMTOM used JASPAR2018, which contains some plant entries but is predominantly vertebrate-derived.
- Only 4,726 of 7,445 promoter-target GLYMA identifiers mapped into the revised GO background, leaving 36.5% outside the foreground analysis.
- Functional annotation includes electronically inferred terms and can vary in specificity and completeness.
- The study contains two biological replicates, one tissue stage, and one ChIP target; weak sites, other developmental contexts, and co-binding partners remain unresolved.
- Downstream annotation, motif scanning, and GO enrichment used the 21,067 consensus set rather than the 7,693-peak IDR subset.
- Nearby-gene assignment does not establish direct transcriptional regulation, and ChIP-seq does not establish whether binding is direct or mediated through a protein complex.
The run provides strong library-level QC and reproducible ChIP enrichment, a broader consensus peak set, and a stricter IDR subset. Confidence is lower for complete GLYMA target assignment, plant-specific motif identity, and causal biological mechanisms.
Peak sets, annotations, enrichment tables, signal tracks, and machine-readable flags
The files below are recorded in the session inventory. The website publishes the complete session PDF and selected report figures; the remaining artifacts are documented for provenance.
peaks/peaks.filtered.bed21,067-peak consensus set used for downstream analysispeaks/idr_output.txt7,693 peaks passing IDR ≤ 0.05peaks/annotated_peaks_glyma.csvPeak annotations with post-hoc GLYMA identifiersresults/GO_BP_enrichment_glyma_promoter.csv47 significant promoter-target GO Biological Process termsresults/pearson_correlation_matrix.txtEight-library coverage correlation matrixresults/qc_metrics.jsonMachine-readable alignment, duplication, and enrichment QCresults/analysis_flags.jsonUpdated active and resolved methodology warningsmotif_results/streme_out/streme.txtSixteen de novo motifs in MEME formatmotif_results/tomtom_out/tomtom.tsvJASPAR2018 motif comparisonsmotif_results/fimo_out/fimo.tsvPer-site motif occurrences across consensus summitsresults/chip_vs_input_pooled.bwPooled log₂ ChIP/Input genome-browser signal track