Back to Case Studies

Case study · Single-cell RNA-seq

Single-Cell Clustering and Marker Analysis of Human Pancreatic Islets

The GSE85241 CEL-seq2 matrix was audited, filtered, normalized, and clustered with Scanpy. The article reports marker-supported working cell-type labels while separating rare-cell hypotheses from an ERCC-dominated technical cluster.

Homo sapiens3,072 cells4 donorsGSE85241Scanpy

Original 20-page session report · PDF · 4.1 MB

At a glance

A cell-atlas analysis with explicit QC and annotation boundaries

This is a cluster-discovery and marker-identification study, not a treatment-versus-control test. Donor and plate structure were reconstructed before the decision not to integrate batches.

Input cells3,072
After QC2,681
Leiden clusters18
Donors4
01 / Scientific context

Which transcriptionally distinct populations are present in the islet dataset?

The four donor and 32 plate identifiers encoded in cell barcodes were reconstructed and assessed before clustering. The objective was to find cell populations and their marker genes.

Which major and rare cell populations are supported by the GSE85241 expression matrix, and which clusters are better explained by technical artifacts?
SourceNCBI GEO GSE85241
Input19,140 genes × 3,072 CEL-seq2 cells
Sample structure4 donors × 8 plates × 96 wells
ObjectiveQC, clustering, working annotations, and per-cluster markers
02 / Analysis workflow

Count-integrity, QC, batch assessment, clustering, and markers

Count-integrity gateNon-integer values were identified as CEL-seq2 UMI-collision-corrected transcript counts and treated as count-like data.
Quality controlCells with fewer than 200 detected genes and genes in fewer than three cells were removed; 2,681 cells and 17,792 genes remained.
Normalize and reduceScanpy normalization, log1p, 2,000 highly variable genes, scaling, PCA, a 20-neighbor graph, and UMAP were applied.
Assess batch structureThe donor silhouette score on the top 10 PCs was −0.071, so Harmony or BBKNN integration was not applied.
Cluster and annotateLeiden at resolution 1.0 produced 18 clusters; Wilcoxon tests and a pancreas marker dictionary provided working labels.
Decision gateThe unintegrated PCA and neighbor graph were retained because donor identity did not dominate the leading components. This does not prove that every donor or plate effect is absent.
03 / Results

Major pancreatic populations were resolved—with one clear technical artifact

Marker profiles supported working annotations for alpha, beta, delta, PP/gamma, acinar, ductal, stellate, and endothelial populations. The run retained 36,833 significant marker gene–cluster associations after Bonferroni correction.

Expected populations

Large clusters carried canonical markers including INS, SST, PPY, PRSS1, CFTR, SPARC, and PLVAP.

Artifact cluster

Cluster 6 contained 247 low-complexity cells dominated by ERCC spike-ins and was not treated as a biological delta-cell population.

Rare-cell hypothesis

Cluster 13 contained 10 KIT/CPA3/TPSB2-positive cells consistent with candidate mast cells, pending validation.

Batch decision

Donor identity did not dominate the leading PCs, so clustering used the uncorrected representation.

Pre-integration UMAP of pancreatic islet cells colored by donor
Figure 1. Pre-integration UMAP by donor. Donors overlap in the displayed embedding; the PCA-based silhouette score informed the decision not to integrate.
UMAP of 2,681 pancreatic islet cells colored by 18 Leiden clusters
Figure 2. Leiden clustering. The 2,681 QC-passed cells form 18 computational clusters. Separation alone does not validate the cell-type labels.
Dot plot of top marker genes across 18 Leiden clusters
Figure 3. Top cluster markers. Dot size and color summarize detection and expression patterns used to review working annotations.
04 / Limitations

Working labels and QC decisions remain hypotheses to review

  • The matrix contains collision-corrected, non-integer counts rather than raw molecule counts.
  • No mitochondrial genes were available, so a conventional mitochondrial-fraction QC gate could not be applied.
  • Several small clusters are ambiguous and the 10-cell mast-cell cluster needs orthogonal validation.
  • Marker-dictionary annotations are working labels, not reference-mapping or experimental confirmation.
  • The dataset has no treatment contrast; cluster markers should not be described as differential treatment responses.
  • Skipping integration is a documented choice, not proof that all plate-level effects are negligible.
Annotation boundary

The report supports a cluster atlas with marker-based working labels. It does not support definitive discovery of a new cell type or disease mechanism.

05 / Reproducibility and outputs

Artifacts recorded in the Pipette session

adata_raw.h5adImported expression matrix
adata_qc.h5adQC-filtered cells and genes
adata_clustered.h5adEmbedding, clusters, and annotations
cluster_markers.csv36,833 significant marker associations
cluster_annotations.csvWorking labels and match scores
analysis_flags.jsonMachine-readable caveats
Original Pipette session report20-page PDF · 4.1 MB · generated July 6, 2026. The HTML article is the primary case study; the PDF is the preserved session record.

Analyze a single-cell expression matrix

Bring a count matrix or public accession and specify the QC, clustering, and annotation questions you need to audit.

Start in Pipette