Skip to main content
A user group segmentation model uses unsupervised clustering algorithms to group users based on shared behavioral, demographic, or transactional patterns (Univio, Travis — Data Science for Human-Centered Product Design, overview talk). “Unsupervised” is the whole difficulty: there is no label telling me the right segments, so the model finds structure and I have to decide whether that structure is real and useful. This article walks the pipeline I actually run, from signals to a served segment label.

The signals I start from

Segmentation quality is decided before the model, by what goes into the feature vector. The standard collection gathers attributes like age, location, purchase history, app usage frequency, and session length. I organize those into three families, because each pulls out different behavior:
  • Demographic: age, location, account tenure, plan tier. Stable, coarse, good for reportable personas.
  • Behavioral: app usage frequency, session length, active days, feature adoption, recency of last visit. Fast-moving, this is where engagement signal lives.
  • Transactional: purchase history, order value, category mix, returns. Directly tied to revenue, usually sparse.
My default behavioral backbone is RFM — recency, frequency, monetary — because it is interpretable, cheap, and surprisingly strong. Everything else augments it. The one rule I follow: only use signals I can also compute for a brand-new user, or my offline segments will be impossible to assign at signup time.

Feature engineering

The preprocessing step is mechanical and unforgiving: clean missing values, encode categorical features, and scale numeric features using StandardScaler. I add the reasoning that makes it hold up:
  • Missing values are themselves a signal for sparse transactional data. I often fill with a sentinel and add a “was-missing” indicator rather than dropping the user, because “has no purchase history” is a segment, not a bug.
  • Categorical encoding. One-hot for low cardinality (country, plan); target/frequency encoding for high cardinality; and for free-text tags, I collapse near-duplicate tags into canonical clusters first — the method is on Tag clustering algorithm — so five spellings of “beach” become one interest feature instead of five diluted ones.
  • Scaling. Clustering minimizes distance, so a raw “age” (0-100) will drown a “sessions/day” (0-5). StandardScaler z-scores every column to zero mean / unit variance so each feature contributes on its own scale. I do not scale binary flags this way; scaling one-hot columns wrecks their meaning.
  • Binning and derived features. I convert continuous RFM into quantile bins, add ratios (spend per session, sessions per active day), and sometimes reduce dimensions with PCA or t-SNE — but only to simplify high-dimensional user data for plotting and analysis, never as a silent pre-step before K-Means on PCA components, which changes which distance I am actually optimizing.

Choosing the method

Each popular algorithm trades something specific, and I pick against the shape of the data, not familiarity.
  • K-Means Clustering: divides users into a fixed number (K) of groups by minimizing the distance between data points and cluster centers. My default when I need clean, reportable, roughly-equal personas. It assumes convex, isotropic blobs and is sensitive to scaling and outliers.
  • Hierarchical Clustering: builds a multi-level tree of clusters to analyze user segments at various levels of granularity. When the business wants “broad personas that drill down,” this beats K-Means because I cut the dendrogram at whatever depth a stakeholder needs. Costly at scale (O(n^2)+).
  • DBSCAN: finds arbitrarily shaped clusters and handles noise or outlier users effectively. For behavioral data with a huge one-off tail (bots, single-session tourists), DBSCAN’s explicit “noise” label is a feature, not a fallback.
  • Gaussian Mixture Models (GMM): soft clustering where a user belongs to multiple segments with assigned probabilities (Deolesh Panaskar, Customer Segmentation in ML). When a user genuinely is 60% deal-seeker / 40% researcher, a hard label lies; GMM gives me the mixture and a likelihood threshold for “confident enough to target.”

Finding K and evaluating

The Elbow Method — calculating Within-Cluster Sum of Squares (WCSS) and looking for the bend where adding a cluster stops cutting distortion much — is the standard way to pick the ideal number of segments. I treat the elbow as a hint, not an answer, because it is frequently a smooth curve with no clean elbow. I triangulate:
  • Silhouette score across a K range to check separation, and for GMM/DBSCAN the metric changes (no WCSS), which is exactly why I pick the metric after the algorithm.
  • Stability: rerun on different days/seed and confirm the same segments survive. A segment that shuffles every night is noise, and no marketer should target it.
  • Interpretability: I profile each cluster’s feature means. A cluster I cannot describe in one sentence to a business owner is not a segment, it is an artifact.
  • Business validation: segments must differ on a metric I care about that was not a clustering feature (e.g., later conversion or churn). If they do not, the segmentation is decorative.
  • Drift: I monitor population share per segment over time with a PSI-style check, because “the model is fine” and “the users moved” look identical if I am not watching counts.

Offline computation, online serving

This is where most segmentation projects quietly fail, and it is the part I design first. The heavy clustering runs offline — a nightly or weekly batch job over the full population. For a large user base, that batch is exactly the embarrassingly-parallel scoring problem I work through on Calculating time to run a pipeline, and the storage budget for keeping per-user profiles is on User stats. The batch job writes user_id -> segment_id (and, for GMM, the probability vector) into a key-value store. Serving is a separate, online path, and it has three cases:
  1. Existing user: at request time I look up the precomputed segment label — a single fast key-value read, not a re-cluster. This is what powers personalized marketing and recommendation engines in the workflow’s deployment step.
  2. Existing user, features changed but before next batch: if behavior is fast-moving, the offline label is stale. I either accept the staleness for coarse personas, or run a lightweight online scorer — a frozen logistic/centroid model that maps the live feature vector to the nearest centroid without touching the batch pipeline.
  3. New user at signup: there are no behavioral features yet. Cold-start users go to an explicit “new/unknown” segment and get assigned to a real one only after enough signals accumulate. I never force-cluster a user with two features.
The key mental model: the cluster model is offline; the assignment can be online. I persist the trained model (centroids, scaler parameters, PCA basis) so the online scorer reconstructs the exact feature space the model was trained in. A mismatch between offline and online feature computation is the most common reason served segments diverge from the ones I validated.
Never re-fit the clustering inside the serving path, and never let online features drift from the offline definitions. If your offline pipeline z-scores with StandardScaler fit on last month’s data, the online scorer must reuse those frozen mean/std — recomputing them per request silently changes the distance geometry and renumbers your segments under the marketing team’s feet.

Failure modes I design against

  • Scaling bugs: forget StandardScaler and monetary values with big magnitudes own the clusters. I check feature means after scaling.
  • K over-fit to one day: pick K on a stability window, not a snapshot.
  • Uninterpretable segments: a cluster nobody can name never gets used; fold tags and drop features that only add noise.
  • Stale labels: offline-only serving for fast-moving behavior serves yesterday’s segment; add an online scorer or shorten the batch cadence.
  • Uneven segments: with a power-law user distribution, huge hubs distort K-Means centroids; either switch to DBSCAN/GMM or pre-balance with dimensionality reduction.
The standard workflow — collect, preprocess, optionally reduce dimensionality, find the optimal clusters, deploy labels to drive personalization — is well covered end to end (GeeksforGeeks, Machine Learning Mastery), with interpretable-representation approaches in the research literature, and worked walkthroughs in video form, here, and again in the overview talk — the differentiator between a demo and a shipped system is the offline/online split above and the evaluation discipline before it.