Get the basics right
A stationery retailer managed two linked customer databases containing a total of 3.7 million contacts, but never ran a systematic data quality check on them. The result: customers who weren’t being reached, and duplicate profiles that no one noticed. Stratics’ Data Quality Services set the record straight.
Objective
The retailer wanted to set one thing straight: a customer base you can rely on. Specifically, this meant systematically cleaning up the entire database in four areas—address, email, name, and phone number—while also improving deduplication. The goal: fewer unreachable contacts and more reliable segmentation, since each customer would appear only once in the system.
Challenge
Two customer databases, linked together, containing a total of 3.7 million valid contacts. Neither database had ever undergone a structured verification of addresses, email addresses, names, or phone numbers.
Inaccurate and incomplete data had two mutually reinforcing consequences. Contacts were unreachable due to incorrect addresses and invalid email addresses. And duplicate customer profiles went unrecognized because the underlying data differed too much from one another. As a result, a single customer appeared in the system as two or three separate, smaller profiles. The result: fragmented communication and segmentation built on quicksand.
The Stratics Approach
We activated Data Quality Services in phases: first on the smaller database, then on the larger one. Four checks ran simultaneously, each targeting a single data type. At the same time, we activated deduplication rules so that the correct underlying data could finally merge duplicate profiles into a single customer record.
Result
The 7.56 million corrections, across both databases combined (initial cleaning plus four months of follow-up), broken down by service:
| Address (DQA) | 3.436.133 |
| Phone (DQP) | 2.789.098 |
| Name (DQN) | 1.299.509 |
| Email (DQE) | 38.487 |
| Total | 7.563.227 |
Every correction affects accessibility and selection:
The most telling figure is the deduplication rate. In the 24 months prior to activation, an average of 1,500 to 2,000 duplicate profiles were automatically merged each month. During the months of the DQ activation, that number shot up to eight to ten times that level, peaking at 17,692 merges in a single month. Over the course of the entire data cleanup, more than 164,000 duplicate profiles were merged into a single correct customer profile—something that had previously been impossible due to conflicting or erroneous underlying data.
In terms of segmentation, that’s the real result. A customer who previously existed as two or three separate profiles is now a single, fully-fledged, correctly segmented profile that can be used for CV groups, RFM, and any subsequent selections.
"Segmentation based on contaminated data is segmentation based on noise. The benefits start with a solid foundation."
A 360° customer view stands or falls on data quality. Without a clean, deduplicated dataset, you’re not measuring your customers—you’re measuring your errors. Data quality is therefore not a side issue for MIP; it is the foundation upon which it is built.