Customer Case Study · MIP

Get the basics right

From Dirty Data to a Single, Accessible Customer View

A stationery retailer managed two linked customer databases containing a total of 3.7 million contacts, but never ran a systematic data quality check on them. The result: customers who weren’t being reached, and duplicate profiles that no one noticed. Stratics’ Data Quality Services set the record straight.

Objective

A file that's accurate and customers you reach

The retailer wanted to set one thing straight: a customer base you can rely on. Specifically, this meant systematically cleaning up the entire database in four areas—address, email, name, and phone number—while also improving deduplication. The goal: fewer unreachable contacts and more reliable segmentation, since each customer would appear only once in the system.

Challenge

3.7 million contacts, never verified

Two customer databases, linked together, containing a total of 3.7 million valid contacts. Neither database had ever undergone a structured verification of addresses, email addresses, names, or phone numbers.

Inaccurate and incomplete data had two mutually reinforcing consequences. Contacts were unreachable due to incorrect addresses and invalid email addresses. And duplicate customer profiles went unrecognized because the underlying data differed too much from one another. As a result, a single customer appeared in the system as two or three separate, smaller profiles. The result: fragmented communication and segmentation built on quicksand.

The Stratics Approach

Four checks, one clean foundation

We activated Data Quality Services in phases: first on the smaller database, then on the larger one. Four checks ran simultaneously, each targeting a single data type. At the same time, we activated deduplication rules so that the correct underlying data could finally merge duplicate profiles into a single customer record.

DQA · Address Standardization, add missing ZIP code and city, correction.
DQE · Email Syntax checking, detection of invalid domains and typos.
DQN · Name Validation and normalization of name and gender.
DQP · Phone Detection and formatting of landline and mobile numbers.

Result

7.5 million corrections and duplicates that finally matched up

3.7 million
Valid contacts
in two linked databases
7.56 million
Data Corrections
address, phone number, name, and email address
up to 10×
Peak in recognized doubles
compared to the normal monthly level
164.000+
Profiles Merged
in total, over the entire cleanup process

The 7.56 million corrections, across both databases combined (initial cleaning plus four months of follow-up), broken down by service:

Address (DQA)3.436.133
Phone (DQP)2.789.098
Name (DQN)1.299.509
Email (DQE)38.487
Total7.563.227

Every correction affects accessibility and selection:

  • 3.4 million address corrections mean fewer returned pieces of mail and a wider direct-mail reach.
  • 38,487 email corrections mean fewer bounces and a greater effective email reach.
  • 2.8 million phone number corrections make the data usable for telemarketing and text message targeting.

The most telling figure is the deduplication rate. In the 24 months prior to activation, an average of 1,500 to 2,000 duplicate profiles were automatically merged each month. During the months of the DQ activation, that number shot up to eight to ten times that level, peaking at 17,692 merges in a single month. Over the course of the entire data cleanup, more than 164,000 duplicate profiles were merged into a single correct customer profile—something that had previously been impossible due to conflicting or erroneous underlying data.

In terms of segmentation, that’s the real result. A customer who previously existed as two or three separate profiles is now a single, fully-fledged, correctly segmented profile that can be used for CV groups, RFM, and any subsequent selections.

"Segmentation based on contaminated data is segmentation based on noise. The benefits start with a solid foundation."

A 360° customer view stands or falls on data quality. Without a clean, deduplicated dataset, you’re not measuring your customers—you’re measuring your errors. Data quality is therefore not a side issue for MIP; it is the foundation upon which it is built.

Curious to know how many duplicate records are hidden in your database? Book a data quality scan, and we’ll show you based on your own data.
Ready to get your customer profile right? Contact us.