Marketing & Content

Data cleaning and deduplication:
one record per customer, not five

A customer list, a product catalog or a CRM export almost always has the same entity spelled three different ways and rows nobody dares delete. An agent finds the duplicates and standardises the formatting. It brings the genuinely ambiguous merges to a person instead of guessing.

from$500
Timeline3 to 10 days
What is includedAudit of the current data: duplicate rate, format inconsistencies, missing fieldsFuzzy matching across names, emails, phone numbers and addressesStandardisation of formats (dates, phone numbers, currency, units)A review queue for merges below a confidence thresholdAudit log of every merge and edit with a rollback option
85-95%of clear duplicates merged automatically at a high confidence threshold, typical range
0destructive merges applied without a human review on anything below that threshold
3-10 daysfrom a raw export to a cleaned, deduplicated dataset

Why the same customer has three names

A dataset fed by more than one source for more than a year almost never stays clean on its own. A customer signs up once through a web form and once through a sales call logged by a different manager. Now the same person exists as “J. Smith,” “John Smith” and an email-based variant in three separate rows, each with slightly different order history attached. A product catalog imported from two suppliers has the same SKU under two different internal codes with different unit labels. Nobody’s fault, exactly. It is what happens when data arrives from different places over time, with nobody’s job being to reconcile it.

The cost is not abstract. Marketing sends the same promotion to the same person three times under three record variants, which looks sloppy and burns send quota. A support agent pulls up one version of a customer’s record and misses the order history under the duplicate. They give an answer that contradicts what the customer was told last time. A finance report counts revenue against a product twice, because two SKUs that are really one product were never merged. The discrepancy takes an afternoon to track down.

Manual deduplication does not scale past a few hundred rows. People start making the same judgment calls over and over, inconsistently. Fatigue sets in, and the twentieth “is this the same person” decision gets less care than the first. Nobody wants to run a bulk merge by hand on live customer or financial data, because one wrong merge is a genuinely bad day. So the cleanup gets postponed until the dataset is too painful to work with, which only makes the postponement worse.

What the audit finds, and what happens next

The agent starts with an audit of the dataset as it actually exists: duplicate rate by match type, formatting inconsistencies, inconsistent units or currencies. It also flags which fields are missing often enough to matter. This runs on a copy of your data, never the live system, so the audit itself carries zero risk.

Matching runs on fuzzy logic across names, emails, phone numbers and addresses rather than exact string matching. That catches the “John Smith” versus “J. Smith” cases simple tools miss. Every match comes with a confidence score. Above a threshold you set, matches merge automatically, keeping the most complete and most recent values and logging exactly what was merged from where. Below that threshold, the match goes to a review queue with both candidate records shown side by side. A person makes the genuinely ambiguous call instead of the agent guessing on a coin flip.

Formatting gets standardised at the same time. Dates into one format, phone numbers into what your system expects, currency and units normalised so a report does not silently mix pounds and kilograms. The whole pass gets written to an audit log, every automatic merge, every queued decision. You can see and reverse any single change, not just restore from a full backup.

For data that keeps accumulating duplicates from an ongoing source, like a form that is not enforcing uniqueness, the same matching logic runs on a schedule. New duplicates get caught as they arrive, instead of letting the dataset drift dirty again a month after the first cleanup.

Where a person still has to decide

Every merge below the confidence threshold stays with a person. So does any decision to delete rather than merge a record. Two duplicates can disagree on something important, a different address, a different price. A person makes that final call, since they know the business context the agent does not have.

How a bad merge gets undone

Every automatic merge is logged with both original records attached and can be reversed individually, not just restored from a full backup. The first pass on any new dataset runs as a dry run against a copy. Nothing touches the live system until that dry run has been reviewed and the confidence threshold tuned against real examples from your own data. Access during the project is scoped to the fields needed for matching, not the full dataset.

Price and timeline

Option Price What it covers Timeline
Single automation from $500 One dataset cleaned and deduplicated, up to ~50,000 rows, full audit log 3 to 10 days
Department package from $2,500 Cleanup plus spreadsheet workflow automation and a recurring monthly clean for the team that owns the data 2 to 4 weeks

Running cost for a recurring monthly clean is usually $10 to $50 in model usage depending on dataset size, with a budget cap set before launch.

This pairs directly with spreadsheet workflows automation once the underlying data is clean enough to automate on top of. Weekly reports in plain language and dashboard commentary are both only as trustworthy as the data feeding them. See the automation-everything overview and the AI agents service page for the broader catalogue. We did a version of this at real scale in the self-hosted ERP recovery for a supplements factory, exporting and rebuilding 51 tables and 12,039 rows. The two-brand analytics warehouse depends on the same kind of clean, joined data underneath its dashboards.

If a dataset has reached the point where nobody fully trusts a report built on it, get in touch. We will send back an audit of what is actually wrong with it before touching anything.

Tired of doing this by hand? We can take the whole routine off your team, not only this step: Routine takeover, from $400 →

FAQ

How much does data cleaning and deduplication cost?

From $500 for a one-time cleanup of a dataset up to around 50,000 rows, delivered with a full audit log. Larger datasets, multiple source systems, or a recurring monthly clean are $1,500 and up.

How long does a cleanup take?

3 to 10 days depending on dataset size and how many fields need standardising. Most of that time goes into setting the matching rules correctly for your data, not running the match itself.

Which tools and formats does it connect to?

CSV and Excel exports, direct database connections (PostgreSQL, MySQL), and API pulls from CRMs like HubSpot, amoCRM or Salesforce. Output goes back into the same system or to a clean file, whichever you need.

What if the AI merges two records that should stay separate?

Every merge below a confidence threshold you set goes to a review queue instead of happening automatically. Every automatic merge is logged with both original records attached, so it can be reversed. We never delete the pre-cleanup data.

Is our customer data safe during the process?

The first pass always runs on a copy, never your live system, and access is limited to the fields needed for matching. Your data stays in your own database or spreadsheet. We do not retain a copy after the project closes unless you ask us to keep maintaining it.

Start here

Tell us the problem.
We bring the system.

A 30-minute call, then a written plan with numbers within 48 hours. No obligation. If we are not the right fit, we will say so and point you to someone who is.

LIKE WHAT YOU SEE?

This site is our work.
Want one like it?

Ten languages, no page builder, launched in 2026 by a team working since 2015. We can build the same quality into your site.

  • 10 languages
  • Since 2015
Get a site like this →