A customer appears twice in the CRM. A supplier has one name in accounts and another in the purchasing spreadsheet. The operations team keeps its own list because they no longer trust the data elsewhere. If this sounds familiar, learning how to fix duplicate data is not just a clean-up exercise. It is a way to stop routine work becoming slower, less reliable and harder to manage.
Duplicate records are usually a symptom of a process problem. Delete a few obvious copies without addressing that problem and they will return, often within weeks. The practical aim is to make each record trustworthy, give it a clear owner, and prevent the same information being created in several places.
Why duplicate data causes more than admin irritation
At first, duplication looks harmless. Two contact records might mean a few extra rows in a report. But the cost grows as the business relies on that information for quoting, delivery, invoicing, customer service and forecasting.
A sales report can overstate the number of active customers. A warehouse team might dispatch to an old address. Finance may chase an invoice against the wrong account. Staff spend time comparing spreadsheets and asking which version is right, rather than completing the work in front of them.
The less obvious effect is loss of confidence. Once people suspect the main system is inaccurate, they start maintaining private spreadsheets, email folders and handwritten notes. Those workarounds create further copies and move the business further away from a single view of its operation.
Find the source before merging records
The first job is not to run a bulk de-duplication tool. It is to understand how duplicates are getting in. In a growing business, there are normally a few familiar causes.
The same customer may be entered by sales, accounts and operations because each team uses a different system. Staff may create a new record when a search returns no obvious match, perhaps because one entry says “ABC Ltd” and another says “A.B.C. Limited”. Imports from spreadsheets, website forms and external software can add records without checking what already exists.
Sometimes the issue is more structural. A system has no agreed unique identifier, or it allows fields such as company name, email address and postcode to be entered inconsistently. In other cases, integrations are set up to create new records every time rather than update an existing one.
Map the journey for the data that matters most. Ask where a record starts, who adds to it, which systems receive it, and what triggers an update. A simple process map is often enough to reveal that the duplicate is not being created by careless staff. The system is making the wrong action the easiest action.
Decide what counts as a duplicate
Not every similar record should be merged. Two people at the same company may need separate contact records. A customer may have several trading locations. A supplier could use different legal entities for different services.
Before changing anything, set clear matching rules. For individual contacts, email address may be a sensible primary match, with name and telephone number as supporting checks. For companies, a Companies House number, account code or VAT number is usually more dependable than the trading name alone. For stock, a SKU or internal product code should carry more weight than a description.
It also helps to define which record wins when two records genuinely represent the same thing. The newest entry is not automatically the best one. One record may contain current contact details while another holds the transaction history. The right approach may be to retain the record linked to orders or invoices and copy across verified information from the other entry.
This is where judgement matters. Automatic matching saves time on large datasets, but overly aggressive rules can join records that should remain separate. Start cautiously, test the logic on a small sample and allow uncertain matches to be reviewed by someone who understands the day-to-day operation.
How to fix duplicate data without creating new errors
Once the rules are agreed, the clean-up can be handled in a controlled sequence. Keep a copy of the original data and make the changes in a test environment where possible. If the system does not offer one, work in manageable batches and retain a record of every merge, deletion and amendment.
First, standardise the fields used for matching. This may mean removing extra spaces, applying consistent capitalisation, separating first and last names, or agreeing whether company suffixes such as Ltd are included. Standardisation does not make data perfect, but it makes comparison far more reliable.
Next, identify likely matches using the agreed rules. Exact matches, such as identical email addresses or account numbers, can often be dealt with quickly. Near matches need more care. “Smith & Sons Ltd” and “Smith and Sons Limited” may be obvious to a person but not to a basic spreadsheet formula.
Then merge or retire records in line with the agreed policy. Do not simply delete every duplicate row. Check for linked orders, support cases, invoices, notes, permissions and audit history. A record that appears redundant may be holding information another team relies on.
Finally, validate the result. Check a sample of merged records, rerun key reports and ask the people who use the system most whether customer history, balances and operational information still look correct. The clean-up is not finished until normal work can continue safely.
Put one system in charge of each type of data
A clean database will not stay clean if several tools are treated as the master record. Decide where each important item of information should live.
For example, the CRM may own customer and contact details, the accounts package may own invoice status, and an operations system may own job progress. Other systems can use that information, but they should not all be able to overwrite it independently.
This does not necessarily mean replacing every application. Many businesses need separate systems because sales, finance and delivery have different needs. The important point is that the handover between them is deliberate. If data moves through an integration, define whether it should create, update or ignore a record, and what identifier it uses to find the correct match.
A bespoke system can be particularly useful where the real process sits between several off-the-shelf tools. Rather than asking staff to copy and paste details between them, a well-designed workflow can validate records, assign internal IDs and pass only the required information to the right place.
Make good data the easy option for staff
Prevention should happen at the point of entry. A clear search screen that shows likely matches before a new record is created will prevent more duplicates than a monthly clean-up ever can. So will required fields where they are genuinely necessary, sensible validation and simple guidance on naming conventions.
Avoid turning this into bureaucracy. Requiring twenty fields before someone can log a new enquiry will encourage shortcuts and incomplete entries. Capture what the business needs at that stage, then collect further detail as the relationship or job progresses.
Access rules also matter. Some teams need to create records, while others only need to update existing ones. Where several people can alter key information, an audit trail helps identify what changed and why. This is useful for accountability, but it also helps improve the process when the same mistakes keep appearing.
Monitor the problem instead of waiting for the next clean-up
Duplicate data is easier to manage when it is treated as an operational measure, not an annual IT task. Set up a simple exception report for records with matching emails, phone numbers, account numbers or addresses. Review it regularly, with ownership assigned to someone who can make a decision.
The right frequency depends on volume. A business adding a handful of customers each month may only need a quarterly review. A busy team taking orders from several channels may need daily checks built into the workflow. The key is to spot a pattern before hundreds of records need attention.
When duplicates keep returning, look at the workflow rather than blaming the person entering them. A repeated issue usually means the system is unclear, the integration is unreliable or the existing process is too slow for the pressure staff are under.
Good data is not about making every field immaculate for its own sake. It is about giving people enough confidence to act without cross-checking three spreadsheets first. Build that confidence into the process, and the business gains time back every day.

