Skip to content
AI & Automation7 min read

How Entity Resolution Improves AI: Connecting the Same Customer, Product and Company Across Systems

How entity resolution links the same customer, product and company across CRM, ERP, ecommerce, support and marketing, and why AI answers depend on it.

01

Quick answer

Entity resolution works out which records refer to the same real-world entity (a customer, a product, a company) across systems that each use their own IDs, and links them under one canonical identity. Entity matching is the comparison step inside it: scoring whether two records describe the same thing.

AI depends on it because assistants and agents assemble answers from many systems. If the CRM, ERP, ecommerce platform, helpdesk and marketing tool each hold a different version of the same customer, the AI sees four partial customers. Its answers become incomplete, and its actions can be wrong.

02

Why the same entity ends up with many identities

Every system creates records for its own purpose. Sales creates an account in the CRM with the company's trading name. Finance creates a customer in the ERP with the legal name and tax ID. The online store creates a customer per email address, so a buyer who used a work and a personal email exists twice. Support creates contacts from incoming emails. Marketing imports leads from events with misspelled names. Products suffer the same: a supplier's part number, an internal SKU, a marketplace listing ID and a bundle code for the same item.

None of these systems is wrong. They simply do not share an identity, and most were never designed to.

One customer across five systems (diagram)
CRM          account  "Acme Industrial"      A-1007
ERP          customer "ACME Industrial Ltd"  C-55120  VAT GB123...
Ecommerce    customer "j.doe@acme.co"        S-88812
Support      contact  "John Doe <jd@acme.co>" T-3391
Marketing    lead     "Jon Doe, Acme Ind."   M-40012
                     │
                     ▼  entity resolution
        CANONICAL ACCOUNT  ent:acme-industrial
        └─ person ent:john-doe (2 emails, 3 records)
03

How inconsistent identities degrade AI answers and actions

The failures are rarely dramatic, which is why they go unnoticed. Some hypothetical but typical examples:

SituationWhat the AI doesWhy
Support assistant summarises a customerMisses an open complaint and a recent refundTicket and refund sit under different IDs
Sales agent drafts an upsell emailPitches a product the customer bought last week onlineOnline orders are linked to a personal email
Finance assistant reports top accountsRanks a group's subsidiaries separately and too lowNo parent-child link across entities
Marketing agent sends a win-back offerTargets someone who renewed through a resellerReseller orders use the reseller's account
Product assistant answers stock questionsSays an item is out of stockInventory is under the supplier part number, not the SKU

Key takeaway

For agents, bad identity is a safety issue, not only a quality issue. An agent that acts on a partial customer can send the wrong offer, refund twice or contact someone who opted out.

04

Customer, product and company identity

Customer identity usually combines verified identifiers (account number, verified email, phone) with fuzzy attributes (name, address). In B2C, households and multiple emails complicate it; consent and privacy rules limit what can be linked. Product identity relies on GTINs, manufacturer part numbers and supplier codes where they exist, and on attribute matching (brand, model, size, colour) where they do not; variants and bundles need explicit relationships rather than merges. Company identity uses legal names, registration and tax numbers, domains and addresses, and needs hierarchies: parent groups, subsidiaries, sites and trading names.

In each case, decide what 'the same' means for your business. A parent company and its subsidiary are different entities with a relationship, not duplicates; that decision belongs in your business ontology.

05

How entity resolution works

Most pipelines follow the same stages. Standardize names, addresses, phone numbers and codes. Block records into candidate groups (same postcode, same email domain) so you avoid comparing everything with everything. Compare candidate pairs on several attributes. Score each pair, with rules for exact identifiers and a probabilistic or ML model for fuzzy evidence. Decide using thresholds: auto-link, auto-reject or send to review. Cluster linked pairs into entities and assign canonical IDs. Maintain links as new records arrive, and keep the ability to split wrong merges.

The classic statistical approach is the Fellegi-Sunter model, which weighs how much each agreeing or disagreeing attribute changes the odds of a match. The UK Ministry of Justice's open-source library Splink implements it at scale and is used to link people across courts, prisons and probation data (Splink documentation).

Entity resolution pipeline (diagram)
CRM ─┐
ERP ─┤
Shop ├─▶ standardize ─▶ block ─▶ compare pairs ─▶ score
Help ─┤                                           │
Mktg ─┘                         ┌─────────────────┼──────────┐
                                ▼                 ▼          ▼
                           auto-link        human review  no match
                                └────────┬────────┘
                                         ▼
                        cluster ─▶ canonical IDs + crosswalk
                                         │
                 ┌───────────────────────┼─────────────────────┐
                 ▼                       ▼                     ▼
          data products            retrieval metadata     agent tools
06

Architecture: where resolved identity lives

The output of entity resolution is a crosswalk: a table mapping every source record ID to a canonical entity ID, with match confidence and the date linked. Keep source records intact. Downstream, the canonical ID should appear everywhere AI consumes data: in data products, as metadata on retrieval chunks (so a search can filter to one customer's documents across systems) and in agent tool inputs and outputs. Agents should look up a customer once and receive the canonical ID plus the linked source IDs, rather than searching each system by name.

Some organizations run this as a master data management (MDM) programme with golden records; others keep a lighter crosswalk. For AI, the essential parts are the same: a stable canonical ID, links back to sources and a way to fix mistakes.

07

Practical examples

Ecommerce brand (hypothetical). Shopify customers, marketplace orders, helpdesk contacts and email subscribers are resolved on verified email, phone and address. The support assistant now sees every order and ticket for a person, and the marketing agent suppresses offers to anyone with an open complaint.

B2B distributor (hypothetical). CRM accounts and ERP customers are linked on tax ID and domain, then grouped by legal parent. Account managers' assistants report group-level revenue that matches finance, and the renewal agent sees all sites under a contract.

Product catalogue (hypothetical). Supplier part numbers, internal SKUs and marketplace listings are linked through GTINs and attribute matching, so the product assistant answers stock and compatibility questions with the right item.

08

Using LLMs in entity resolution

Language models are useful for the hard middle: normalizing messy company names, judging ambiguous pairs with free-text evidence and explaining why two records probably match. They are a poor fit for the full comparison workload because of cost, latency and inconsistency. A sensible pattern is to block and score with conventional methods, then send only uncertain pairs to an LLM with a structured output ('match', 'no match', 'unsure', with reasons), and route 'unsure' to a person. Measure precision and recall on a labelled sample before trusting it.

09

Implementation checklist

  • Choose the entities that matter for your first AI use case
  • Define what 'same entity' means, including hierarchies and exclusions
  • Inventory identifiers in each source and how reliable they are
  • Standardize, block, compare and score; start with exact identifiers
  • Set thresholds for auto-link, review and reject from a labelled sample
  • Keep a crosswalk with confidence and link dates; never overwrite sources
  • Support un-merging when a link is wrong
  • Propagate canonical IDs into data products, retrieval metadata and tools
  • Respect consent and privacy rules when linking personal data
  • Monitor match rates and review queues as new data arrives
10

Common mistakes

Over-merging is worse than under-merging: linking two different customers can expose one person's data to another and lead agents to act on the wrong account. Tune for precision first. Other mistakes include matching on names alone, treating parent and subsidiary as duplicates, running resolution once instead of continuously, and resolving identities in the warehouse while the AI still queries source systems by name. Our guides to data quality for AI and CRM automation cover related clean-up work.

Connecting customer and product data for AI?

ZSpace Labs integrates CRM, ERP, ecommerce and support systems so AI assistants and agents see one consistent view. See AI automation and Shopify development.

Start a Project
11

Conclusion

Entity resolution is quiet infrastructure with a large effect on AI. It decides whether an assistant sees one customer or four fragments, and whether an agent acts on the whole picture. Define what 'same' means, match with rules and scores, review the uncertain middle, keep a reversible crosswalk and carry canonical IDs into everything AI reads.

FAQ

Common questions.

Entity resolution is the process of working out which records, across one or many systems, refer to the same real-world thing, such as the same customer, product or company, and linking them under one identity.

Get in touch

Have a project in mind?

Whether you're building a new digital product, improving an existing website, or looking to automate part of your business — let's talk.