How Entity Resolution Improves AI: Connecting the Same Customer, Product and Company Across Systems
How entity resolution links the same customer, product and company across CRM, ERP, ecommerce, support and marketing, and why AI answers depend on it.
Quick answer
Entity resolution works out which records refer to the same real-world entity (a customer, a product, a company) across systems that each use their own IDs, and links them under one canonical identity. Entity matching is the comparison step inside it: scoring whether two records describe the same thing.
AI depends on it because assistants and agents assemble answers from many systems. If the CRM, ERP, ecommerce platform, helpdesk and marketing tool each hold a different version of the same customer, the AI sees four partial customers. Its answers become incomplete, and its actions can be wrong.
Why the same entity ends up with many identities
Every system creates records for its own purpose. Sales creates an account in the CRM with the company's trading name. Finance creates a customer in the ERP with the legal name and tax ID. The online store creates a customer per email address, so a buyer who used a work and a personal email exists twice. Support creates contacts from incoming emails. Marketing imports leads from events with misspelled names. Products suffer the same: a supplier's part number, an internal SKU, a marketplace listing ID and a bundle code for the same item.
None of these systems is wrong. They simply do not share an identity, and most were never designed to.
CRM account "Acme Industrial" A-1007
ERP customer "ACME Industrial Ltd" C-55120 VAT GB123...
Ecommerce customer "j.doe@acme.co" S-88812
Support contact "John Doe <jd@acme.co>" T-3391
Marketing lead "Jon Doe, Acme Ind." M-40012
│
▼ entity resolution
CANONICAL ACCOUNT ent:acme-industrial
└─ person ent:john-doe (2 emails, 3 records)How inconsistent identities degrade AI answers and actions
The failures are rarely dramatic, which is why they go unnoticed. Some hypothetical but typical examples:
| Situation | What the AI does | Why |
|---|---|---|
| Support assistant summarises a customer | Misses an open complaint and a recent refund | Ticket and refund sit under different IDs |
| Sales agent drafts an upsell email | Pitches a product the customer bought last week online | Online orders are linked to a personal email |
| Finance assistant reports top accounts | Ranks a group's subsidiaries separately and too low | No parent-child link across entities |
| Marketing agent sends a win-back offer | Targets someone who renewed through a reseller | Reseller orders use the reseller's account |
| Product assistant answers stock questions | Says an item is out of stock | Inventory is under the supplier part number, not the SKU |
Key takeaway
For agents, bad identity is a safety issue, not only a quality issue. An agent that acts on a partial customer can send the wrong offer, refund twice or contact someone who opted out.
Customer, product and company identity
Customer identity usually combines verified identifiers (account number, verified email, phone) with fuzzy attributes (name, address). In B2C, households and multiple emails complicate it; consent and privacy rules limit what can be linked. Product identity relies on GTINs, manufacturer part numbers and supplier codes where they exist, and on attribute matching (brand, model, size, colour) where they do not; variants and bundles need explicit relationships rather than merges. Company identity uses legal names, registration and tax numbers, domains and addresses, and needs hierarchies: parent groups, subsidiaries, sites and trading names.
In each case, decide what 'the same' means for your business. A parent company and its subsidiary are different entities with a relationship, not duplicates; that decision belongs in your business ontology.
How entity resolution works
Most pipelines follow the same stages. Standardize names, addresses, phone numbers and codes. Block records into candidate groups (same postcode, same email domain) so you avoid comparing everything with everything. Compare candidate pairs on several attributes. Score each pair, with rules for exact identifiers and a probabilistic or ML model for fuzzy evidence. Decide using thresholds: auto-link, auto-reject or send to review. Cluster linked pairs into entities and assign canonical IDs. Maintain links as new records arrive, and keep the ability to split wrong merges.
The classic statistical approach is the Fellegi-Sunter model, which weighs how much each agreeing or disagreeing attribute changes the odds of a match. The UK Ministry of Justice's open-source library Splink implements it at scale and is used to link people across courts, prisons and probation data (Splink documentation).
CRM ─┐
ERP ─┤
Shop ├─▶ standardize ─▶ block ─▶ compare pairs ─▶ score
Help ─┤ │
Mktg ─┘ ┌─────────────────┼──────────┐
▼ ▼ ▼
auto-link human review no match
└────────┬────────┘
▼
cluster ─▶ canonical IDs + crosswalk
│
┌───────────────────────┼─────────────────────┐
▼ ▼ ▼
data products retrieval metadata agent toolsArchitecture: where resolved identity lives
The output of entity resolution is a crosswalk: a table mapping every source record ID to a canonical entity ID, with match confidence and the date linked. Keep source records intact. Downstream, the canonical ID should appear everywhere AI consumes data: in data products, as metadata on retrieval chunks (so a search can filter to one customer's documents across systems) and in agent tool inputs and outputs. Agents should look up a customer once and receive the canonical ID plus the linked source IDs, rather than searching each system by name.
Some organizations run this as a master data management (MDM) programme with golden records; others keep a lighter crosswalk. For AI, the essential parts are the same: a stable canonical ID, links back to sources and a way to fix mistakes.
Practical examples
Ecommerce brand (hypothetical). Shopify customers, marketplace orders, helpdesk contacts and email subscribers are resolved on verified email, phone and address. The support assistant now sees every order and ticket for a person, and the marketing agent suppresses offers to anyone with an open complaint.
B2B distributor (hypothetical). CRM accounts and ERP customers are linked on tax ID and domain, then grouped by legal parent. Account managers' assistants report group-level revenue that matches finance, and the renewal agent sees all sites under a contract.
Product catalogue (hypothetical). Supplier part numbers, internal SKUs and marketplace listings are linked through GTINs and attribute matching, so the product assistant answers stock and compatibility questions with the right item.
Using LLMs in entity resolution
Language models are useful for the hard middle: normalizing messy company names, judging ambiguous pairs with free-text evidence and explaining why two records probably match. They are a poor fit for the full comparison workload because of cost, latency and inconsistency. A sensible pattern is to block and score with conventional methods, then send only uncertain pairs to an LLM with a structured output ('match', 'no match', 'unsure', with reasons), and route 'unsure' to a person. Measure precision and recall on a labelled sample before trusting it.
Implementation checklist
- Choose the entities that matter for your first AI use case
- Define what 'same entity' means, including hierarchies and exclusions
- Inventory identifiers in each source and how reliable they are
- Standardize, block, compare and score; start with exact identifiers
- Set thresholds for auto-link, review and reject from a labelled sample
- Keep a crosswalk with confidence and link dates; never overwrite sources
- Support un-merging when a link is wrong
- Propagate canonical IDs into data products, retrieval metadata and tools
- Respect consent and privacy rules when linking personal data
- Monitor match rates and review queues as new data arrives
Common mistakes
Over-merging is worse than under-merging: linking two different customers can expose one person's data to another and lead agents to act on the wrong account. Tune for precision first. Other mistakes include matching on names alone, treating parent and subsidiary as duplicates, running resolution once instead of continuously, and resolving identities in the warehouse while the AI still queries source systems by name. Our guides to data quality for AI and CRM automation cover related clean-up work.
Connecting customer and product data for AI?
ZSpace Labs integrates CRM, ERP, ecommerce and support systems so AI assistants and agents see one consistent view. See AI automation and Shopify development.
Conclusion
Entity resolution is quiet infrastructure with a large effect on AI. It decides whether an assistant sees one customer or four fragments, and whether an agent acts on the whole picture. Define what 'same' means, match with rules and scores, review the uncertain middle, keep a reversible crosswalk and carry canonical IDs into everything AI reads.
Common questions.
Entity resolution is the process of working out which records, across one or many systems, refer to the same real-world thing, such as the same customer, product or company, and linking them under one identity.