AI Shopping Assistant: How to Build One for Your Online Store
How to build an on-site AI shopping assistant: scope, architecture, catalog retrieval, tools, guardrails, handoff, accessibility, evaluation, costs and measurement.
Quick answer
An AI shopping assistant helps shoppers find and choose products through conversation on your store. Build it on four layers: inputs (the message, page context, cart and consent state), retrieval over accurate catalog, stock, policy and fit data, a language model that can call tools such as search and add to cart (with confirmation), and guardrails including scope limits, no invented facts, human handoff and logging. Start with a narrow scope, clean your data first, evaluate on real questions and measure against a holdout.
Assistant, Agent or Chatbot?
Terms blur, so define the scope. A rules-based chatbot follows scripted flows. An AI shopping assistant uses a language model to understand open questions and guide product choice on your store, grounded in your data. An AI agent is a system that plans and takes actions with tools towards a goal, with varying autonomy (AI agents for ecommerce). External AI shopping agents act for consumers across stores (AI shopping agents).
This article covers the on-site assistant. It may use agent-like tool calling (searching the catalog, checking stock, adding to cart), but it acts within your store, for the shopper in front of it, with the shopper confirming actions.
Define the Scope First
Assistants that try to do everything do most things poorly. Start with the jobs your shoppers most need help with, based on support questions, search logs and research. Typical first scopes are product finding for a need, comparisons within a category, sizing or compatibility, and pre-purchase policy questions. Leave order changes, refunds and account issues to support flows with proper identity checks.
| Job | Needs | Include in first release? |
|---|---|---|
| Find products for a need | Search, attributes, stock | Yes |
| Compare products | Structured specs | Yes, within categories |
| Size and fit help | Size charts, fit notes, reviews data | Yes, if data is good |
| Delivery and returns questions | Policies, delivery rules | Yes |
| Add to cart | Cart API, confirmation | Yes, with confirmation |
| Order status and changes | Order systems, identity checks | Often later, via support |
| Refunds and complaints | People and policy judgement | Hand off |
Architecture
Most assistants share a similar architecture. The interface collects the message and page context (which product or category the shopper is viewing). The orchestration layer decides what to retrieve and which tools to call. Retrieval fetches relevant products, policies and help content. The language model writes the answer using retrieved content and tool results. Guardrails check inputs and outputs. Logs record conversations for review, with personal data handled carefully.
| Component | Role | Notes |
|---|---|---|
| Chat interface | Input, output, product cards | Accessible, fast, mobile-friendly |
| Context | Page, cart, market, consent | Only what's needed |
| Retrieval | Catalog, stock, prices, policies, help | Live data for price and stock |
| Tools | Search, filter, get product, add to cart | Explicit, limited permissions |
| Language model | Understanding and writing | Choose for quality, latency, cost |
| Guardrails | Scope, safety, no invented facts | Input and output checks |
| Logging and review | Quality monitoring | Retention and masking |
| Handoff | Route to people | With conversation history |
Retrieval: The Assistant Knows What You Give It
The assistant should answer from your data, not from the model's general knowledge. Product questions need structured attributes (materials, dimensions, compatibility, care), not just marketing copy. Sizing help needs size charts and fit notes. Policy questions need current, consistent policy text. Prices and stock should be fetched live through tools rather than stored in an index that can go stale.
Poor data is the most common reason assistants disappoint. Audit product data and help content before building, and plan to keep them current. See product data for AI search.
Planning an AI shopping assistant?
ZSpace builds on-site assistants grounded in your catalog, with guardrails, handoff and measurement from day one.
Tools and Permissions
Tool calling lets the model take defined actions: search the catalog with filters, fetch product details, check stock for a size, add an item to the cart. Define each tool narrowly, validate its inputs, and limit what it can do. The assistant should never complete a purchase or payment; adding to cart should require the shopper's confirmation, and checkout remains the normal flow.
search_products(query: string, filters?: {category, size, colour, price_max}) -> ProductSummary[]
get_product(id: string) -> ProductDetail # live price and stock
check_size_stock(id: string, size: string) -> {in_stock: boolean}
add_to_cart(variant_id: string, quantity: 1..5) -> CartLine # only after shopper confirms
get_policy(topic: "delivery" | "returns" | "warranty") -> PolicyTextGuardrails
Guardrails keep the assistant useful and safe. Scope limits keep it on shopping topics. Instructions and checks prevent invented facts: if information isn't in retrieved data, it says so and offers alternatives. Output checks catch prices or claims that don't match data. Input handling addresses prompt injection and abuse. Sensitive topics (medical claims, safety issues, complaints) route to people. The OWASP Top 10 for LLM Applications is a useful checklist of risks to design against (OWASP GenAI Security Project).
- Stays on shopping and store topics
- Answers only from retrieved data; says when it doesn't know
- Prices and stock from live tools
- No medical, legal or safety claims beyond approved content
- Prompt injection and abuse handling
- Handoff on low confidence, frustration or sensitive topics
- Conversations logged with personal data masked
Experience Design
Make the assistant easy to find but not intrusive: a clear entry point, suggested prompts relevant to the page, short answers with product cards, and links to product pages. Show what it used ("based on the size chart") where helpful. Let shoppers continue browsing while the conversation stays available. On mobile, avoid covering the add-to-cart button or content.
Accessibility is essential: keyboard operation, focus management, labels, announcements of new messages to screen readers, contrast and resizable text. See ecommerce accessibility.
Evaluation Before Launch
Build a test set of real shopper questions from support logs, search logs and research, with expected answers or acceptable products. Run the assistant against it and score accuracy, relevance, policy correctness and tone. Include tricky cases: products you don't sell, ambiguous sizes, questions outside scope and attempts to make it misbehave. Repeat evaluation after every significant change to prompts, data or model.
| Test case type | What to check |
|---|---|
| Common product questions | Accurate, grounded answers |
| Needs-based requests | Relevant products, in stock |
| Comparisons | Correct specs, fair comparison |
| Policy questions | Matches current policy exactly |
| Out-of-range requests | Honest "we don't stock that" with alternatives |
| Out-of-scope and adversarial | Declines politely; no data leakage |
Build or Buy
Buying a vendor assistant or platform app is faster; building gives more control. When evaluating vendors, check how they ground answers in your data, whether you can control scope and tone, how handoff works, what analytics they provide, how they handle personal data and model providers, and whether the interface is accessible. When building, budget for data preparation, evaluation and ongoing review, not only the initial integration.
Costs
Ongoing costs include model usage (which scales with conversations and message length), hosting, vendor fees, data maintenance and the time to review conversations and improve answers. Estimate from expected conversation volume. Use smaller models or rules for simple tasks and reserve larger models for complex questions to manage cost and latency.
Measuring Impact
Shoppers who use an assistant are often more engaged, so their conversion rates overstate its effect. Measure with a holdout: randomly withhold the assistant from some visitors and compare conversion, revenue per visitor, returns and support contacts. Track quality through reviewed samples, satisfaction ratings and handoff reasons. See personalization testing.
A Phased Build Plan
| Phase | Scope | Exit criteria |
|---|---|---|
| 1. Data readiness | Product attributes, policies, size data, help content | Test questions answerable from data |
| 2. Prototype | Retrieval and answers for one category | Accuracy on test set above agreed bar |
| 3. Limited launch | Small traffic share, one or two categories | Quality reviews, holdout comparison |
| 4. Expand | More categories, add to cart, more pages | Stable accuracy, positive holdout result |
| 5. Operate | Ongoing review, content updates, re-evaluation | Owner and cadence in place |
Assistants and Agentic Commerce
The data that powers an on-site assistant (structured attributes, accurate policies, live stock) is the same data external AI agents and commerce protocols need to represent your products elsewhere. Investing in it serves both. The external channels are emerging and vary by platform and market, so treat them as an extension of good product data rather than a replacement for your own assistant or store experience. See agentic commerce.
Handling Sensitive Categories
Some products need extra care: health and supplements, beauty products with claims, children's products, age-restricted goods, and anything with safety implications. Limit the assistant to approved product information, avoid medical or safety advice, route questions to qualified staff where appropriate, and follow advertising and product claims rules for your markets. Test these categories specifically before launch.
Common Mistakes
- Launching before product data and policies are clean
- Letting the model answer from general knowledge
- Broad scope on day one
- No handoff to people
- Inaccessible chat widget
- Measuring usage instead of incremental impact
Ready to build a shopping assistant?
Talk to ZSpace about AI assistant and agent development, catalog and cart integration and assistant UX.
Conclusion
A useful AI shopping assistant is narrow at first, grounded in good data, limited in what it can do, transparent, accessible and measured against a holdout. Related: conversational ecommerce and AI product discovery.
Common questions
An on-site or in-app assistant that helps shoppers find and choose products through conversation, using a language model grounded in the store's catalog, stock, policies and help content.