Skip to content
UI/UX6 min read

AI Assistant App UX: Designing Interfaces That Run Inside ChatGPT, Claude and Other Assistants

How to design apps that render inside AI assistants with MCP Apps: when to show UI, inline vs fullscreen, state, actions and the limits of a sandbox.

01

Quick answer

An AI assistant app is your product's functionality made available inside an AI assistant. The assistant calls your tools through the Model Context Protocol and, where the host supports it, renders interactive components you provide: a product carousel, a booking form, a map, an approval card. MCP Apps, an official MCP extension announced in January 2026, standardizes how a server declares these interfaces as ui:// resources and how hosts render them in sandboxed iframes (MCP blog).

Designing them is different from designing a website. Your UI is a guest inside someone else's conversation. Show the smallest interface that helps the user decide or act, let the model handle explanation, and route every consequential action through your normal backend controls.

02

How assistant apps work

The assistant (the host) connects to your MCP server, reads your tools and their metadata, and decides when to call them. A tool can point to a UI resource. When the tool runs, the host fetches that resource and renders it in a sandboxed iframe next to the conversation. The UI and the host exchange messages over MCP's JSON-RPC, so the UI can request further tool calls (subject to host rules and user consent) and update the model's context. OpenAI's Apps SDK, which helped shape the standard, documents display modes such as inline cards, inline carousels, fullscreen and picture-in-picture (OpenAI Apps SDK UI guidelines).

AI assistant app architecture (diagram)
User ◀──▶ AI assistant (host: chat UI + model)
                │  calls tool: search_rooms(city, dates)
                ▼
          Your MCP server ──▶ your APIs (inventory, pricing,
                │               booking, auth, audit)
                │ tool result + _meta: ui resource
                ▼
   Host renders ui://rooms/results in a sandboxed iframe
                │  user picks a room
                ▼
   UI requests tool call: hold_room(id)  → host may ask consent
                │
                ▼
   Model continues the conversation with the new state
03

When to show UI, and when text is enough

Text is the default in an assistant. UI earns its place when it makes a decision or action easier than words would.

Show UI when...Let text handle it when...
People compare several options with images, prices or attributesThere is one clear answer
The user needs to pick, edit or fill structured fieldsThe model can ask one clarifying question
Spatial or visual content matters (maps, seats, media)The result is an explanation or summary
A consequential action needs a clear review and confirmationThe action is trivial and reversible
Status of an ongoing task should stay visibleThe task finishes immediately
04

Choosing a display mode

Start with the smallest presentation that lets someone understand the result or complete the task, and escalate only when needed. An inline card suits one decision or a small amount of structured data. An inline carousel suits scanning a few similar options. Fullscreen suits tasks that need room: a map, an editor, detailed browsing. Picture-in-picture suits an ongoing activity. Avoid nested scrolling and long multi-tab interfaces inline; offer an inline summary with a way to expand.

05

Design principles for guest UI

  • Be conversational-first: the model explains; your UI shows and lets people act
  • Do one job per view: one decision, one form, one comparison
  • Follow the host's look: respect host theming and spacing so the UI feels native; keep your brand to content and small accents
  • Keep state in your backend: the iframe can be reloaded or discarded; carts, holds and drafts must survive
  • Make actions explicit: label buttons with outcomes ('Book room for 2 nights, $340'), not 'Submit'
  • Confirm consequential actions: show what will happen and its cost before it happens; see AI action confirmation UX
  • Degrade gracefully: every tool must also work without UI, in hosts that do not render it
  • Design for small widths and both themes
06

State, context and the model

Three parties hold state: your backend (authoritative), the UI (what the user sees now) and the model (what it believes happened). Keep them aligned. When the user acts in the UI, send the change to your backend first, then report the new state to the host so the model's next message reflects it. Never let the model assume an action succeeded because a button was shown. Tool results should carry identifiers and timestamps so the model and the UI refer to the same objects.

07

Security and trust

MCP Apps run in sandboxed iframes with restricted permissions and declared content-security rules, and hosts can require user consent before UI-initiated tool calls. That protects the host, not your business logic. Enforce authorization, limits and validation on your server for every tool call, authenticate users with delegated OAuth where accounts are involved, and never put secrets in UI resources. See APIs and MCP servers for AI agents and MCP security.

08

Example: a hypothetical booking app

A hotel group exposes search_rooms, hold_room and book_room. 'Find me a quiet room in Lisbon for two nights next week' calls search_rooms, which returns an inline carousel of three rooms with photos, price and cancellation terms. The user taps one; the UI calls hold_room and shows a confirmation card with total price, dates and policy. Booking happens only after an explicit 'Book for $340' tap, which calls book_room on the server with idempotency and payment checks. In a host without UI support, the same tools work through text: the model lists the rooms and asks for confirmation in words.

Bringing your product into AI assistants?

ZSpace Labs designs and builds MCP servers and assistant app interfaces on top of existing APIs. See AI automation and UI/UX design.

Start a Project
09

Common mistakes

Porting a whole website into an iframe. Showing UI for every tool call. Keeping important state only in the iframe. Relying on the host's consent prompt instead of server-side authorization. Fighting the host's styling. And forgetting that the model can describe your results in its own words, so tool outputs must be accurate and unambiguous.

10

Conclusion

Assistant apps put your product where people already are, but on the assistant's terms. Design small, purposeful views that help people decide and act, keep state and rules on your server, confirm consequential actions clearly and make every tool useful without UI. It is the most practical form of generative UI many businesses will ship first. For the wider pattern set, see AI interface patterns.

FAQ

Common questions.

It is a service you make available inside an AI assistant such as ChatGPT or Claude. The assistant calls your tools through the Model Context Protocol and, where supported, renders interactive interface components you provide alongside the conversation.

Get in touch

Have a project in mind?

Whether you're building a new digital product, improving an existing website, or looking to automate part of your business — let's talk.