Skip to content
AI & Automation

AI Application Threat Modeling: How to Identify Risks Before Deployment

A structured method for threat modelling AI applications: describing the system, assets, actors, data flows and trust boundaries, AI-specific attack surfaces, abuse cases, mitigations and residual risk, with a worked template.

Quick answer

Threat model an AI application by first drawing the system: components, data flows, assets and actors. Mark trust boundaries, treating user input, retrieved documents, web content, tool outputs and model providers as untrusted or external. For each boundary and component, list threats using STRIDE plus AI references such as the OWASP Top 10 for LLM Applications, write concrete abuse cases, choose layered mitigations with owners, and record residual risk for an accountable person to accept. Revisit whenever tools, data, models or autonomy change.

Where This Fits

Threat modelling comes before red teaming and security testing, which verify that mitigations work. An overview of AI threats is in AI security for business applications, and governance of residual risk in AI governance framework.

The Method

Trust boundaries are where most AI threats concentrate, so they deserve the most time.

Step 1: Describe the System

Draw a data flow diagram covering users, front ends, backend services, the model provider or self-hosted model, retrieval stores, tools and the systems they call, logs and traces, and administrators. Note what data moves along each flow and in which direction. Many risks become obvious once the diagram shows, for example, that an assistant reading customer emails also has a tool that sends email.

Step 2: Assets and Actors

AssetsActors
Customer and employee personal dataLegitimate users making mistakes
Confidential documents in retrievalMalicious users and fraudsters
Ability to take actions (refunds, emails, changes)Authors of content the AI reads (web, email, documents)
Credentials and API keysCompromised third-party tools or providers
Model and prompt configurationInsiders with excessive access
Budget and availabilityAutomated abuse such as scraping and cost attacks

Step 3: Trust Boundaries

Mark every point where data crosses from less trusted to more trusted context. In AI applications the critical boundaries are: user input entering prompts; retrieved documents, web pages, emails and tool outputs entering the model's context; model output turning into tool calls, rendered HTML or database queries; and data leaving to model providers and third-party tools. Treat everything crossing into the model's context from outside as potentially containing instructions, and everything leaving the model as untrusted output.

Designing an AI feature with tools or sensitive data?

ZSpace Labs runs threat modelling workshops for AI applications and turns findings into concrete controls. See AI development services.

Start a Project

Step 4: Threats and Abuse Cases

Use STRIDE (spoofing, tampering, repudiation, information disclosure, denial of service, elevation of privilege) for the conventional view, then walk through AI-specific risks from the OWASP Top 10 for LLM Applications and tactics in MITRE ATLAS. Turn them into concrete abuse cases written from the attacker's perspective, such as: 'As a customer, I paste instructions into the chat so the assistant approves a refund above policy' or 'As an outsider, I email instructions that make the inbox assistant forward a summary to me.'

Step 5: Mitigations and Residual Risk

For each abuse case, choose layered mitigations: prevention (least privilege, validation outside the model, isolation of untrusted content), detection (monitoring tool calls, anomaly alerts), response (kill switches, rollback) and limits (budgets, approval thresholds). Prompt instructions can be one layer, never the only one. Assign owners and target dates, then record residual risk and have an accountable owner accept it explicitly.

Example: threat model entry (illustrative)
id: TM-07
component: support assistant -> refund tool
boundary: user chat input -> model context -> tool call
abuse_case: customer instructs assistant to issue refund above policy
impact: financial loss; likelihood: medium
mitigations:
  - refund tool enforces policy limits server-side (prevent)
  - refunds above 50 require agent approval (limit)
  - alert on refund rate per account (detect)
  - regression test TM-07 in red team suite (verify)
residual_risk: low; accepted_by: head of support ops; review: 2027-01

Step 6: Verify

A threat model is a hypothesis about where the system is weak. Verify mitigations with targeted tests, red team exercises and automated regression cases tied to threat IDs. Update the model when tests reveal unexpected paths.

Common AI Threats to Consider

  • Direct and indirect prompt injection
  • Sensitive information disclosure through retrieval, outputs or logs
  • Excessive agency: tools or permissions beyond the task
  • Improper output handling: model output executed or rendered unsafely
  • Supply chain compromise of models, packages or tools
  • Data and model poisoning through writable sources
  • Unbounded consumption: cost and denial-of-service abuse
  • Misinformation and over-reliance on wrong outputs

Advantages and Limitations

Threat modelling finds design-level risks when they are cheapest to fix, and it gives security, product and engineering a shared picture. It depends on participants' knowledge and imagination, and it ages as the system changes. Keep it lightweight, tied to the architecture diagram and updated on every significant change.

Threat Modelling Agents and Multi-Agent Systems

Agents add threats around autonomy: goals being hijacked by injected content, tools combined in unintended ways, permissions inherited across agents, runaway loops and actions taken without the user's awareness. In multi-agent systems, messages between agents are another trust boundary; one compromised or confused agent can pass harmful instructions to others. Model each agent's identity, tools, data access and the channels between agents explicitly. See AI agent architecture and agent-to-agent communication.

Keeping the Threat Model Current

Tie threat model updates to change triggers: adding a tool, connecting a new data source, changing the model or provider, increasing autonomy or exposing the feature to new users. Keep the model next to the architecture diagram in the repository, review it in design reviews for these changes and link each threat to its tests. A short, current threat model is far more useful than a long one written once.

Running a Threat Modelling Session

A practical session for one AI feature takes two to three hours. Before it, the engineering lead prepares the data flow diagram and a list of tools, data sources and providers. In the session, walk the diagram from user input to final action, stopping at each trust boundary to ask what could enter, what could leave and what the model could be persuaded to do. Capture abuse cases on the diagram itself. Spend the last part agreeing mitigations, owners and which threats need tests.

Keep the group small but varied: the engineers who build it, a security specialist, the product owner and someone who knows the users and data. Afterwards, write up the results in the repository next to the architecture, and schedule a short review when the next significant change lands. Testing that follows is covered in AI red teaming.

Worked Example

An illustrative scenario, not a client case: a property management company designs a tenant assistant that can read lease documents and log maintenance requests. The threat model shows the assistant could retrieve other tenants' leases because documents are indexed by building, not by tenant. The design changes to per-tenant document filters before any code is written, and a canary test is added to the release suite.

Common Mistakes

Most weak threat models share the same gaps.

  • Modelling only the chat input, not retrieved content and tools
  • Relying on prompt instructions as the main mitigation
  • No owner or acceptance for residual risk
  • Threat models written once and never updated
  • No tests linked to the identified threats

Want a threat model before you build?

Talk to ZSpace Labs about an AI threat modelling session for your planned AI feature or agent.

Start a Project

Conclusion

Threat modelling makes AI risks visible while they are still design decisions. Draw the system, mark trust boundaries, write abuse cases, layer mitigations, accept residual risk explicitly and verify with tests.

FAQ

Common questions

A structured analysis of how an AI system could be attacked or misused, done before and during development: describing the system and its data flows, identifying assets, actors and trust boundaries, listing threats and abuse cases, choosing mitigations and accepting or reducing residual risk.

Related services
Relevant industries
Get in touch

Have a project in mind?

Whether you're building a new digital product, improving an existing website, or looking to automate part of your business — let's talk.

Keep exploring
AI & Automation
6 min read

AI Security Testing: A Practical Checklist for Testing AI Systems

A practical, defensive checklist for testing AI systems: authentication and authorization, prompt injection, data leakage, tool execution, output handling, logging and privacy, dependencies, resource limits and incident response readiness.

Read article
AI & Automation
8 min read

AI Red Teaming: How to Test AI Applications for Security Risks

A defensive guide to red teaming AI applications: scoping, threat discovery, test scenarios for prompt injection, tool misuse and data exposure, automated and manual testing, evaluation, remediation and repeat testing.

Read article
AI & Automation
7 min read

AI Security for Business Applications: How to Protect AI Systems

How to secure AI applications: threat model, prompt injection, tool permissions, data exposure, secrets, authorization, model and supply chain risks, runtime monitoring and AI incident response.

Read article