How Much Autonomy Should You Give an AI Agent? A Five-Level Framework
A five-level framework for AI agent autonomy, from suggesting to acting alone, with the conditions, controls and evidence each level requires.
Quick answer
Give an AI agent the lowest level of autonomy that still delivers value, then raise it per case type as evidence accumulates. A practical scale has five levels: 1. Suggest (agent recommends, a person acts), 2. Draft (agent prepares the work, a person sends or applies it), 3. Execute with approval (agent acts after explicit sign-off), 4. Execute within boundaries (agent acts alone inside hard limits on amount, scope and reversibility), 5. Autonomous (agent acts under monitoring with sampled review). Choose the level from three factors: how often the agent is right on that case type, how costly and reversible a mistake is, and how quickly someone would notice.
Autonomy is a design decision, not a capability
It is tempting to set autonomy by what the model can do: if it handles a refund correctly in a demo, let it issue refunds. Feng, McDonald and Zhang's 2025 paper *Levels of Autonomy for AI Agents* makes the opposite argument: autonomy should be chosen deliberately, separately from capability and environment, by deciding what role the user plays (operator, collaborator, consultant, approver or observer). That framing is useful for businesses because it turns a vague debate into a configurable decision per task.
The five levels
| Level | What the agent does | What the person does | Typical use |
|---|---|---|---|
| 1. Suggest | Recommends an action or answer | Decides and acts | Research, triage hints, next-best-action |
| 2. Draft | Prepares the email, record, order or change | Reviews, edits and sends | Customer replies, quotes, reports |
| 3. Execute with approval | Acts after explicit sign-off on a preview | Approves or rejects each action | Refunds above a limit, vendor payments, data changes |
| 4. Execute within boundaries | Acts alone inside hard limits | Sets limits; reviews samples and exceptions | Small refunds, rescheduling, tagging, routing |
| 5. Autonomous | Plans and acts end to end | Monitors outcomes; handles escalations | High-volume, low-risk, well-measured tasks |
Key takeaway
Most valuable business agents run at levels 2 to 4. Level 5 is rarely justified for anything that touches money, customers or records unless the actions are trivially reversible.
Choosing the level: three questions
Score each case type the agent handles, not the agent as a whole.
| Question | Points to lower autonomy | Points to higher autonomy |
|---|---|---|
| How often is the agent right on this case type? | Below your target on a representative test set | Consistently at or above target over weeks of real cases |
| What does a mistake cost, and can it be undone? | Money, legal commitments, customer harm, irreversible | Small, reversible, internal |
| How quickly would someone notice an error? | Days later, or only if a customer complains | Immediately, through validation or monitoring |
What each level requires
Autonomy is earned with controls. Each step up removes a human check, so something else has to take its place.
| Level | Minimum controls |
|---|---|
| 1–2 | Grounded sources, logging, easy feedback from the reviewer |
| 3 | Clear previews of what will happen, approval records, idempotent actions |
| 4 | Hard limits enforced in code (amounts, quantities, scope), validation after each action, sampled review, alerts |
| 5 | All of the above plus outcome monitoring, cost budgets, a tested kill switch and a regular evaluation cycle |
Deciding how far to trust an agent?
ZSpace Labs designs agents with autonomy set per case type, the limits and approvals that go with it, and the measurements needed to raise it safely. See AI automation services.
Mixed autonomy inside one agent
A customer service agent illustrates why autonomy should be set per action, not per agent (an illustrative design, not a client case):
- Answer order status and delivery questions: level 5, grounded in order data
- Reschedule a delivery: level 4, within carrier rules
- Refund under a small threshold: level 4, once per order, logged
- Refund above the threshold or outside policy: level 3, preview sent for approval
- Respond to a complaint mentioning legal action: level 1, route to a person with a summary
Raising and lowering autonomy safely
Move up one level at a time, for one case type at a time, after a defined period meeting targets. Run the higher level in shadow mode first (the agent decides, a person still acts) and compare decisions. Move down immediately when accuracy drops, inputs change, a model is updated, or a security concern appears; this should be a configuration flag, not a code change. Record each change and the evidence behind it, which also supports accountability; see who is responsible when an AI agent makes a mistake.
Common mistakes
- Setting one autonomy level for the whole agent
- Granting autonomy because a demo worked, without a test set
- Approval steps that show too little to judge, so reviewers approve everything (see human-in-the-loop AI)
- Limits written in the prompt instead of enforced in code (see AI agent guardrails)
- No way to lower autonomy quickly when something changes
Conclusion
Autonomy is the main dial in any agent deployment. Set it per case type, start low, raise it with evidence and lower it without hesitation. The goal is not maximum independence but the most useful level the evidence supports. For how agents can operate software directly, and the autonomy questions that raises, see computer-use agents.
Common questions.
As much as the evidence supports for that specific task, and no more. Start at a level where people review the agent's output, measure its accuracy on real cases, and raise autonomy only for case types where errors are rare, reversible and cheap.