Context Engineering: What Information an AI Agent Needs to Work Reliably
What context engineering is, what information an AI agent needs at each step, why more context can hurt, and how to supply it reliably.
Quick answer
Context engineering is deciding what information an AI agent sees at each step. An agent needs five things: instructions and limits, the current task and who it is for, the specific facts required for the next decision (fetched from systems of record, not pasted in bulk), clear descriptions of its tools, and a compact record of what has happened so far. More is not better: models attend less reliably as context grows, so the goal is the smallest set of high-signal information. Most agent failures that look like reasoning errors are actually missing, stale or noisy context.
From prompt engineering to context engineering
Prompt engineering focused on wording a single instruction well. Agents changed the problem: they run for many steps, call tools, read documents and accumulate history, so what the model sees at step twelve is mostly not the prompt you wrote. Anthropic's engineering team described this shift in September 2025 as context engineering: curating and maintaining the optimal set of information during inference, including everything that lands in the context window beyond the prompt.
The same article names the core constraint. Models have a finite attention budget; as the number of tokens grows, the ability to recall and use specific information declines, an effect Anthropic calls context rot. Large context windows help, but they do not remove the need to choose.
The five kinds of context an agent needs
| Context | What it contains | Where it should come from |
|---|---|---|
| Instructions | Role, goal, constraints, escalation rules, tone | System prompt; short and specific |
| Task and user | The request, who is asking, their permissions and account | Application state, identity system |
| Facts for the next decision | Order status, contract terms, product specs, policy clauses | Tools and retrieval over systems of record, fetched just in time |
| Tools | What actions exist, when to use each, what they return | Tool names, descriptions and schemas |
| History | What has been tried, decided and learned in this task | Compacted summaries and structured notes, not full transcripts |
Fetch facts just in time instead of pasting them in
The most common mistake is loading everything up front: the whole policy manual, the full customer record, every past ticket. It is expensive, it buries the relevant line, and it is out of date by the time it is used. Better: give the agent tools to look up what it needs when it needs it ("get_refund_policy(region)", "get_order(order_id)"), and make those tools return compact, relevant results. This is why tool design and context engineering are the same discipline; see AI agent tool design.
Key takeaway
The quality of an agent is constrained by the quality and accessibility of the systems it can actually query. If the facts are not reachable through a tool or retrieval, no prompt will supply them.
Capture the knowledge people never wrote down
Every process runs on unwritten rules: which customers get flexibility, which supplier formats are unreliable, what "urgent" means for a particular account. People apply them without thinking; agents cannot. Before building, interview the people who do the work and review cases where they made non-obvious decisions. Then decide where each rule belongs: a short instruction if it applies everywhere, a retrievable document if it applies sometimes, a field in a system if it is per customer, or a deterministic check if it must always hold.
Managing long tasks
Agents working for many steps accumulate tool outputs and reasoning that crowd out what matters. Three techniques keep context useful:
- Compaction: periodically summarize progress and decisions, then continue from the summary instead of the full transcript
- Structured notes: have the agent keep a small working file (findings, open questions, next steps) that survives compaction
- Sub-agents: delegate focused sub-tasks to separate agents with clean context, returning only their conclusions
- Trim tool output: return summaries, top results and identifiers, with a way to fetch detail on request
Memory across sessions
Some information should persist beyond one task: a customer's preferences, a project's conventions, lessons from previous cases. Store it deliberately, scoped per user or tenant, with its source recorded, and retrieve it when relevant rather than injecting all of it every time. Poorly governed memory is also a security risk, since poisoned entries can steer later runs. See AI agent memory and the memory risk in the OWASP agentic Top 10.
Agents getting answers wrong in production?
ZSpace Labs diagnoses agent failures from transcripts, then fixes the context: retrieval, tools, instructions and memory. See AI automation services.
Diagnosing context problems
When an agent makes a bad decision, read the exact context it had at that step and ask four questions.
| Question | If yes | Fix |
|---|---|---|
| Was a needed fact missing? | Missing context | Add a tool or retrieval source; capture the unwritten rule |
| Was the fact present but outdated? | Stale context | Read from the system of record; add freshness checks |
| Was it present but buried? | Noisy context | Trim tool outputs, compact history, narrow retrieval |
| Were there conflicting instructions or facts? | Conflicting context | Resolve the source of truth; simplify instructions |
Context Engineering vs Prompt Engineering
Prompt engineering shapes how the model should behave: role, instructions, tone, examples. Context engineering designs the whole information environment the model works in at each step: which facts, tools, history and memories are present, in what form and when. Prompts are one part of context. For single-turn features, a well-written prompt may be enough; for agents that run many steps and call tools, context design decides most of the quality.
| Prompt engineering | Context engineering | |
|---|---|---|
| Focus | Instructions and behaviour | Information available at each step |
| Scope | System and user prompts | Prompts, retrieval, tool definitions and results, memory, history |
| Changes during a task? | Mostly static | Changes every step |
| Typical failure | Vague or conflicting instructions | Missing, stale, buried or conflicting facts |
| Main tools | Wording, examples, output format | Retrieval, tool design, compaction, memory policies |
| Matters most for | Single-turn generation and classification | Agents and long-running, tool-using tasks |
Conclusion
Reliable agents are mostly well-fed agents. Decide what each step needs, fetch facts just in time from systems of record, keep instructions short, return compact tool results, compact long histories and capture the rules people never wrote down. For the data preparation underneath, see AI data readiness; for retrieval techniques, retrieval-augmented generation.
Common questions.
The practice of deciding what information goes into a model's context window at each step (instructions, task details, retrieved data, tool results, history) and keeping it relevant as a task progresses. Anthropic describes it as the natural progression of prompt engineering for agents.