Skip to content
AI & Automation7 min read

Securing AI Coding Agents: Sandboxes, Permissions, Secrets and Prompt Injection

How to use Claude Code, Codex, Cursor and Copilot safely: the threat model, sandboxing, permissions, secrets, MCP servers and a team policy.

01

Quick answer

A coding agent combines a model that can be tricked with hands that can act: it reads files, runs commands, installs packages, calls tools and pushes code. Secure it by containing those hands. Run agents in an OS-level sandbox or isolated container that restricts file and network access; use permission rules so risky commands need approval; keep production credentials and real customer data out of reach; treat issues, web pages, dependencies and tool outputs as untrusted input; vet MCP servers; and require every change to land on a branch and pass review and CI before reaching the main branch. Write this down as a team policy that differs by environment.

02

Why coding agents are a different security problem

An autocomplete suggestion can only be accepted or rejected by a developer. An agent acts. Modern coding agents edit many files, run test suites and shell commands, install dependencies, browse documentation, call MCP tools and open pull requests, often for minutes at a time without a person watching each step. That autonomy is the point, and it is also the attack surface.

The core risk is the combination the security community calls the lethal trifecta for agents: access to private data, exposure to untrusted content, and a way to send data out or take actions. A coding agent on a developer laptop often has all three: source code and local credentials, a repository full of third-party files, and a terminal with network access.

03

The threat model

ThreatHow it happensExample impact
Prompt injection via contentInstructions hidden in files, issues, PR comments, docs, web pages or package READMEsAgent runs a command or edits config that the developer did not ask for
Secret exfiltrationAgent reads .env files, credentials or tokens and sends them out via a command, URL or commitLeaked API keys or cloud credentials
Malicious dependenciesAgent installs a hallucinated or lookalike packageAttacker code running on developer machines or in CI
Configuration tamperingAgent edits its own settings, CI files or editor config to reduce safeguardsSafeguards silently disabled
Over-broad tool accessMCP servers or CLIs with write access to production systemsChanges to live data or infrastructure
Unreviewed changesAgent pushes directly to main or auto-mergesVulnerable or broken code in production
04

This has already happened

These are not theoretical. In August 2025, Microsoft disclosed CVE-2025-53773 in GitHub Copilot and Visual Studio: researchers showed that instructions injected into content the agent processed could get it to modify workspace settings to auto-approve its own actions and then execute commands. It was patched, and similar issues have been reported and fixed in other agents and editors. The pattern is consistent: untrusted content plus the ability to change configuration or run commands equals code execution.

The OWASP Top 10 for Agentic Applications captures the same risks under Agent Goal Hijack (ASI01), Unexpected Code Execution (ASI05) and Agentic Supply Chain Vulnerabilities (ASI04). Our OWASP agentic guide explains each.

Key takeaway

Assume any file, issue or web page an agent reads could contain instructions written by someone else. Design the environment so that following them cannot cause serious harm.

05

Control 1: Sandbox the agent

Sandboxing is the most effective single control because it limits what an agent can do even when it is manipulated. The major tools now include it. Claude Code's sandbox uses operating-system mechanisms on macOS, Linux and WSL2 to restrict which files and network hosts shell commands can reach, which lets it run sandboxed commands without asking for approval each time. Codex documents sandboxing and approval modes for its CLI and IDE extension. Cursor's cloud agents run in isolated virtual machines with controls for secrets and allowed domains.

Check exactly what the sandbox covers. Claude Code's documentation, for example, notes that its sandbox applies to shell commands, while its file tools, MCP servers and hooks run outside it, governed by permission rules instead. Where tools offer no sandbox, run the agent inside a dev container or VM with no access to your home directory and a network allowlist. Security controls differ between products, which is one of the main criteria in our Claude Code vs Codex vs Cursor comparison.

  • File writes limited to the project directory
  • Network access limited to package registries and approved hosts
  • No access to SSH keys, cloud credential files or browser profiles
  • Sandbox enabled by default in team settings, not left to each developer
06

Control 2: Permission rules and approvals

Use the agent's permission system to separate safe actions from risky ones. Allow read-only operations, running tests and linting freely; require approval for installing packages, network calls outside the allowlist, editing CI or agent configuration, and anything touching git remotes. Deny outright commands that should never run, such as deploying to production or reading credential directories.

Manage these settings centrally where the tool supports it (managed or organization-level settings) so individual developers cannot accidentally weaken them, and protect the settings files themselves from edits by the agent.

07

Control 3: Keep secrets out of reach

Anything an agent can read can end up in a prompt, a log, a commit or an outbound request. Keep production credentials off developer machines used with agents, use development credentials with minimal scope and short lifetimes, load secrets from a manager at runtime rather than `.env` files in the repository, and add secret scanning with push protection. If you suspect a key was exposed to an agent session, rotate it.

08

Control 4: Treat inputs as untrusted

Be deliberate about what you point agents at. Asking an agent to "fix the issue in this GitHub link" or "follow the setup in this README" feeds it third-party text. For public repositories and issues from unknown users, run agents in the most restricted mode, and do not combine untrusted input with access to secrets or the ability to push. Review new dependencies the agent proposes before installing; hallucinated package names are a known attack route.

For background on the attack technique, see indirect prompt injection and prompt injection prevention.

09

Control 5: Vet MCP servers and integrations

MCP servers give agents new tools (databases, ticketing, cloud consoles, browsers), and each one adds both capability and attack surface. Use official or well-maintained servers, pin versions, read the tool list and descriptions, grant read-only access unless writes are needed, and never connect an agent to production databases or infrastructure through MCP without strong approvals. See MCP security.

10

Control 6: Review before merge, always

Agents should work on branches and open pull requests, never push to protected branches. Branch protection, required CI checks (tests, static analysis, dependency and secret scanning) and human review for sensitive areas ensure that a manipulated or simply wrong change does not ship. AI review tools are a useful extra reader; see AI code review and AI-generated code security.

Rolling out coding agents across a team?

ZSpace Labs helps teams set up coding agents with sandboxing, managed permissions, secrets handling and review workflows that keep velocity without the risk. See AI automation services.

Start a Project
11

A team policy by environment

EnvironmentAllowedRequired controls
Developer laptopInteractive agent on company repositoriesSandbox on, managed permissions, no production secrets, approvals for installs and network
Dev container or VMLonger autonomous runsIsolated file system, network allowlist, development credentials only
Cloud agentBackground tasks producing pull requestsScoped repo token, environment secrets limited to test services, domain allowlist
CI (agent in pipeline)Review, triage, small fixesRead-only by default, no deploy credentials, outputs as PR comments or PRs
ProductionNot for coding agentsNo direct access; changes only through reviewed deploys
12

Conclusion

Coding agents are worth using, and they are a new kind of privileged process on your machines and in your pipelines. Contain them the way you would any untrusted automation: sandbox their execution, restrict their permissions, keep secrets out of reach, treat everything they read as potentially hostile, and let nothing they produce reach production without passing the same checks as human code. For how teams use agents day to day, see AI coding agents.

Turn these controls into company rules with an AI coding policy, and govern which MCP servers coding agents may connect to with MCP governance.

FAQ

Common questions.

They can be, with the right setup. The risk comes from combining a model that can be manipulated with the ability to run commands, read files and use credentials. Sandboxing, limited permissions, no production secrets and review before merging contain most of that risk.

Get in touch

Have a project in mind?

Whether you're building a new digital product, improving an existing website, or looking to automate part of your business — let's talk.