AI Coding Agents: How They Work and How Development Teams Use Them
How AI coding agents work: repository exploration, planning, editing, running tests, iterating and opening pull requests, plus task design, permissions, sandboxing, review and where agents struggle.
Quick answer
An AI coding agent takes a software task and works through it like a junior engineer with fast hands: it explores the repository, plans a change, edits files, runs tests and builds, reads the failures, iterates and hands back a result, usually a pull request. Agents run interactively in an IDE or terminal, or in the background in a cloud sandbox or CI. They work best on well-specified, testable tasks, need least-privilege access and sandboxing, and every change should pass CI and human review before merge.
Where This Fits
The broader picture is in AI software development. How agents compare with completion and chat is in AI-assisted vs agentic coding, and reviewing their output in AI code review. General agent design principles apply too; see AI agent architecture.
How a Coding Agent Works
Underneath, a coding agent is the same loop as any AI agent, specialised with developer tools: search the codebase, read files, edit files, run shell commands (tests, linters, builds), and sometimes browse documentation or call issue trackers. The model decides which tool to use next; the environment decides what each tool is allowed to touch.
- Explore: search for relevant files, read code, tests and project instructions
- Plan: outline the change, sometimes shown to the developer for approval
- Edit: modify files across the codebase
- Verify: run tests, type checks, linters and builds
- Iterate: read failures and fix them, within a step or time budget
- Hand off: summarize changes and open a pull request or present a diff
Interactive vs Background Agents
| Interactive (IDE or terminal) | Background (cloud or CI) | |
|---|---|---|
| Where it runs | Developer's machine or dev container | Hosted sandbox or CI runner |
| Supervision | Developer watches and steers | Developer reviews the pull request |
| Good for | Exploratory work, debugging, larger changes with guidance | Backlog issues, routine fixes, parallel tasks |
| Risks | Commands run with the developer's local access | Less steering; vague tasks drift |
| Examples | Agent modes in IDE assistants, Claude Code, Codex CLI | GitHub Copilot coding agent, Codex cloud tasks, Claude Code in GitHub Actions |
Writing Tasks Agents Can Complete
Most failed agent runs start with a vague task. Write issues as you would for a capable new teammate who knows nothing about the history: what the problem is, where the relevant code is, what done looks like, how to verify it and what must not change. Many tools also read repository instruction files (for example conventions, commands and architecture notes), which improves results across all tasks.
Title: Reject expired coupon codes at checkout
Problem: Expired coupons are accepted; discount applied after expiry date.
Where: src/checkout/coupons.ts (validateCoupon), tests in tests/checkout/
Acceptance criteria:
- validateCoupon returns { valid: false, reason: "expired" } when now > expires_at
- Timezone: compare in UTC
- Existing valid-coupon tests still pass
Verify: npm test -- tests/checkout
Do not change: coupon schema, public API response shapeWant coding agents working on your backlog safely?
ZSpace Labs can set up agent workflows, repository instructions, CI guardrails and review practices on your codebase.
Permissions and Sandboxing
- Branch-only write access; branch protection on main with required reviews and checks
- Scoped, short-lived tokens; no production credentials or customer data in the environment
- Sandboxed execution for commands, with network restrictions where possible
- Approval prompts for risky commands in interactive tools
- Treat issue text, comments and repository files as untrusted input that may contain injected instructions
- Log agent actions and keep pull request history as the audit trail
Reviewing Agent Pull Requests
Review agent output as you would a new contributor's: check that it solved the stated problem, did not change unrelated code, added or updated tests that actually exercise the change, follows conventions and does not introduce dependencies without reason. Keep pull requests small; ask the agent to split large changes. AI review tools can help triage, but a person approves. See AI code review.
Where Agents Struggle
Agents are weaker when requirements are ambiguous, when the right answer depends on knowledge outside the repository (business rules, customer commitments), when changes span many services, and in areas where tests are thin. They can also loop on failing tests or 'fix' tests to pass instead of fixing code. Step budgets, clear acceptance criteria and reviewers watching for test changes address most of this.
Advantages and Limitations
| Advantages | Limitations |
|---|---|
| Handle well-scoped tasks end to end | Need precise tasks and good tests |
| Run in parallel on backlog items | Review load shifts to developers |
| Explore unfamiliar code quickly | Can make confident, wrong changes |
| Iterate against tests automatically | Security exposure if permissions are loose |
How to Introduce Coding Agents Step by Step
- 1. Add repository instructions: build and test commands, conventions, architecture notes
- 2. Make tests fast and reliable so agents can verify work
- 3. Configure permissions: branch-only, scoped tokens, sandbox
- 4. Create an agent-ready issue template
- 5. Pilot on labelled small issues and track merge rate and review time
- 6. Expand task types as results justify
- 7. Review security and cost monthly
Repository Instructions That Help Agents
Most agent tools read a project instruction file (names vary by tool) before starting work. A short, accurate file saves every agent run from rediscovering the basics and steers it toward your conventions.
# Project notes for AI agents
## Commands
- Install: npm ci
- Test all: npm test
- Test one area: npm test -- tests/checkout
- Lint and types: npm run lint && npx tsc --noEmit
## Conventions
- TypeScript strict; no `any` in new code
- Money in integer minor units (pence), never floats
- Errors: throw typed errors from src/errors.ts; never swallow
## Boundaries
- Do not edit generated files in src/generated/
- Do not change database migrations that are already merged
- Payment and auth code (src/payments, src/auth) needs a human-led changeCost, Throughput and Parallel Work
Background agents can work on several issues in parallel, which raises throughput but also review load and usage costs. Track cost per merged pull request and review time per agent change, cap concurrent agent tasks per team to what reviewers can handle and stop runs that exceed step or time budgets. Parallelism is only useful if the review pipeline keeps up; see AI code review for first-pass help.
Choosing a Coding Agent
The main options in 2026 include GitHub Copilot's coding agent, which works from assigned issues and opens pull requests; Claude Code, which runs in the terminal, IDEs and GitHub Actions; and OpenAI Codex, available as a cloud agent and a CLI. Several IDEs also include agent modes. Capabilities change quickly, so evaluate on your own repositories rather than on published comparisons.
Practical selection questions: Where does the agent run, and can you control its network access and secrets? Does it work with your repository host and CI? Can administrators set policies, view audit logs and limit which repositories it can touch? How is usage priced, and how predictable is cost when agents run long tasks? What happens to your code and prompts under the provider's data terms? Teams often end up with more than one tool: an interactive agent for developers and a background agent for well-specified issues.
Official documentation: GitHub Copilot coding agent and OpenAI Codex.
Handling Agent Failures
Agents fail in recognizable ways: they loop on a failing test, make a change that passes tests by weakening them, misread the task and solve a different problem, or produce a sprawling diff touching unrelated files. Step and time limits stop loops; review checks catch weakened tests; small, explicit tasks reduce misreading.
When an agent's pull request is wrong, resist the urge to fix it by hand on the same branch. Close it, record why, and improve the task description or repository instructions so the next run succeeds. Over time these notes reveal which task types agents handle well in your codebase. Debugging techniques for agent-produced changes are in AI debugging, and the wider lifecycle in AI in the SDLC.
Agents in CI and Automation
Beyond working on issues, agents can run inside CI workflows: proposing fixes for failing builds, updating dependencies, triaging new issues or responding to review comments. Claude Code, for example, has a GitHub Actions integration, and other tools offer similar hooks. These workflows are powerful because they run unattended, which is also why they need the strictest permissions.
Run CI agents with tokens limited to the repository and to creating branches and pull requests, never to merging or deploying. Restrict which events can trigger them, since comments from outside contributors can carry prompt injection. Log every run and review a sample regularly. Security principles for agents are in AI security for business applications.
Worked Example
An illustrative scenario, not a client case: a team labels 30 small issues for a background agent. Twenty produce pull requests merged after light review, six need significant rework and four are abandoned because the issue lacked context. The team updates its issue template, adds a missing test command to the repository instructions and excludes issues touching billing logic, where reviewers found the agent's changes risky.
Common Mistakes
- Vague one-line tasks
- Giving agents broad tokens or production access
- Merging without reading the diff
- Not noticing agents edited tests to make them pass
- Large pull requests that are hard to review
Planning an agentic development workflow?
Talk to ZSpace Labs about AI-assisted development and agent workflow design.
Conclusion
Coding agents are productive when tasks are clear, tests are strong, permissions are narrow and people review every change. Related: assisted vs agentic coding, AI code review and AI test generation.
Common questions
An AI system that takes a software task, explores the repository, plans changes, edits files, runs commands such as tests and builds, iterates on failures and produces a result such as a pull request, with tools and permissions defined by the environment it runs in.