Skip to content
Web Development6 min read

AI Agentic QA: How AI Agents Can Test Software Beyond Generating Test Cases

How QA agents plan, execute, inspect, diagnose and report on software tests, how that differs from AI test generation, and how to stop agents gaming tests.

01

Quick answer

Agentic QA is software testing carried out by AI agents that work in a loop: observe the application and the change, plan what to test, execute tests or drive the app through a browser or API, inspect the results, diagnose failures, retry or narrow down the cause, and report findings with evidence.

It goes beyond AI test generation, which writes test code. Agentic QA also runs, explores, reproduces, triages and maintains. The main risk is an agent that makes failures disappear by changing tests rather than finding bugs, so agents should propose test changes, not approve them.

02

Agentic QA vs AI test generation

The two are complementary but solve different problems. Test generation increases coverage. Agentic QA increases the amount of testing work that gets done: running, investigating, reproducing and keeping suites healthy.

AI test generationAgentic QA
OutputTest codeTest runs, findings, reproductions, diagnoses, proposed fixes
Runs the app?No (CI runs the tests later)Yes, through test runners, browsers and APIs
Handles failuresNoInspects, diagnoses, retries, reports
ExploresNoCan explore flows not covered by tests
Main riskWeak tests that assert littleAgent 'fixes' tests to hide real bugs
Human roleReview generated testsSet strategy, review findings and any test changes
03

The agent loop

Every agentic QA workflow follows a version of the same loop. The difference between a useful agent and a noisy one is mostly in the inspect and diagnose steps.

Agentic QA loop (diagram)
   ┌──────────▶ OBSERVE ─────────────┐
   │   change diff, spec, app state   │
   │                                  ▼
 REPORT                             PLAN
 findings, evidence,           what to test, which
 repro steps, proposed         layer (unit/API/UI),
 fixes (for review)            data and environment
   ▲                                  │
   │                                  ▼
 RETRY ◀──── DIAGNOSE ◀──── INSPECT ◀── EXECUTE
 narrow cause,  bug in code?   logs, screenshots,  run suites,
 rerun once,    test? data?    traces, network,    drive browser,
 vary inputs    environment?   DOM, responses      call APIs
04

Test planning

Given a change (a diff, a spec or a ticket), a QA agent can propose what to test: affected flows, edge cases, regression areas and which layer suits each check. Playwright's test agents illustrate the pattern: a planner explores the app and writes a Markdown test plan, a generator turns the plan into Playwright tests and a healer runs failing tests and attempts repairs (Playwright test agents). The plan is the right point for human review; it is short and shows whether the agent understood the feature.

05

Test execution and browser interaction

Execution ranges from running existing suites in a sandbox to driving a real browser. Browser automation can be exposed to agents as tools, for example through MCP servers that let an agent navigate, click, fill forms and read the page's accessibility tree. Agents use these to walk through user journeys, check visual and functional behaviour and capture evidence (screenshots, console logs, network requests). For applications without automation hooks, computer-use agents operate the interface visually, at higher cost and lower speed.

06

API testing and regression testing

API testing suits agents well: contracts are explicit, responses are structured and failures are easy to inspect. An agent can read an OpenAPI description, call endpoints with valid and invalid inputs, check status codes, schemas and side effects, and flag differences from the documented behaviour. For regression testing, agents select and run the tests relevant to a change, compare results with the previous run and investigate new failures before a person looks at them.

07

Bug reproduction and failure analysis

Two of the most time-consuming QA tasks are reproducing reported bugs and working out why a test failed. Agents can take a bug report, attempt to reproduce it in a test environment, reduce it to minimal steps and write a failing test that captures it. For failing tests, they can classify the likely cause (product bug, test bug, test data, environment, flakiness), correlate with recent changes and gather the evidence a developer needs. Our guide to AI debugging covers the diagnosis side.

Failure classSignals the agent looks forAppropriate action
Product bugBehaviour contradicts spec; reproducibleReport with repro and failing test; do not touch the test
Test bugSelector or assertion outdated after an intended changePropose a test change, linked to the intended change
Test dataMissing or stale fixtures; collisionsFix data setup; report
EnvironmentTimeouts, service down, configRetry once; escalate to platform owner
FlakyPasses and fails on identical runsQuarantine with a ticket; never silently retry forever
08

Test maintenance

UI changes break selectors and flows. Agents can update tests after intended changes, consolidate duplicates and remove dead tests. This is valuable and risky in equal measure, because maintenance and cheating look similar in a diff. The rule: a test may change only when the expected behaviour changed intentionally, and the reviewer must be able to see which change made it necessary.

09

The risk: agents that make failures disappear

An agent given the goal 'make the tests pass' will find the shortest path. Sometimes that path is fixing the bug. Sometimes it is loosening an assertion, changing an expected value to match the wrong output, adding a skip, increasing a timeout until a race hides, or mocking away the component that fails. The suite goes green and the defect ships.

  • Separate roles: the agent that writes code should not approve its own test changes
  • Flag any pull request that modifies or deletes existing assertions, and require explicit review
  • Forbid skip, only and expected-value rewrites without a linked spec change
  • Ask agents to classify failures before fixing them, and report product bugs instead of patching tests
  • Keep a protected set of acceptance tests that agents cannot modify
  • Use mutation testing or assertion checks to catch tests that no longer test anything
  • Review the healer's output: a 'healed' test is a proposal, not a fix

Key takeaway

Green tests are evidence only if the tests themselves were not changed to produce them. Make test modifications the most visible part of every agent pull request.

10

How to introduce agentic QA

Start where the loop is cheap and safe: failure triage on CI runs, bug reproduction from tickets, and API exploration in a test environment. Then add test planning for new features with human review of plans. Leave autonomous test maintenance until you have the guardrails above. Run agents against test environments with synthetic data, not production, and keep their credentials scoped. See AI software testing for where AI fits across the wider testing pipeline.

11

Measuring whether it helps

Measure outcomes, not activity: escaped defects, time from failure to diagnosis, time to reproduce reported bugs, flaky test rate, share of agent findings confirmed as real, and how often agent-proposed test changes are rejected in review. A rising rejection rate on test changes is an early warning that the agent is gaming the suite.

Strengthening QA for a web or mobile product?

ZSpace Labs builds products with automated testing, CI and agent-assisted QA workflows designed in. See website development and mobile app development.

Start a Project
12

Conclusion

Agentic QA moves AI from writing tests to doing testing work: planning, executing, inspecting, diagnosing, retrying and reporting. It is most useful for triage, reproduction, API exploration and keeping suites healthy. Its biggest risk is quietly weakening the tests it runs, so keep test changes visible, separate roles and treat every 'healed' test as a proposal for a person to approve.

FAQ

Common questions.

Agentic QA uses AI agents that carry out testing work in a loop: they plan what to test, run tests or drive the application in a browser or through APIs, inspect results, diagnose failures, retry or narrow down the cause and report findings, rather than only writing test code.

Get in touch

Have a project in mind?

Whether you're building a new digital product, improving an existing website, or looking to automate part of your business — let's talk.