AI-Generated Code Security: What to Check Before You Launch
What research shows about AI-generated code security, the vulnerabilities that appear most often, and a pre-launch checklist with CI controls.
Quick answer
AI-generated code is functional far more often than it is secure. Independent testing in 2026 found models chose insecure options in roughly 44 percent of security-relevant tasks, with cross-site scripting and log injection handled worst, and studies show models invent package names attackers can exploit. Before launch, treat AI-written code as untrusted: run static analysis, dependency and secret scanning in CI, have a human review authentication, authorization, input handling and data access, test access controls directly, verify every dependency exists and is the real package, and give your coding tools written security requirements so fewer problems are generated in the first place.
What the research shows
Three findings matter most for anyone shipping AI-written code.
Security has not improved as fast as capability. Veracode's 2026 GenAI Code Security Report (July 2026) tested 11 models on tasks where a secure and an insecure implementation were both possible. Code compiled almost every time, but the average security pass rate was 56 percent, barely changed from 55 percent in its first report. The best model reached 68 percent. Coding-specialized models were no more secure than general-purpose ones.
Some vulnerability classes are handled much worse than others. In the same report, SQL injection and cryptographic algorithm choice were handled correctly most of the time (83 and 87 percent), while cross-site scripting (15 percent) and log injection (12 percent) were usually wrong. Context-dependent issues, where safety depends on how data flows through the app, are the hardest for models.
Models invent dependencies. Spracklen and colleagues (USENIX Security 2025) analysed 576,000 code samples from 16 models and found 19.7 percent of suggested packages did not exist. Many hallucinated names recurred across runs, which makes them predictable targets for attackers who register them.
| Finding | Source | Implication |
|---|---|---|
| Average security pass rate 56% | Veracode 2026 GenAI Code Security Report | Assume generated code needs security review |
| XSS 15%, log injection 12% pass rates | Veracode 2026 | Focus review on output encoding and logging |
| 19.7% of package suggestions hallucinated | USENIX Security 2025 | Verify every new dependency |
| 46% of developers distrust AI output accuracy | Stack Overflow Developer Survey 2025 | Teams already sense the problem; formalize the checks |
Why AI code goes wrong
Models learn from public code, much of which is insecure or written for tutorials. They optimize for code that runs and satisfies the prompt, and prompts rarely state security requirements. They do not see your threat model, your other services or which inputs are attacker-controlled. And in agentic workflows they make many changes quickly, which increases the volume reviewers must check; DORA's 2025 research describes larger batches and review load as a key reason AI is still associated with lower delivery stability.
The vulnerability patterns to check first
| Pattern | What it looks like in generated code | Check |
|---|---|---|
| Broken authorization | Routes check login but not ownership; permissions enforced only in UI | Test as a second user and logged-out against the API |
| Cross-site scripting | Raw HTML rendering of user content; unsafe use of innerHTML-style APIs | Search for raw HTML sinks; rely on framework escaping |
| Injection | String-built SQL, shell commands or queries | Parameterized queries; no shell with user input |
| Log injection and leaks | User input logged unsanitized; tokens or personal data in logs | Sanitize log fields; redact secrets |
| Hard-coded secrets | API keys in source, front end or config committed to Git | Secret scanning; move to secret manager; rotate |
| Insecure defaults | CORS set to allow everything, debug mode on, permissive cookies | Review configuration explicitly |
| Weak crypto and randomness | Home-made token generation, outdated hashing | Use platform libraries; verified password hashing |
| Hallucinated or stale dependencies | Packages that do not exist, lookalikes, outdated versions | Verify each package; lockfiles; vulnerability scanning |
| Unsafe file and data handling | Path traversal in uploads, unsafe deserialization | Validate paths and types; avoid unsafe parsers |
A pre-launch checklist
Run through this before any AI-assisted code reaches real users. It complements the OWASP Top 10 and ASVS, which remain the right references for web application security.
- Every route and query enforces authentication and object-level authorization on the server
- Access control tested directly against the API as multiple users and as an anonymous visitor
- No secrets in source, front-end bundles, logs or Git history; exposed keys rotated
- All user input validated on the server; outputs encoded for their context (HTML, URL, SQL, shell)
- Every dependency verified as the intended, maintained package; lockfile committed
- Static analysis, dependency and secret scanning run in CI and block on high-severity findings
- Security headers, CORS, cookie flags and error handling reviewed in configuration
- Logging excludes passwords, tokens and sensitive personal data
- Rate limits on login, sign-up, password reset and any endpoint calling paid APIs
- A human who understands the code has reviewed authentication, authorization and data handling
Shipping code your team built with AI?
ZSpace Labs runs pre-launch security reviews of AI-assisted codebases and sets up the CI checks that keep them clean. See web application development.
Build the checks into CI, not memory
Manual checklists decay. Put the deterministic parts into your pipeline so they run on every pull request, whoever (or whatever) wrote it.
| Control | Catches | Notes |
|---|---|---|
| Static analysis (SAST) | Injection, XSS sinks, unsafe APIs | Tune rules to your stack to limit noise |
| Dependency scanning | Known vulnerable or unmaintained packages | Enable automatic update pull requests |
| Secret scanning | Keys and tokens in code and history | Also enable push protection where available |
| Tests for access control | Broken authorization | The highest-value tests for most apps |
| AI code review | Logic issues, missing checks, risky patterns | Useful second reader; not a gate on its own |
| Human review | Design flaws and context-specific risk | Required for authentication, payments and data access changes |
Pro tip
Require human review for changes touching authentication, authorization, payments, cryptography or data deletion, even when everything else can merge with lighter review.
Generate fewer vulnerabilities in the first place
Coding agents follow repository instructions. Write your security rules down where they will be read (an AGENTS.md or CLAUDE.md file, Cursor rules or equivalent): which libraries to use for authentication, database access and HTML rendering, that secrets must never be hard-coded, that every route needs an authorization check, and that new dependencies require approval. Point agents to secure examples in your own code. Veracode's results suggest that giving models explicit security guidance improves outcomes, though not enough to skip the checks above.
For agent-specific risks (agents running commands, reading untrusted files or holding credentials), see securing AI coding agents. For using AI as a reviewer, see AI code review.
Conclusion
AI makes code cheap to produce and does not make it secure. Treat generated code as untrusted input to your codebase: scan it automatically, verify its dependencies, test access control directly and keep a knowledgeable human on the changes that matter. With those controls, AI-assisted development can be as safe as any other; without them, it ships the same well-known vulnerabilities faster.
Common questions.
Not reliably by default. Veracode's 2026 GenAI Code Security Report found an average security pass rate of 56 percent across models on security-relevant tasks, with the best model at 68 percent. It must be reviewed and tested like any other code, with extra attention to known weak spots.