AI-Generated Code Provenance: How Teams Should Track Code Written by AI
How to record human, AI-assisted and agent-written code through commits, trailers, pull requests and build provenance, and what you cannot reliably detect.
Quick answer
AI-generated code provenance is the record of how a change came to exist: whether it was human-written, AI-assisted (a person accepting suggestions) or agent-generated (an agent writing and editing files), which tool and model were involved where that is known, who gave the instruction, who reviewed and approved it, and which commits, pull requests and builds it passed through.
The key point is that provenance must be recorded at creation, not detected later. No tool can reliably tell AI-written code from human-written code after the fact, especially once both have edited it. Use agent identities, commit trailers, pull request metadata, session logs and build provenance to capture what actually happened.
Why teams need to know how code was produced
| Need | Question provenance answers |
|---|---|
| Auditability | Who or what made this change, on whose instruction, and who approved it? |
| Security | Which changes came from agents that read untrusted input or had broad permissions? |
| Compliance | Can we show regulators or customers how AI was used and reviewed in this code? |
| Software supply chain | What went into this build, and was every step trusted? |
| Licensing | Did we enable public-code matching filters; was anything flagged? |
| Debugging | How was this bug introduced, and what context did the author (human or agent) have? |
| Accountability | Which person owns this change and its consequences? |
| Improvement | Do agent-produced changes cause more rework or incidents in certain areas? |
Human, AI-assisted and agent-generated: useful categories
Most teams need only a few categories, recorded at the commit or pull request level. Line-level attribution is possible with some tools but rarely necessary for governance.
| Category | Typical source | How to record it |
|---|---|---|
| Human-written | Developer writes code, maybe with search or docs | Default; developer identity |
| AI-assisted | Developer accepts inline suggestions or chat snippets and edits them | PR checkbox or label; tool named in team policy |
| Agent-generated, human-supervised | Interactive agent edits files under a developer's direction | Developer identity plus trailer naming the agent |
| Agent-generated, autonomous | Background or cloud agent produces a branch or PR | Agent bot identity; PR links to task and session |
| Agent modification of human code | Agent fixes review comments or failing tests on a PR | Separate commits under agent identity or trailer |
Worth noting
Do not claim precision you do not have. 'Agent-generated' at commit level is honest and useful. 'This file is 63 percent AI-written' usually is not, because people and tools edit each other's lines.
Commits: identities and trailers
Identities. Autonomous agents should commit under their own bot account or app identity, never a developer's personal account, so git log shows the agent as author and permissions can be scoped. GitHub's Copilot coding agent, for example, works on its own branch and opens a pull request for a person to review, and approvals from the person who asked for the work may not count toward required reviews, which keeps an independent reviewer in the loop (GitHub on reviewing Copilot pull requests).
Trailers. For interactive agent work under a developer's identity, add Git trailers at the end of the commit message. Co-authored-by is widely understood and displayed by GitHub (GitHub on co-authored commits); some agents add one automatically, and their settings control it. Teams that want machine-readable detail add custom trailers, which Git can parse with git interpret-trailers.
Add retry with backoff to webhook sender
Retries 5xx and timeouts up to 5 times with jittered
exponential backoff. Idempotency key reused across retries.
Co-authored-by: Coding Agent <agent-bot@example.com>
AI-Tool: <agent name> <version>
AI-Model: <model identifier, if exposed>
AI-Session: https://internal.example.com/agent-sessions/8f2c
Reviewed-by: A. Developer <a.dev@example.com>Pull requests: the best place for provenance
The pull request is where intent, change, evidence and review meet, so it is the most useful provenance record. Use a template with an 'AI involvement' section (none, assisted, agent-generated), the task or spec it implemented, the agent session link, test evidence and any files the agent was told not to touch. Labels (ai-agent, ai-assisted) make reporting easy. Review history then records who examined and approved the change. Our guide to AI coding agents and Git covers the full pull request workflow.
Agent session logs
For agent-generated changes, the session log (the instructions, files read, commands run, tool calls and intermediate failures) is the richest provenance there is. Store it, or a summary, for a defined retention period and link it from the pull request. It answers questions commits cannot: what context the agent had, whether it read untrusted content, and whether it was asked to weaken a test. Apply the same access controls and secret scanning to these logs as to the code itself.
Emerging attribution formats
Tool vendors are starting to standardize attribution. Agent Trace is a draft open specification, proposed by Cursor, for recording which code ranges came from AI, humans or both, with optional model identifiers, stored alongside version control in files, Git notes or a database (Agent Trace). It is a proposal, so check adoption by the tools you use before depending on it. Separately, build provenance frameworks such as SLSA describe verifiable information about how an artifact was built, from which source and by which builder (SLSA provenance). They complement each other: code provenance covers authorship; build provenance covers the path from source to artifact.
Security and the software supply chain
Agent-produced code adds supply chain questions: did the agent add dependencies, were they verified, did it read instructions from untrusted sources, did it run with network access? Provenance helps you answer them after an incident and target reviews before one. Pair it with dependency review, secret scanning and signed builds. Our guides to AI supply chain security and AI-generated code security cover the checks.
What you cannot reliably do
Be clear internally about limits. You cannot reliably detect AI-generated code by style. Developers can paste AI output from tools outside your control. Inline suggestions are accepted and edited in ways no log captures line by line. Provenance policies therefore rely partly on honest self-reporting for assisted work, and on enforced identities and logs for agent work. That is still valuable: it covers the changes with the most autonomy, which are the ones that matter most for risk.
Implementation checklist
- Define a small set of categories (human, assisted, agent-supervised, agent-autonomous)
- Give autonomous agents their own bot identities with scoped permissions
- Standardize commit trailers for agent involvement
- Add an AI involvement section and labels to the pull request template
- Link agent session logs from pull requests; set retention
- Require an independent human approval for agent pull requests
- Record dependency additions and their review
- Generate build provenance for release artifacts
- Report rework, incidents and review time by category
- Write the policy down; see our AI coding policy guide
Common mistakes
Buying an 'AI code detector' instead of recording provenance. Letting agents commit as the developer who started them, which hides autonomy. Tracking percentages of AI code as a productivity metric, which invites gaming (see measuring AI coding impact). Storing session logs with secrets in them. And treating provenance as blame: its purpose is to understand and improve the process.
Setting up governed AI-assisted development?
ZSpace Labs builds software with coding agents under clear identities, review gates and traceable pull requests. See full-stack development.
Conclusion
Code provenance for AI is a recording problem, not a detection problem. Use agent identities, commit trailers, pull request metadata, session logs and build provenance to capture how each change was produced and reviewed. Be honest about what cannot be tracked, focus on the changes with the most autonomy, and use the record for audits, security, debugging and improving how your team works with agents. Our AI coding policy guide covers the surrounding rules.
Common questions.
It is the record of how code came to exist: whether a person wrote it, an AI assistant suggested it, or an agent generated it; which tool and model were involved where known; who instructed and reviewed it; and which commits, pull requests and builds it went through.