
A definition of done for AI coding agents is a shared rule for when a requested change can be accepted, backed by evidence for the behavior that changed. It names the expected outcome before coding, the smallest relevant checks, who accepts the result, and which claims remain unverified. A patch is code, a passing test is evidence for a case, a merge puts a revision on a branch, a deploy records what reached an environment, and a live check observes what a user can actually do. Those are separate facts. This guide gives engineering leads and agent operators a reusable checklist and an illustrative evidence record.
The term comes from Scrum: the official Scrum Guide defines a Definition of Done as the product quality measures an Increment must meet and says it gives everyone a shared understanding of completed work. An AI agent is not a Scrum Team or an Increment. Here the useful adaptation is the shared, inspectable completion standard, not a claim that Scrum prescribes an agent workflow. As of October 2026, each team still needs to choose checks appropriate to its product and risk.
Start with the changed behavior
Write the promise in terms a reviewer can observe. ‘Fix export’ leaves room for an agent to change a button and announce success while the downloaded file still has the wrong rows. ‘When a user selects two report rows and downloads CSV, the file has those two rows in the documented column order’ identifies the behavior. Add what must remain true: downloading with no selection still uses the existing all-rows path, if that is the agreed product rule. Name the environment and identity when they affect the result. These statements become acceptance criteria, not merely implementation instructions.
Then attach evidence to each criterion. A focused component test might show the selection is passed to an export handler. A file assertion might show headers and row count. A browser download might show the complete user path. None of these proves the others automatically. Choose the cheapest evidence that actually covers the promise, and label a substitute honestly when the intended check is unavailable. The acceptance criteria generator can format a draft from details you enter; it is deterministic and does not inspect code, execute a test, or decide whether a result meets the criteria.
A reusable definition-of-done checklist
Use the following list as a template for a specific task, not as a demand to run an entire repository suite for every edit. For a small text correction, the relevant content check and a review may suffice. For a security boundary or data migration, the proof and authorization will be deeper. Write ‘not applicable’ with a reason when a stage genuinely does not apply; write ‘not checked’ when it does apply but has not been verified.
- Outcome: State the changed behavior and at least one preserved behavior in observable terms.
- Scope: Identify permitted paths, systems, and actions; record who may approve a wider change.
- Revision: Identify the code revision or diff being reviewed and the environment used for each check.
- Evidence: Pair each acceptance criterion with the command or observation and its actual result.
- Failures and gaps: Include failed, skipped, blocked, and untested cases with their impact on the claim.
- Review: Record who accepted the result and which scope or risk questions they resolved.
- Merge: If required, record the merged revision and branch separately from local proof.
- Deploy: If required, record the deployed revision and target environment separately from the merge.
- Live result: If promised, observe the agreed user path on the deployed revision and record what happened.
- Handoff: Leave a current work record and a next action for any remaining gap.
The checklist is useful only if its entries can be inspected. ‘Tests pass’ is weaker than ‘CSV export test passed on revision abc123 in local development; browser download not checked.’ Attach the command, observed output or artifact, revision, and environment when they matter. A reviewer can then decide whether evidence covers the promise without trusting an agent’s fluent summary. Keep the record next to the task, where the next agent can refresh it. The context engineering guide shows how to carry current sources and unresolved questions into a handoff.
Separate patch, test, merge, deploy, and live
An agent may finish a patch before the work is accepted. A test run proves only the cases it exercised on the revision and environment where it ran. A review can accept a proposed change but does not make that revision present on the target branch. A merge records that code entered a branch; it does not prove a deployment occurred. A deployment record says which revision reached an environment; it does not prove a user can complete the path there. A live check supplies that last observation, with its own date, environment, and limits.
This distinction is operational, not ceremonial. GitHub’s status-check documentation describes checks as information about whether a pull request is ready to merge, and required checks apply only when configured for a protected branch. That platform behavior supports a narrow point: green status is evidence for a gate, not evidence of deployment or user success. Other merge systems have different controls. Inspect the actual pipeline and record its result instead of inferring it from a test badge.
Do not turn every task into an automatic production release. The definition of done should follow the request: a library refactor may be complete when its contract checks and review are accepted, while a request to restore a broken production flow remains open until that flow is checked live. If deployment requires a separate person or window, mark the engineering change accepted and the release or live verification pending. This preserves the truth of both states without forcing an unrelated action. The human review guide explains how to place an approval at the consequential decision.
Example: prove a CSV export change
Consider an invented reports page where selecting rows should limit a CSV download. The team agrees on two outcomes: selected rows yield a file with exactly those rows and the documented headers; with nothing selected, the old all-rows download still works. The agent may edit the export component and its focused test, but the lead owns merge and release. The following record is illustrative: task name, revision, commands, observations, and outcomes are invented to show how a real report should read. They are not AppHandoff customer evidence or results from this article’s development.
Task: selected-row CSV export (illustrative)
Revision under review: abc123, local branch
Criterion A: selected rows only; headers in documented order
PASS: focused export test on abc123, local Node environment;
sample file has 2 selected rows and expected headers.
Criterion B: no selection keeps the all-rows path
FAIL: focused test on abc123 returned an empty file.
Browser download: NOT TESTED; preview was unavailable.
Review: PENDING; lead has not accepted the changed behavior.
Merge: NOT DONE. Deploy: NOT DONE. Live result: NOT VERIFIED.
Next action: fix Criterion B, rerun both focused cases,
then exercise the download path before requesting review.This is useful despite the failure. The reader can see that Criterion A has bounded evidence, Criterion B blocks acceptance, and the full download path is still unknown. The agent should not rewrite the report as ‘CSV export done’ because one path passed. After fixing the fallback, it should replace the result with the actual new revision and rerun cases affected by the edit. If a reviewer later approves it, that decision can be recorded without pretending merge, deploy, or live proof occurred. The record follows the work rather than a preferred success story.
Keep evidence attached to the right revision
Evidence ages when code, configuration, data, or acceptance criteria change. A test from yesterday can still explain why a bug was found, but it is not proof that today’s edited branch passes. A live observation of an older deployment does not verify a new release. Before accepting a result, compare the revision used for each check with the revision proposed for review. Repeat the affected check if intervening edits could change that behavior; there is no need to rerun unrelated broad suites merely to make the report look thorough.
Separate the evidence types in the task record. Store a failing assertion and later passing assertion with their revisions, not a single ‘green’ label. For visual or interactive work, name the path a person exercised and what they saw. For a backend contract, name the request and response observed. A bounded local handler fixture is evidence about that handler, not proof that a remote client or modern MCP connection works end to end. The MCP server testing guide develops that distinction for protocol-facing work.
Make acceptance and handoff visible to the next agent
A definition of done is easiest to use when the work item holds the outcome, scope, evidence, decision, and next action together. If a second agent arrives after a context reset, it can read the current criteria before touching code and distinguish a pending release from a verified live result. The record must be updated from the actual branch and system state; copying a polished prior summary is insufficient. If a check was blocked, state what blocked it and who can clear it. If a person needs to authorize a step, identify that decision without treating silence as approval.
AppHandoff provides a shared task record through its documented MCP endpoint, https://api.apphandoff.com/mcp. Its current catalog contains bootstrap, get, find, ticket, plan, message, project, and decide_lifecycle_proposal. An authorized agent can resolve the project and read a work item before acting; writes depend on action rules and account scope. A lifecycle approval card is a signed-in human decision for a specific proposal, not a blanket acceptance of the code. The official July 2026 MCP tools specification defines tool discovery and calls; it does not decide a product team’s acceptance policy. The MCP overview explains the product connection surface.
When closing a handoff, write the most precise claim the evidence allows: ‘focused checks passed, review pending,’ ‘merged, deployment not recorded,’ or ‘deployed revision verified through the agreed live path.’ State the remaining question and next owner when the claim is narrower than the original request. This lets a human accept a bounded result and lets another agent continue without rediscovering whether ‘done’ meant code written, code merged, or behavior observed.
Frequently asked questions
What is a definition of done for an AI coding agent?
It is a shared completion standard that asks the agent to show evidence for the changed behavior, the applicable quality checks, and any unresolved limits. A patch, passing test, merge, recorded deploy, and verified live result are different states; report only the ones actually reached.
Does every agent change need a deployment?
No. Match completion to the authorized task and environment. A local documentation change may be accepted after review and the relevant checks. If the task promises a live behavior, record the deployed revision and observe the agreed user path before calling that live behavior verified.
Is a green test run enough to call the work done?
A green run supports the cases it exercised on a particular revision and environment. Compare those cases with the acceptance criteria, report failures and skipped checks, and verify any promised live behavior separately. A green run does not prove an untested outcome or authorize a merge or deployment.