BlogProcess
AI SDLC: what coding agents change in each phase of the lifecycle
An AI SDLC is the lifecycle your team already runs, with coding agents doing part of the work in every phase and named people deciding at every handoff. This guide covers what agents can take on from planning to operations, which artefact passes between phases, who owns it, and how to measure and govern the change.
Why the lifecycle is the unit of change
Most teams meet coding agents in the editor, and early measurement stays there: seats, sessions, tokens and accepted suggestions. That view stops being useful once agents produce whole diffs and pull requests. The implementation step gets faster, and work piles up on either side of it. Tasks arrive underspecified, while review, test pipelines and release processes receive more changes than they were built for.
DORA's 2025 report on AI-assisted software development finds that AI adoption now improves software delivery throughput, a shift from the previous year, but still increases delivery instability. The report describes AI's primary role as that of an amplifier, magnifying an organisation's existing strengths and weaknesses. An already strained review or release process shows its weaknesses sooner.
AWS takes a lifecycle view in its AI-Driven Development Life Cycle (AI-DLC), described in July 2025. In AI-DLC, AI creates plans, asks clarifying questions and implements only after human validation, across three phases: Inception, Construction and Operations. Teams validate the AI's proposals together in sessions AWS calls Mob Elaboration and Mob Construction, plans and design artefacts are stored in the project repository, and sprints give way to shorter cycles called bolts. The method's open-source workflows now add initialisation and ideation phases, run in several coding agents and stop at human approval gates.
You do not need a named method to bring AI into the software development lifecycle. For each phase, write down two things: how work is split between agents and people, and which artefact passes to the next phase, with its owner. That is what we mean by an AI SDLC, or agentic SDLC: the lifecycle you already run, with those decisions made explicit. These are the handoff artefacts and their owners.
- Spec, owned by the task writer
- Plan, owned by the engineer accountable for the change
- Diff (the pull request), owned by the engineer who delegated the task
- Test report, produced by CI and read by a named person
- Review record, owned by the named reviewer
- Release note, owned by the release owner
- Incident record and follow-up issues, owned by the service team
From request to diff: planning, design and implementation
Problems that start in the early phases get expensive later, because nobody can review a diff against requirements that were never written down. These phases decide whether an agent's output is worth a reviewer's time.
Requirements and planning
Agents are useful before anyone writes code. They can check an issue against the codebase and its history, list the modules a change would touch, ask the questions the issue leaves open and draft acceptance criteria. Deciding what to build, for whom and what counts as done stays with people. The spec is the handoff, and each task should be cut so it produces one diff a reviewer can read in one sitting.
Design
For anything that crosses a module boundary, ask the agent for a plan before the first edit: files to change, interfaces and data shapes, migration order and how the change will be tested. An agent can also build a throwaway prototype to settle a design question cheaply. People approve plans that touch public APIs, data models, security boundaries or another team's code, and reject plans that touch more than the task needs. The approved plan stays in the repository or the pull request, so the reviewer can check the diff against it.
Implementation
Here agents do most of the work, and the harness matters most. The agent writes the change, runs build, lint and tests from a clean checkout and iterates on failures, while independent tasks run in parallel in isolated workspaces. The engineer who delegated the task stays its owner: they steer the session, restart it when it drifts from the brief, or take over by hand. The pull request states what was run and what was not verified.
Task PAY-412 Retry failed webhook deliveries Spec docs/specs/PAY-412.md, approved by the task writer Plan approved in the pull request before the first edit Agent tool, model and harness revision Ran lint, unit and integration tests: passing Not run load test against the provider sandbox Risk retries pile up during a long provider outage Owner engineer who delegated the task Reviewer named person, merges after CI passes
From diff to production: testing, review, release and operations
Later phases absorb whatever the implementation step produces. Their capacity, more than the agent's speed, sets how quickly changes reach users.
Testing
Agents write tests quickly, which helps and is also the main risk in this phase. A test written by the agent that wrote the code shares the agent's reading of the task, including its misreadings. Watch for updated snapshots, loosened assertions and skipped tests that turn a red build green. Agents do well at reproducing a reported bug as a failing test before the fix, and at adding coverage around code they are about to change.
People decide what must be tested, and they keep the acceptance tests that matter out of the agent's reach. The test report is the handoff: CI results mapped to the acceptance criteria, plus anything checked by hand.
Review
Review fills up first when agents produce diffs faster than people can read them. An agent can do a first pass in a fresh context or with a different model, check the diff against the plan and your conventions, and summarise it for the reviewer. Approval stays with a named reviewer, who is accountable for what merges. The review record, findings included, stays with the pull request.
Release
Agents can draft release notes from merged pull requests, check that migrations and feature flags match the plan, and prepare the deployment change. When to release, to whom and whether to roll back are decisions for the release owner, under your existing change management. The release note is the handoff: what changed, known risks and the rollback path.
Operations and maintenance
In operations, start agents on read-only work, such as correlating logs, traces and recent deployments during an incident, and have them propose fixes as ordinary pull requests. Recurring maintenance suits agents well, because dependency updates, deprecations and small migrations are repetitive and checkable. Incident command, production changes and customer communication stay with people. Follow-up issues from the incident record go back into planning.
Packaged workflows that span several phases
Some teams adopt a packaged workflow instead of designing every handoff themselves. One example is pstack, a Cursor plugin by Lauren Tan, published in the cursor/plugins repository under the MIT licence. Its playbooks cover recurring engineering work, from investigation and bug fixes to features, refactoring and performance work, and some drive pull requests through CI and review comments to merge. Further skills have several models try to break a diff, and a rule written at setup assigns models to roles such as coding, judgement and review.
A packaged workflow carries its author's view of where a person steps in. The principles in pstack tell the agent to proceed on reversible work and ask before irreversible actions, and its README includes an overnight run that lands a stack of pull requests. A sibling playbook builds and verifies a stack, then leaves review and landing to a person. Whether merging to a branch that deploys is reversible is your organisation's decision.
Playbooks instruct the agent and enforce nothing. When we adopt a packaged workflow with a team, we map each step to its phase, stop it at your approval points, and adapt, restrict or leave out steps that assume more autonomy than your setup grants. The approval gates built into the AI-DLC workflows help as well, and they do not replace yours. Approval points that must hold belong in systems the agent cannot change on its own.
- The agent's permissions and credentials, scoped to the task
- Branch protection and required reviews
- CI checks that block a merge when they fail
- Deployment approvals for production
pstack workflows in your repositoriespstack on Cursorpstack source and README
Roles that change
Job titles rarely need to change. What changes is where people spend their week and which artefacts they own.
- Task writers produce briefs an agent can execute without asking, with acceptance criteria and an out-of-scope list.
- Engineers who delegate own every diff they hand to an agent, including the call to restart a session or take over.
- Reviewers spend more time on review and need small diffs, the approved plan and the test report.
- The platform team owns the harness: instructions, build and test commands, permissions, hooks, model defaults and budgets.
- Security and release owners decide which task types may run unattended and approve the exceptions.
Version the harness with a change process, and cap how many agent pull requests can wait on one reviewer. DORA's AI Capabilities Model, a companion to the 2025 report, lists quality internal platforms among seven capabilities shown to magnify the positive impact of AI. The harness belongs to that platform.
Metrics that still make sense
DORA's software delivery performance metrics still describe the outcome. Change lead time, deployment frequency and failed deployment recovery time measure throughput, while change fail rate and deployment rework rate measure instability. DORA applies them at the application or service level, which suits agent adoption: the question is whether the service delivers better as a whole.
These metrics move slowly and respond to many causes. Add measures closer to the agent's work and track them per task type.
- Cost per accepted change: model and tool spend for a task type, divided by the changes that merged and stayed merged.
- Review time per agent pull request, from ready for review to approval.
- Rework: follow-up commits that fix an agent change soon after merge.
- Reverts, and agent pull requests closed without merging.
- Time spent on briefs and plans, so the upstream cost stays visible.
Report by team and task type, never by individual engineer. Bug fixes, refactors and dependency updates behave differently, and an average across them hides where agents help. DORA warns that turning metrics into goals invites gaming, and individual scores also cost trust. Leave out lines of code, suggestion acceptance rates and pull request counts, which can all rise without better delivery.
Governance for unattended runs and approvals
Set governance per task type. A dependency update in a well-tested service and a schema migration in a payment system need different rules, even with the same tool and model. An unattended run makes sense when the task can be handed over in writing, the test suite runs on its own without flaking, and a person reviews every change before it merges. The run itself needs permissions bounded by the task, an isolated workspace and scoped credentials. Anything that changes production, data or access needs a named approver, enforced outside the agent.
- Merges to protected branches: a named reviewer, enforced through branch protection and required reviews.
- Production deployments and rollbacks: the release owner, through your deployment approvals.
- Schema and data migrations: the owning team, with a tested rollback path.
- New tools, models, MCP servers or credentials for agents: the platform and security teams.
Audit trail
For each change, keep the brief, the approved plan, the agent configuration (tool, model and harness revision), the test report, the review record and who merged. Store it with the pull request, where it outlives any tool's session history. Our rule is that the agent proposes, and deterministic checks and people decide. Our founder's open-source project Runmill follows it: deterministic code decides pushes, pull requests and merges, never the agent.
Introducing an agentic SDLC one phase at a time
None of this needs a new process or a reorganisation. DORA's AI Capabilities Model lists working in small batches among its seven capabilities, and a rollout benefits from the same discipline.
- 1
Pick one phase and one task type
Choose work with reliable tests and a small blast radius, such as small bug fixes, test coverage for one service or dependency updates.
- 2
Write down the handoffs
Name the artefact the agent receives, the one it returns and the owner of each.
- 3
Take a baseline
Record the DORA metrics for the service, plus review time, rework and cost for that task type, before anything changes.
- 4
Run with checks unchanged
Keep review and CI as they are, and record the tool, model and harness revision for every run.
- 5
Compare and decide
Widen to the next task type or an adjacent phase, fix the harness, or stop. Unattended runs come last.
How Cloudsail helps
We work with one team at a time, in its own repositories. We are vendor-independent, and no model or tool vendor pays us.
- We map the phases, handoff artefacts and owners for one task type, then widen.
- We set up the harness and the approval points, versioned and with an owner.
- We run workshops remotely or on-site in Germany and Poland, in English, German or Polish, as team sessions on your repositories and backlog.
- We measure the effect by replaying merged changes against held-out tests, and report cost per accepted change by task type.
Coding agent adoptionHarness engineeringEvaluations on your repositories
Questions
Is an AI SDLC the same as AWS's AI-DLC?
No. AI-DLC is a specific method from AWS, with its own phases, team rituals and open-source workflows. The questions in this guide, such as who decides in each phase and who owns each handoff, apply whether you adopt AI-DLC, a packaged workflow such as the pstack plugin for Cursor, or no named method at all.
Can coding agents run the whole lifecycle on their own?
Agents can take on work in every phase, and we do not recommend taking people out of the decisions. What to build, merge approval, production releases and incident command stay with named people. Unattended runs make sense for task types with written briefs, reliable tests and a review before merge.
Which phase should we start with?
Usually implementation or testing, for one narrow task type in a service with reliable tests, because your existing review and CI catch mistakes there. Release and operations come later, starting with read-only work such as incident investigation.
How do we know whether it is working?
Compare the same task type before and after: the DORA metrics for the service, plus cost per accepted change, review time, rework and reverts. If throughput rises while rework or change fail rate rise with it, fix testing and review before you widen.
Will these metrics be used to rate individual engineers?
Not in any setup we build. We report by team and task type, because individual scores invite gaming and cost trust. It also makes it easier to brief a works council early, where you have one.
Related services
Coding agent adoption
Rolling out one tool in one team, or several tools across many, in a form your platform team can run without us.
Evaluations on your repositories
Tools, models and settings compared on work your team has already merged.
Harness engineering
The instructions, commands, permissions and checks that decide how an agent behaves in your repositories.
More from the blog
Spec-driven development
How to write specs coding agents can follow, take one change from spec to merge, and choose between Spec Kit, OpenSpec, BMad, GSD and Kiro.
AI code review
How to set up AI code review and keep a person accountable for every merge when coding agents write more of your pull requests.
Coding agent security
Practical security and governance for coding agents: prompt injection, secrets, permissions, sandboxing, unattended runs and ownership.
Sources
- AWS DevOps Blog: AI-Driven Development Life Cycle: Reimagining Software Engineering
- AWS DevOps Blog: Open-Sourcing Adaptive Workflows for AI-Driven Development Life Cycle (AI-DLC)
- awslabs/aidlc-workflows on GitHub (README and phase guide)
- DORA: DORA's software delivery performance metrics
- DORA: State of AI-assisted Software Development 2025
- State of AI-assisted Software Development 2025 (full report, PDF)
- Google Cloud Blog: Announcing the 2025 DORA Report
- Google Cloud Blog: Introducing the DORA AI Capabilities Model
- pstack in the cursor/plugins repository (README, playbooks and principles)
Talk to an engineer
A 30-minute call with one of our engineers about your coding agent setup, what it costs and where it can improve. No access to your systems and no data shared.
Not using coding agents yet? Use the same form and tell us what you are planning.
Our team has built software and production AI for trivago, SAP, Tonies, EWE and tecRacer.
Check your email program
We tried to open a draft in your email program. Nothing is sent until you send it, and this page cannot tell whether the draft opened or the email was delivered. The form stays editable, and you can also write to miki@cloudsail.com directly.