BlogDelivery practice

Spec-driven development with coding agents

Spec-driven development means agreeing in writing what a change must do, and how you will know it does, before an agent writes the code. This guide walks one change from specification to merge and compares the frameworks teams use for it.

Updated

What spec-driven development is and what it is not

Spec-driven development is a way of working with coding agents in which each meaningful change starts from a short written specification. It states the problem, the scope and non-goals, the acceptance criteria, the constraints and the test plan. The agent plans and implements against it, reviewers check the result against it, and it lives in the repository next to the code.

It is a practice for single changes. Waterfall froze the requirements of a whole project before anyone wrote code. Here a spec covers one change and is corrected as soon as implementation shows it was wrong.

It does not mean a design document for every change either. A spec is sized to its change: a few lines for a contained fix, a page for a feature across several services, nothing for a typo.

Birgitta Böckeler of Thoughtworks distinguishes three levels: spec-first, where the spec guides one task and may then be discarded; spec-anchored, where it is kept for later changes to the feature; and spec-as-source, where people edit only the spec. We recommend starting spec-first and anchoring specs for the parts of your system that change often.

Why it matters when agents write the code

An agent works from what it is given. Where a request leaves something open, it fills the gap with a plausible guess, and plausible guesses are hard to catch in review because the code looks reasonable. A spec hands those decisions to a person before the work starts: which endpoints are affected, which behaviour must not change, what the client sees on failure.

A spec also makes review against intent possible. Without one, the reviewer reconstructs the purpose of a change from the diff and a chat they never saw. With one, they check that each acceptance criterion has a test and spend the rest of their attention on what tests cannot show.

It is also what makes delegation workable. A task that can be handed over in writing, with completion criteria a reviewer can check without replaying the session, can go to an agent running in the background. Anything less clear should stay interactive.

  • The agent gets scope, non-goals and constraints it would otherwise guess
  • The reviewer gets criteria to check the diff against
  • The team keeps a record of why the code behaves as it does
  • The next session inherits context that outlived the chat

One change from specification to verified merge

Take an ordinary change: your public API has a password reset endpoint without any limit, and scripts use it to send reset emails in bulk. The fix is rate limiting, and a spec of roughly this size is enough.

Example spec for a rate limit on the password reset endpoint
# Rate limit for POST /v1/password-reset

## Problem
No limit today. Scripts use the endpoint to send reset emails in bulk.

## Scope
- Limit requests per email address and per client IP
- Return 429 with a Retry-After header when a limit is hit

## Non-goals
- Limits on other endpoints, CAPTCHA, changes to the email itself

## Acceptance criteria
1. 4th request for one email within 15 minutes returns 429
2. 21st request from one IP within 15 minutes returns 429
3. Responses do not reveal whether an account exists
4. Every 429 is logged with a hashed email and the client IP

## Constraints
- Counters in the existing Redis cluster, limits configurable without a deploy

## Test plan
- Integration tests for criteria 1 to 4 against a local Redis
- Load test in staging: added p95 latency below 5 ms
  1. 1

    Write and agree the spec

    An engineer drafts the spec from the ticket, with or without the agent, and someone who knows the system reads it before anything is planned. Non-goals and acceptance criteria need the most attention, because that is where the agent would otherwise guess.

  2. 2

    Let the agent plan, then review the plan

    The agent reads the spec and the relevant code and proposes a plan: where the limiter sits in the request pipeline, how counter keys are built, what changes in configuration. A person approves it or sends it back. Fixing a plan takes minutes, while fixing finished code takes a review cycle.

  3. 3

    Break the plan into tasks

    Each task ends in a state that builds and passes its tests, so the work can be checked or stopped at any point. Here that means the limiter with its configuration, the 429 response with logging, and the load test.

  4. 4

    Implement with tests from the criteria

    The agent writes a test for each acceptance criterion with the code, ideally first. A criterion that cannot become a test is rewritten or becomes a manual check with a named owner.

  5. 5

    Verify against the acceptance criteria

    Run the full test suite, then check each criterion against the test that covers it. If your framework has the agent compare its work with the spec, treat that as a first pass, since the agent is marking its own work.

  6. 6

    Review and merge

    The pull request carries the spec or links to it. The reviewer checks scope and non-goals against the diff, confirms each criterion has a passing test and reads the code for what tests cannot show, such as how counter keys are built. Branch protection and required checks apply as usual.

  7. 7

    Keep the spec current

    After the merge, the spec becomes part of the system's description or stays as a dated record of the change, as each repository decides. When behaviour changes later, update the spec in the same pull request, because a spec that contradicts the code misleads every agent that reads it.

How Spec Kit, OpenSpec, BMad, GSD and Kiro differ

These tools keep Markdown artefacts in your repository and drive the agent through slash commands or skills. They differ in process, in what happens to a spec after the merge and in the agents they support. We checked each against its own repository or documentation in October 2026. They release often, so check again before you standardise.

GitHub Spec Kit

Spec Kit is GitHub's open source toolkit, installed as a Python command line tool. After a one-off constitution of project principles, each feature goes through specify, plan, tasks, implement and converge, repeating the last two until convergence reports the feature as converged. Clarification, checklists and consistency analysis are optional gates. Its integration list names more than 40 agents, including Claude Code, Codex CLI, Cursor, GitHub Copilot and Gemini CLI.

OpenSpec

OpenSpec, an npm package from Fission AI, describes itself as fluid, iterative and built for existing code as well as new projects. Current behaviour lives in openspec/specs, and each change gets a folder under openspec/changes with a proposal, delta specs, a design and tasks. Delta specs record only which requirements are added, modified or removed, and archiving a finished change merges them into the main specs. The default loop is propose, apply and archive, and more than 30 tools are supported.

BMad Method

BMad Method is free, MIT-licensed and distributed as agent skills, at version 6 with version 7 in development. Its planning skills produce a product brief, a PRD, a UX design, an architecture or a spec as the work requires, and a ticket skill orders larger work into stories. Implementation runs through one build skill, one session per unit of work. It needs a coding tool that supports skills and installs through the Skills CLI or the Claude Code and Codex plugin marketplaces.

GSD Core

GSD Core continues the project formerly published as Get Shit Done and is maintained by Open GSD under the MIT licence. Each phase of a milestone runs through discuss, plan, execute, verify and ship, with the heavy work done by fresh-context subagents and decisions kept in files such as CONTEXT.md and STATE.md. It has separate entry points for new projects and existing codebases, lighter commands for small tasks, and an installer for Claude Code, OpenCode, Codex, GitHub Copilot, Cursor and others.

Kiro specs

Kiro builds specs into its IDE, CLI and web version. A feature spec has requirements, a design and a task list, starting from either the requirements or the design, while a bugfix spec begins with current, expected and unchanged behaviour. Requirements use EARS notation (Easy Approach to Requirements Syntax), naming a triggering condition and the response the system shall give, which keeps them testable. Tasks run singly or in dependency-ordered waves, and a quick spec writes all three documents in one pass without approval gates.

pstack for Cursor

pstack, Lauren Tan's Cursor plugin in the cursor/plugins repository, is sceptical of planning, and its README says so: it ships no planning skills, relies on Cursor's plan mode, and its author holds that the best spec is code. Its multi-phase plan playbook, for work spanning several pull requests, lists the files, build steps, expected result and unit, live and performance checks for each one. A script checks the plan, and execution waits for the operator's go.

Which framework fits your team

We make this choice with engineering leads, based on the repositories, the agents in use and the existing process. These questions usually settle it.

  • Existing code or new product: OpenSpec's delta specs and GSD Core's onboarding suit existing codebases, and BMad Method's planning documents suit a new product that needs a PRD and an architecture first.
  • Agents in use: Spec Kit, OpenSpec and GSD Core support many agents and BMad Method any tool with skills, while Kiro's specs belong to Kiro and pstack to Cursor.
  • Process overhead: Spec Kit has five steps per feature plus optional gates, OpenSpec's default loop three, and the others have lighter paths for small changes.
  • Several repositories: OpenSpec can keep specs in a separate planning repository shared by several code repositories, a feature still in beta.
  • Data handling: OpenSpec collects anonymous usage statistics unless you switch them off, and specs reach your agent's model just as code does.

If none fits, a spec template in the repository and a review rule take you a long way. A framework earns its setup once the habit exists and you want the same commands and checks in every repository.

Where spec-driven development goes wrong

The practice fails in a few predictable ways. Each has a plain countermeasure.

Specs that are too long

Generated specs grow quickly, and reviewers skim long ones. Reviewing Kiro and Spec Kit for martinfowler.com in October 2025, Birgitta Böckeler describes a small bug that Kiro turned into four user stories with sixteen acceptance criteria, and writes that with Spec Kit she would rather have reviewed code than all the Markdown. The tools have moved on since, but the lesson holds: keep the spec to what a reviewer will check.

Specs that drift from the code

A spec that is not updated when behaviour changes confidently describes something that no longer exists, and the next agent reads it as fact. Change the spec in the same pull request as the code, and make that part of your definition of done.

Acceptance criteria without tests

A criterion such as “fast enough” cannot fail, so each one needs an observable outcome and a test or a named manual check. Böckeler also saw Spec Kit's agent treat notes about existing classes as a new specification and generate the classes again as duplicates. Specs reduce guessing without guaranteeing that the agent follows them, so verification comes from tests and human review.

Ceremony for small changes

A full sequence for a one-line fix costs more than it saves, and the frameworks agree. BMad Method sends small, clear changes straight to its build step, Kiro offers a quick spec without approval gates, GSD Core has lighter commands for small tasks, and pstack's planning playbook says to skip the plan when the approach is obvious. Decide which change types need a spec and let the rest follow your normal process.

Introducing spec-driven development in a team

We introduce the practice one change type at a time. It runs inside the process your team already has.

  1. 1

    Pick one change type

    Choose changes that are frequent, bounded and testable, such as new API endpoints or validation rules. Everything else stays with your current process for now.

  2. 2

    Put a spec template in the repository

    Keep it to the six headings of the example and point your agent instructions at it, in AGENTS.md or your tool's equivalent. A framework's templates can replace it later.

  3. 3

    Review specs before code

    A second engineer reads the spec before the agent plans, in the pull request or the ticket. For changes to shared code, review the plan as well.

  4. 4

    Tie criteria to checks

    Make it a review rule that every acceptance criterion maps to a test or a named manual check. Where it is cheap, have CI check that changes of this type include a spec.

  5. 5

    Measure on your own work

    After a few weeks, compare changes made with a spec against similar ones made without: review rounds, rework after merge, time from ticket to merge. Move to the next change type only if the comparison supports it.

How Cloudsail helps

We run spec-driven development workshops for engineering teams in their own repositories, on changes from their own backlog. Engineers write specs for real work, review the agent's plans and check results against the acceptance criteria. Workshops run remotely or on-site in Germany and Poland, in English, German or Polish. We do not run public courses or issue certificates.

Around the workshops we set up what the practice depends on: the spec template and agent instructions in your repositories, build and test commands the agent can run, review rules and, if you want one, a framework configured for your agents. We then measure the effect on your own changes before the practice spreads to other teams.

Questions

Do we need Spec Kit, OpenSpec or another framework to start?

No. A short spec template in the repository, an agent instruction that points to it and a rule that specs are reviewed before code are enough to start. A framework is worth adding once the habit exists and you want the same commands and folder layout in every repository.

Which coding agents do these frameworks support?

Spec Kit, OpenSpec and GSD Core each support many agents, Claude Code, Codex, Cursor and GitHub Copilot among them. BMad Method works with coding tools that support skills. Kiro's specs are part of Kiro, and pstack is a plugin for Cursor.

Does spec-driven development work on an existing codebase?

Yes, if you write specs only for the change at hand and leave the rest of the system until a change touches it. OpenSpec's delta specs and GSD Core's onboarding for existing code are designed for this, and Spec Kit has a separate guide for existing projects.

How long should a spec be?

Long enough that a reviewer can check the change against it, and no longer. For a contained change that is usually well under a page. If the spec looks set to be longer than the code change, split the change or cut the spec.

Do you offer spec-driven development training or workshops?

Yes, as workshops for your team in your own repositories, remotely or on-site in Germany and Poland, in English, German or Polish. We do not run public courses or issue certificates.

Talk to an engineer

A 30-minute call with one of our engineers about your coding agent setup, what it costs and where it can improve. No access to your systems and no data shared.

Not using coding agents yet? Use the same form and tell us what you are planning.

Our team has built software and production AI for trivago, SAP, Tonies, EWE and tecRacer.

The button opens a draft in your email program. You send it yourself. Privacy policy (draft)