BlogStrategy
Autonomous software factory: what it takes and where its limits are
An autonomous software factory lets coding agents take scoped work from issue to merge, while people decide what gets built and what may merge. We describe what it takes to build one on proven engineering practice, which work fits first and where the limits are.
What we mean by an autonomous software factory
We use the term for a delivery line in which coding agents take scoped work from an issue to a merged change, under rules your team has written down. A written spec defines the task, the agent works in an isolated environment, and deterministic checks and people decide what merges. Most of the factory is the line around the agent: intake, queue, environments, gates, merge policy and feedback.
Older meanings of the term
The word is older than coding agents. The US Department of Defense describes a software factory as a collection of people, tools and processes that delivers software continuously, with CI/CD pipelines as its assembly lines. Some IT service providers use it for a structured model of outsourced development. Earlier still, between the late 1960s and the late 1970s, Hitachi, Toshiba, NEC and Fujitsu set up software factories to standardise their processes and spread proven tools and techniques among their staff.
A factory in the agent sense keeps the pipeline and the standard process, and adds agents that write the changes. That moves the hard problem to verification, because the factory has to check work that nobody watched being done.
Dark factories in public discussion
Some public writing goes further. In January 2026 Dan Shapiro published a scale for AI-assisted development whose top level he calls the dark factory: a black box that turns specs into software, named after robot-run factories where people are not needed. In February 2026 StrongDM described the software factory of its AI team, where specs and scenarios drive the agents and the stated rules are that humans neither write nor review the code. Validation there relies on end-to-end scenarios, often kept outside the codebase like a holdout set, and on behavioural clones of the third-party services the software depends on.
What a factory does not promise
We do not build dark factories or promise production without people. In task types that suit them, agents take over much of the writing. People still set direction, accept the risk of each task type and own what reaches production, including incidents and reverts.
- Which task types the factory may take, and which it may not
- The spec for each task, written or approved by a person
- The merge policy, its owner and every change made to it
- Review of changes to the harness, the gates and the policy itself
- Spot checks on merged work, including task types that merge after checks
- The decision to pause the line
Autonomy can grow one task type at a time. If dependency patch updates pass review without findings over an agreed period, the owner may let them merge once the checks pass. That decision lives in code, has a named owner and can be reversed with a single change.
The stages that come first
On the path we describe, the software factory is the last of four stages. In agentic coding, agents work alongside every engineer, on a budget, with a shared harness in each repository. In the optimised stage, every change to that setup, from a default model to an instruction file, is measured on your own code. The factory runs the same setup without a person watching each run, so it inherits every gap the earlier stages left open.
Skipping a stage tends to fail in recognisable ways. Without a shared harness, agents cannot build and test from a clean checkout, and the factory produces plausible changes that nobody can verify. Without measurement, nobody can say whether factory changes cost less than an engineer's or cause more rework, so the line either runs unchecked or is switched off after its first bad week.
The 2025 DORA report, announced by Google Cloud in September 2025, still found a negative relationship between AI adoption and software delivery stability. Its authors write that without strong automated testing, mature version control and fast feedback loops, a higher volume of changes leads to instability. A factory raises the volume of changes by design, so those controls need to be in place before it starts.
Building blocks of an AI software factory
In an agentic software factory, the agent is the smallest component. Every task passes through the same stations in the same order, and each station needs an owner on your team.
- 1
Intake with a written spec
A task arrives as an issue with a spec: expected behaviour, files in scope, how to verify it and what is out of scope. Code checks eligibility from labels, paths and task type, and issues without a spec stay with people.
- 2
Queue and orchestration
A queue holds eligible tasks. The orchestrator starts runs, retries within fixed limits and stops any run that exceeds its time or budget.
- 3
An isolated environment per run
Each run gets a fresh workspace, credentials scoped to the task and only the network access it needs. Nothing carries over from one run to the next.
- 4
The harness
Instructions, build and test commands, tools and permissions decide what the agent sees and what it may do. In unattended runs, permissions do the job that a watching engineer would otherwise do.
- 5
Verification gates
Tests, task-specific evaluations and a review by a separate agent session in a fresh context check the exact candidate commit. That reviewer gets the spec, the diff and the test output, without the conversation in which the change was written.
- 6
Merge policy in deterministic code
A program compares the reviewed commit with the candidate, checks the required results and either asks a person to approve or, for a task type you have cleared, merges. The agent has no part in this decision. On GitHub, branch protection can require approving reviews, passing status checks on an up-to-date branch and a merge queue, and can apply to administrators as well.
- 7
Feedback into specs and harness
Every rejected or failed run gets a recorded cause, such as an unclear spec, missing context, a weak test or a wrong permission. The fix goes into the spec template, the harness or a new check.
Cost controls, observability and audit
- Caps per run, per task type and per day, enforced outside the agent
- Cost per accepted change, counting retries, abandoned runs and review time
- A log of every run: inputs, commands, tool calls, results and the policy decision
- An audit trail from each merge back to its issue, spec, checks and approver
# Illustrative only: one task type, not a specific tool's format task_type: dependency-patch-update eligibility: labels: [factory, dependencies] allowed_paths: [package.json, pnpm-lock.yaml] blocked_paths: [.github/, infra/, AGENTS.md] max_files_changed: 4 runs: max_parallel: 2 timeout_minutes: 30 budget_per_run: from measured baseline gates: - build_and_unit_tests - integration_tests - fresh_context_review - reviewed_commit_equals_candidate merge: mode: human_approval # the owner may switch to checks_only later approvers: [payments-maintainers] on_failure: record_cause_and_requeue_once
Parallel runs and packaged workflows
A factory gets most of its throughput from running several agents at once, and that is also where cost and review load grow fastest. Parallel runs help when work splits cleanly: independent tasks from a queue, or several attempts at one task where you keep the best. They help little when the parts depend on each other, and with weak gates every extra run becomes extra review.
Packaged workflows bring these patterns into the tools engineers already use. pstack, a Cursor plugin by Lauren Tan published in the cursor/plugins repository under the MIT licence, bundles several of them. Its skills can run several attempts at the same task and graft the strongest parts into one base, fan work out to parallel workers that return a single report, or have reviewers on different models attack a diff without changing it. A setup skill writes a rule that assigns a model to each role. One playbook takes independent pull requests through to merge, with one agent owning each pull request and merging only after a clean verdict from parallel verifiers and an explicit grant from the person running it.
All of these are instructions to the agent and enforce nothing on their own. Whether anything merges without a person depends on your branch protection, required approvals, CI and the credentials the agent holds. The number of parallel workers is set per task or derived from the work, so limits on parallel runs and spend have to come from your setup.
- A cap on concurrent runs per repository and per team
- A model per role, chosen from results on your own tasks
- Cost per accepted change for each workflow, compared with a single-agent run
- A limit on how many candidate changes a reviewer is asked to read
pstack workflowspstack on CursorCosts of multi-agent workflows
Which work fits first
Start with task types that occur often, are well specified and can be checked by machines. A wrong result should be cheap to detect and cheap to revert.
- Dependency updates, with release notes read and the full test suite run
- Mechanical migrations, such as a renamed API or a changed configuration format
- Test coverage for code whose behaviour is already agreed
- Well-specified maintenance: lint findings, deprecation warnings, small bugs with a failing test
Other work stays with engineers, who can still use agents interactively. Ambiguous product work needs someone who discovers the requirement while building it. Security-critical changes to authentication, authorisation, cryptography or secrets handling need a person who accepts the risk. Architecture decisions shape everything the factory does afterwards, so they belong with the people who will maintain the result.
Readiness checklist and what to measure
Before the first unattended run, we check the repository the factory will work in. Every item below should hold.
- A clean checkout builds and tests with one command, in the environment the agent uses
- Tests are fast and stable enough that a failure means something
- Harness files sit in version control with an owner, and changes to them are reviewed
- Issues of the chosen task type follow a spec template
- Branch protection and required checks apply to everyone, including administrators and bot accounts
- Spend is attributed per run and capped outside the agent
- A named person can pause the line and knows how
What to measure
- Accepted changes per task type, and the share of runs that end in one
- Review time per accepted change
- Rework and reverts within an agreed window after merge
- Escaped defects traced back to factory changes
- Cost per accepted change, including retries, abandoned runs and review time
Keep tracking the delivery metrics your team already uses, such as DORA's change fail rate and deployment rework rate. If throughput rises while those get worse, the factory is moving problems downstream.
Start with one repository, one task type and one queue
- 1
Choose the task type
Pick one from the list above that occurs often in a single repository. Write its spec template and its eligibility rules.
- 2
Run it supervised
Let the factory open pull requests while people review every one as usual. Record why each run was accepted, reworked or rejected.
- 3
Fix the causes
Feed each rejection back into the spec template, the harness or a check. Rerun the same tasks to confirm the fix.
- 4
Decide the merge policy
Once the numbers hold over an agreed period, the owner decides whether this task type may merge after checks or keeps a human approval. Either way, the decision is written in code.
- 5
Add the next task type
Only then add a second task type or a second repository, each with its own caps and measurements.
Runmill, a public project by our founder, follows the rule we apply with clients: the agent proposes, and deterministic checks and people decide. It takes an eligible Linear issue and runs Claude Code or Codex on it in an isolated workspace. The exact candidate commit has to pass the required checks and a review in a fresh context before a GitHub pull request is opened. Pushes, pull requests and merges are decided by deterministic code, never by the agent, and automatic merge is experimental in this developer preview.
How we help
We work in your repositories with the engineers who will own the factory. First we check where your team stands on the path, because a factory built on an unmeasured setup inherits its gaps. Then we set up one task type end to end with your platform team: spec template, harness, isolated runs, gates, merge policy in code, cost caps and measurement.
What your team keeps
- Spec templates and eligibility rules for each task type
- Harness files and the merge policy in version control, each with an owner
- Runbooks for pausing the line, changing caps and adding a task type
- Measurements per task type, including cost per accepted change
Workshops run remotely or on-site in Germany and Poland, in English, German or Polish, as team sessions on your own repositories. Afterwards your platform team can run the factory and add task types without us.
Harness engineeringEvaluations on your repositoriesLLM cost optimisation
Questions
Can an autonomous software factory merge without human review?
Technically yes, for task types where your team decides that the checks are enough. We start with a person approving every merge and relax that only for a task type whose results hold over an agreed period, with spot checks continuing afterwards. Changes to the harness, the gates and the merge policy stay with people.
How does this differ from a DevSecOps software factory?
In the US Department of Defense sense, a software factory is the people, tools and pipelines that build, test and deliver software continuously. An autonomous software factory adds coding agents that write the changes, and it needs that kind of pipeline underneath.
Which coding agents and models can we use?
The building blocks do not depend on one vendor. We work with the tools your teams already use, such as Claude Code, OpenAI Codex, Cursor or GitHub Copilot, and check what each supports for unattended runs before we agree the scope.
How do we know whether the factory pays off?
Compare cost per accepted change and review time with the same task type done by engineers working with agents, and watch rework, reverts and escaped defects. We do not promise savings. If a task type does not hold up, it goes back to people.
We are not using coding agents yet. Where do we start?
With agentic coding in one or two teams, a shared harness and a budget. The factory comes later, once your setup is measured on your own code.
Related services
Harness engineering
The instructions, commands, permissions and checks that decide how an agent behaves in your repositories.
Evaluations on your repositories
Tools, models and settings compared on work your team has already merged.
LLM cost optimisation
Where the money goes, what an accepted change really costs, and the defaults, caps and routing that bring it down.
More from the blog
AI SDLC
How coding agents change each phase of the software development lifecycle, who owns each handoff, and how to measure and govern the change.
Spec-driven development
How to write specs coding agents can follow, take one change from spec to merge, and choose between Spec Kit, OpenSpec, BMad, GSD and Kiro.
AI code review
How to set up AI code review and keep a person accountable for every merge when coding agents write more of your pull requests.
Sources
- pstack README in the cursor/plugins repository (Lauren Tan, MIT licence)
- DoD Enterprise DevSecOps Fundamentals, version 2.5 (US Department of Defense CIO, 2024)
- Japanese Cooperative R&D Projects in Software Technology (Michael A. Cusumano, MIT Sloan working paper, 1989)
- Software Factories: the Smartest Way to Outsource Software Development (MJV Technology & Innovation)
- The Five Levels: from Spicy Autocomplete to the Dark Factory (Dan Shapiro, January 2026)
- Software Factories And The Agentic Moment (StrongDM, 6 February 2026)
- About protected branches (GitHub Docs)
- DORA's software delivery performance metrics (DORA)
- Announcing the 2025 DORA Report: State of AI-Assisted Software Development (Google Cloud blog, 23 September 2025)
Talk to an engineer
A 30-minute call with one of our engineers about your coding agent setup, what it costs and where it can improve. No access to your systems and no data shared.
Not using coding agents yet? Use the same form and tell us what you are planning.
Our team has built software and production AI for trivago, SAP, Tonies, EWE and tecRacer.
Check your email program
We tried to open a draft in your email program. Nothing is sent until you send it, and this page cannot tell whether the draft opened or the email was delivered. The form stays editable, and you can also write to miki@cloudsail.com directly.