BlogSecurity
Coding agent security for engineering teams
A coding agent reads your repositories, runs shell commands, calls tools and pushes branches with credentials your team gave it, and it acts on text that other people wrote. This guide to coding agent security covers the risks that follow from that access, and the controls and ownership that limit them.
What a coding agent can do in your systems
Discussions of AI coding security often start with the model. The more useful starting point is what the agent can do on your systems. On a developer machine, a CLI agent typically runs under the engineer's own user account, with access to everything that account can read and to the credentials in environment variables, configuration files and the git credential helper. Depending on the setup, it can:
- Read source code, configuration and local files such as .env or cloud credentials
- Run builds, tests, package installs and any other shell command
- Fetch web pages and call APIs, including package registries
- Use MCP servers that hold their own credentials for trackers, databases or cloud accounts
- Commit, push branches and open pull requests, and in CI act with the pipeline's token
That makes the agent a new kind of identity in your systems. It acts with borrowed rights and often leaves no record of its own. It also takes instructions from any text it reads, including files, issues, web pages and tool output that someone else wrote. We treat it like a new service account with an unusual input channel: its own narrowly scoped credentials, a log of its actions and a named owner for its configuration.
Prompt injection in coding agents
OWASP lists prompt injection first in its Top 10 for LLM Applications 2025. It separates direct injection, where a user's own prompt changes the model's behaviour, from indirect injection, where the model takes in content from external sources such as websites or files. Taking in external content is most of what a coding agent does.
The untrusted inputs include repository files written by contributors, dependency source code, issues and pull request comments, documentation and web pages read during research, and the output of tools and MCP servers. Any of them can contain text that reads like an instruction to the agent, and that text does not have to be visible to a reviewer.
A typical example is a Markdown comment that the rendered page does not show. It tells AI agents that the tests need the deployment credentials and asks them to copy the environment file into the pull request description. A person reading the rendered README sees nothing unusual.
OWASP states that it is unclear whether fool-proof prevention methods for prompt injection exist, and we plan on that basis. Filters, system prompts and better models reduce how often an injection works, so the design question becomes what a successful one can reach. An agent that reads untrusted content, holds secrets and has a way to send data out (a network call, a commit, a pull request description) gives an attacker everything needed. Remove one of these and a successful injection has much less to work with.
Secrets, destructive commands and dependencies
Three further risks apply to every tool. None needs an attacker, though attackers can use all three.
Secrets in the context
Secrets reach an agent's context in ordinary ways: it opens a .env file to understand the configuration, prints environment variables while debugging, or reads a log that contains a token. The secret then goes to the model provider with the next request and can end up in a commit, a pull request description, a session transcript or a request to an external host. Deny rules for secret files help. Keeping production secrets out of the agent's environment helps more.
Destructive commands
An agent that can run shell commands can also run the wrong one: deleting files outside the task, force-pushing over a shared branch, running a migration against the wrong database, or calling a cloud CLI logged in with administrator rights. Most of these are mistakes that nothing in the environment stopped. Approval prompts catch some. The rest needs an environment without production credentials, with protected branches and with backups that someone has actually restored.
Dependencies and invented package names
Agents add dependencies when a task seems to need one, and models sometimes suggest packages that do not exist. A study presented at USENIX Security 2025 found such package hallucinations to be persistent and systemic in the commercial and open-source models it tested. Whoever registers an invented name gets their code installed by anyone who follows the suggestion, with install scripts running under that user's or pipeline's rights. A new dependency deserves the same review as new code, and a registry proxy with an allowlist turns that review into a rule.
MCP servers, plugins and where your data goes
MCP servers widen what an agent can reach. A server for your issue tracker, database or cloud account usually holds its own credentials, and the agent can do whatever those credentials allow. The MCP specification says clients must treat tool annotations as untrusted unless they come from trusted servers, and that there should always be a human in the loop who can deny tool invocations. Tool output is also text in the model's context, so a server that returns tickets or web pages is a prompt injection channel too. Before a server goes on your allowlist, someone should know:
- Who maintains it, and which version you pin
- Which credentials it holds and what they allow
- Whether it returns third-party content to the model
- Who owns it and reviews its updates
Instruction files, skills and plugins from third parties need the same scrutiny. They are text the agent follows with your permissions, and some ship scripts too. Workflow plugins such as pstack, the Cursor plugin by Lauren Tan, bundle skills and playbooks for recurring engineering work, including overnight runs that take pull requests through to merge. The playbooks instruct the agent and enforce nothing: whether a merge can happen is decided by the agent's permissions and credentials, branch protection, required reviews and CI. Claude Code's documentation says the same about its own instruction files, which shape what the agent tries to do without changing what Claude Code allows.
pstack workflows in practiceAdopting pstack with our Cursor consulting
Every agent also sends data out by design: prompts, file contents, command output and tool results go to a model provider, sometimes through a gateway, and some tools add telemetry or session sharing. Retention, training use and processing region depend on the provider, the plan and your contract, and can differ between a tool's CLI, IDE extension and cloud agent. Write the flow down for each surface and have your data and contract owners check it against your agreements. We document it, but we can't make commitments on a vendor's behalf.
Security controls for coding agents
No single control makes an agent safe, and none below claims to. They work as layers that a mistake or a successful injection has to get through before it causes damage.
- Coding agent permissions per repository, set centrally where possible, with routine commands allowed so the remaining approval prompts get read
- Sandboxing, or a separate container or VM, for unattended runs, with a fresh checkout and no host credentials
- A network egress allowlist covering your git host, package registry proxy and model endpoint
- Short-lived credentials per task, scoped to one repository and the actions the task needs
- Secret scanning in pre-commit hooks and CI, plus deny rules for secret files
- A dependency policy: registry proxy, lockfiles and human review of new packages
- Branch protection, required checks and human review that the agent's credentials cannot change
Two records complete the set. One is an audit log of commands, tool calls and pushes that someone can actually find after an incident, and the other is the MCP allowlist described above.
Reviewing pull requests written by agents
Vendor controls and their limits
The main tools ship such controls and document their limits openly. Claude Code has permission modes and allow, ask and deny rules that administrators can enforce through managed settings. Its documentation says the mode that skips permission prompts belongs only in isolated environments such as containers or VMs, and that a deny rule for a shell command is not a security boundary around the program. OpenAI's Codex combines a sandbox mode with an approval policy, and its default workspace-write mode keeps network access off unless you enable it.
GitHub documents a firewall that limits internet access for GitHub Copilot's cloud agent by default. It covers processes the agent starts through its Bash tool, but not MCP servers or configured setup steps, and GitHub notes that sophisticated attacks may bypass it. Before you rely on any boundary, read the limitations section for the tool you use and check which controls your plan includes.
Unattended runs in CI and cloud environments
Unattended runs change the risk. Nobody watches the commands, the input often comes from an issue or pull request that anyone could have written, and the run holds a pipeline token. Two incidents from 2025 show how much depends on the credentials around automation.
In July 2025, AWS reported that an inappropriately scoped GitHub token in the build setup of its Amazon Q Developer extension for VS Code let an attacker commit malicious code to the extension's open-source repository. The code shipped in version 1.84.0 and did not run because of a syntax error. In August 2025, attackers took the npm token of the Nx build system through a GitHub Actions workflow that ran with the target repository's permissions on pull requests and passed their titles to the shell unsanitised. According to the Nx team's postmortem, the malicious versions published on 26 August attempted to use local AI tools such as Claude and Gemini while scanning machines for sensitive data.
Neither case needed a clever attack on a model. Both turned on credentials with more reach than the job required, and the second shows that attackers will try to use the agent CLIs on developer machines. For unattended runs we work to these rules:
- A fresh, isolated environment per run, discarded afterwards
- Network egress limited to an allowlist with the model endpoint and package proxy
- Credentials issued per run, scoped to one repository and expiring with the task
- The agent pushes only to its own branches, and pipeline code decides whether a pull request is opened or merged
- Required checks and human review that the agent's token cannot change or bypass
- No runs with secrets on events that outside contributors can trigger
Runmill, an open-source developer preview from our founder, follows this pattern. The agent works in an isolated workspace, the exact candidate commit has to pass the required checks and a review in a fresh context, and pushes, pull requests and merges are decided by deterministic code, never by the agent.
The autonomous software factoryHarness engineering for delegated cloud environments
Governance, ownership and incident response
Controls drift without owners. Agent configuration, from instruction files and permission settings to hooks, MCP configuration and plugin versions, belongs in version control with a named owner and a change process. Changes to it need that owner's review, including changes the agent proposes itself. Someone also decides which tools, MCP servers and plugins are approved and who may grant exceptions.
When something goes wrong
- 1
Stop and revoke
Stop running sessions and revoke the agent's credentials, including MCP server tokens and any CI tokens within its reach.
- 2
Reconstruct what happened
Use audit logs, session transcripts and git history to establish what the agent read, ran, sent and pushed.
- 3
Rotate and clean up
Rotate every secret that may have entered the context, and revert or quarantine affected branches, packages and build artefacts.
- 4
Fix the control
Change the permission, credential scope or review rule that allowed it, and test the fix against what happened.
Works council and data protection in Germany
This part is practical guidance, not legal advice. Under section 87(1) no. 6 of the German Works Constitution Act (BetrVG), the works council has a right of co-determination over the introduction and use of technical systems designed to monitor employee behaviour or performance, where no statute or collective agreement already settles the matter. Agent usage dashboards, session logs and per-person metrics can raise that question, so involve the works council before the rollout and agree what is logged, who sees it and for how long. We report pilots by team, task and model, never by individual engineer.
Data protection runs alongside. Code, commit history, tickets and logs often contain personal data, from author names to customer records in test fixtures, and the GDPR covers any of it that reaches a model provider. Your data protection officer will want the written data flow, the provider's processing terms, retention and region, and a decision on what may enter an agent's context.
Agent surface: <agent> CLI on developer laptops Model endpoint: <provider, account, region> Gateway: <none, or name and owner> Sent: prompts, file contents, command output, tool results Kept out: .env, secrets/, customer-data/ Retention: <per contract, checked on YYYY-MM-DD> Training use: <per contract and plan settings> Telemetry: <setting and destination> Credentials held: <token, scope, lifetime, owner> MCP servers: <name, pinned version, owner, credential scope> Config owner: <team> Next review: <date>
How Cloudsail helps
We do this work with your engineers and security team, in your own repositories: harness and permissions for each tool you use, written data flows, and isolated environments with scoped credentials for unattended runs. Each change is checked on representative tasks before it is shared across teams.
Workshops run remotely or on-site in Germany and Poland, in English, German or Polish, as sessions with one team on its own repositories. We don't run public courses or issue certificates. Your platform team keeps the configuration, the runbooks and the reasoning behind each setting.
Harness engineering for coding agentsRolling out coding agentsClaude Code permissions and rollout
Questions
Can prompt injection be prevented completely?
No one can promise that today. OWASP states that it is unclear whether fool-proof prevention methods exist, so we design for the case where an injection succeeds. That means limited credentials, no route out for secrets and review the agent cannot bypass.
Is a vendor sandbox enough for unattended runs?
It is one layer. Vendor documentation describes what each sandbox or firewall covers and what it leaves out, such as MCP servers or setup steps. For unattended runs we add an isolated environment, short-lived scoped credentials and review gates outside the agent's control.
Do we need to involve the works council before rolling out coding agents?
In Germany, often yes, wherever agent logs or usage data could show how individual employees behave or perform. Involve the works council early and agree what is logged and who sees it. This is practical guidance, not legal advice.
Will our code be used to train models?
That depends on the provider, your plan and your contract, and it can differ between a tool's CLI, IDE and cloud surfaces. We document the data flow for each surface, and your data and contract owners check it against your agreements. We can't make commitments on a vendor's behalf.
Who should own agent configuration?
Usually the platform team, with security reviewing changes to permissions, credentials and MCP servers. Each repository's agent files also need a named owner, and changes to them go through review like any other code.
Related services
Harness engineering
The instructions, commands, permissions and checks that decide how an agent behaves in your repositories.
Coding agent adoption
Rolling out one tool in one team, or several tools across many, in a form your platform team can run without us.
Claude Code consulting and rollout
Setup around your codebase, permissions and review process, starting with a small pilot.
More from the blog
pstack workflows
What the pstack plugin for Cursor contains, how its playbooks and model roles work, and how a team adopts it under its own controls.
Autonomous software factory
What it takes for coding agents to carry scoped work from issue to merge, which work fits, and where people stay in control.
AI SDLC
How coding agents change each phase of the software development lifecycle, who owns each handoff, and how to measure and govern the change.
Sources
- OWASP Top 10 for LLM Applications 2025: LLM01 Prompt Injection
- Model Context Protocol specification (2026-07-28): Tools
- Claude Code documentation: Configure permissions
- OpenAI Codex documentation: Agent approvals and security
- GitHub Docs: Customizing or disabling the firewall for GitHub Copilot
- AWS security bulletin AWS-2025-015: Amazon Q Developer extension for VS Code
- Nx blog: S1ngularity, what happened, how we responded, what we learned
- Spracklen et al., package hallucinations by code-generating LLMs (USENIX Security 2025)
- pstack, the Cursor plugin by Lauren Tan (cursor/plugins repository)
- Section 87 of the German Works Constitution Act (BetrVG), co-determination rights
Talk to an engineer
A 30-minute call with one of our engineers about your coding agent setup, what it costs and where it can improve. No access to your systems and no data shared.
Not using coding agents yet? Use the same form and tell us what you are planning.
Our team has built software and production AI for trivago, SAP, Tonies, EWE and tecRacer.
Check your email program
We tried to open a draft in your email program. Nothing is sent until you send it, and this page cannot tell whether the draft opened or the email was delivered. The form stays editable, and you can also write to miki@cloudsail.com directly.