BlogTooling

pstack workflows for Cursor and how a team adopts them

pstack is a Cursor plugin by Lauren Tan that routes engineering tasks to playbooks, assigns models to roles and runs work and review in parallel. We describe what its files contain as of October 2026 and how we would bring it into your team's repositories.

Updated

What pstack is and where it comes from

pstack is a Cursor plugin published in the cursor/plugins repository on GitHub under the MIT licence. Its author, Lauren Tan (poteto), writes in the README that these are the skills used every day to ship code at Cursor, with the aim of writing less code of higher quality. On 10 October 2026 the manifest listed version 0.15.15.

You install it with /add-plugin pstack in a Cursor chat or from the Customize view. Then /setup-pstack chooses the models and /poteto-mode starts the work. According to its help skill, installing alone changes nothing until someone invokes a skill.

  • 27 skills, among them poteto-mode, setup-pstack, how, why, architect, arena, swarm and interrogate
  • 23 playbooks inside poteto-mode, one per kind of task
  • 24 principle skills of one rule each
  • Two subagents: poteto-agent and Comment Sicko, a read-only comment reviewer
  • The dormant benny automation pack for Slack issue reports, which opens draft pull requests only

The README calls poteto-mode the author's own style and invites readers to fork the plugin, and /automate-me drafts a personal mode from your chat history. For a team, most of the work is deciding how much of that style fits yours.

The pstack workflows: one mode and 23 playbooks

The entry point is poteto-mode. It matches the task to one of 23 playbooks, opens a to-do list that starts with that playbook's steps copied in verbatim, and calls other skills as the steps need them. Skipped steps stay on the list with a reason. Large or cross-cutting work goes to the figure-it-out skill, which designs a bespoke playbook and keeps a decision log.

Enter attaches poteto-mode to one message, while Option+Enter or Alt+Enter keeps it on as a Custom Mode. Every code playbook ends with Opening a PR: a git worktree, small ordered commits, a cleaned diff, a Conventional Commits title and a pull request opened ready for review.

  • Understanding: Investigation, Runtime forensics and Trace forensics, which change no code
  • Changing code: Bug fix, Feature, Refactoring, Perf issue, Hillclimb and Visual parity
  • Deciding: Prototype and Multi-phase plan
  • Merging and long runs: Babysit, Shipping, Autopilot-full, Autopilot-stack, Orchestrate and Autonomous run
  • Upkeep: Authoring a skill, Eval, Session pickup, Pause safely and Worktree cleanup

Bug fix, step by step

Bug fix shows the pattern the code playbooks share. The lead agent plans and reviews, while subagents investigate and implement.

  1. 1

    Reproduce

    Reproduce the defect where it occurs, through a control skill, and ask a person only with a stated reason.

  2. 2

    Find the cause

    Rule out hypotheses with runtime evidence until one survives, adding instrumentation where state is unclear.

  3. 3

    Delegate

    Run architect if the fix crosses a function boundary, then hand implementation to a subagent on the bug-fix model.

  4. 4

    Verify

    Rerun the original reproduction on the same surface. An inconclusive result is flagged and does not count as a pass.

  5. 5

    Order the commits

    Land the failing reproduction or test before the fix in history.

  6. 6

    Open the pull request

    Run Opening a PR and paste the failing and passing output.

Feature and refactoring

Feature runs how and then architect, which explores designs in parallel and goes straight on to implementation unless you ask for a checkpoint. When several shapes are valid, delegation through arena is mandatory. Refactoring first pins behaviour with a characterisation test, snapshot or equivalence harness, states that type checks and lint are not a pin, and reverts a change that does not make the code easier to read.

Where a person approves

Most code playbooks have no approval step before the pull request, so pull request review is where your team's judgement comes in. Hillclimb agrees the metric with you, Multi-phase plan waits for the operator's go, and Autopilot-stack hands over a verified stack for the operator to land.

Babysit drives a pull request to merge-ready and never merges. Shipping merges only when asked, after an agent that did not write the code has verified each pull request. Autopilot-full lets owner agents merge once the root agent's verification is clean, under a full-autonomy grant. The mode pauses for force-pushes to shared branches, deploys, data deletion and customer messages, but posts to team chat and updates tickets without asking.

The role rule: which model does what

pstack ships no rule file. /setup-pstack writes one: it detects the models a subagent can use in your session, asks for one of four reasoning budgets from medium to max effort, with xhigh matching the defaults, confirms each role and writes an always-applied rule to ~/.cursor/rules/pstack-models.mdc. It never writes a model it has not confirmed, and the rule applies to new chats.

On 10 October 2026 the defaults sent code delegates, explorers, investigators and swarm workers to grok-4.7-xhigh-fast, and judgement, prose, the hardest changes and synthesis to claude-opus-5-5-xhigh. Each panel entry starts one subagent, and auto or inherit-parent runs a role on the chat's own model.

Excerpt of the role rule written by /setup-pstack, defaults on 10 October 2026
---
description: pstack per-role model choices (overrides skill defaults)
alwaysApply: true
---
# budget: large (xhigh)
feature, refactoring: grok-4.7-xhigh-fast
bug-fix: grok-4.7-xhigh-fast
judgment and prose: claude-opus-5-5-xhigh
hardest tasks: claude-opus-5-5-xhigh
interrogate reviewers: claude-opus-5-5-xhigh, grok-4.7-xhigh-fast
...

Each role is a subagent with its own context. Cursor's documentation puts it plainly: five parallel subagents use roughly five times the tokens of one agent, billed at the list price of each model. pstack's guide lists the levers, which are a smaller budget or cheaper models, auto for some roles, shorter panels, and poteto-mode only where the rigour is needed.

Shorter panels cost something too, because pstack reviews through model diversity. The interrogate skill gives one reviewer per configured model the same diff and treats findings that two models raise independently as the strongest signal, and Orchestrate puts each verifier on a different model family from its worker. A one-model panel saves tokens and loses that check.

For a team, two details matter. The file sits in each engineer's home directory and several skills read it from there, so one plugin revision can run different models on two laptops. The defaults also changed three times between 11 September and 5 October 2026, and the default panels went from four models to two. A rule written before a change keeps the old default, so earlier adopters keep the old panel.

Parallel agents in Cursor: how pstack fans out and checks work

Cursor runs subagents in parallel when the agent sends several Task calls at once. They share the parent's checkout unless each gets its own git worktree or cloud environment, and /in-cloud sends the next task to a cloud subagent. pstack's guide insists on that isolation, because agents in one worktree overwrite each other.

The swarm skill fans N workers out over slices or a declared race, as cloud agents by default, and returns one report. The arena skill gives N candidates the same brief, has a read-only judge score them, then picks a base and grafts in the best parts of the rest. The architect skill runs arena over design sketches, and the read-only reviewers of interrogate sort findings into act on, consider, noted and dismissed without applying anything.

The heavier playbooks multiply this. Shipping verifies each pull request with its own cloud agent, Autopilot-full reruns a swarm after every push that changes the patch, and the Multi-phase plan template asks for ten live verification lanes per pull request. Orchestrate keeps about ten children in flight per sub-coordinator.

These are the plugin's defaults. Before the heavier playbooks run on your repositories, set your own limits:

  • Swarm workers and live lanes per pull request
  • Panel sizes for interrogate, arena and architect
  • Which playbooks may run, and which may merge
  • Reasoning budgets, time boxes and retry caps
  • One writer per worktree or branch
  • Which MCP servers cloud agents reach, with which credentials

What pstack expects from Cursor, other plugins and MCP servers

pstack relies on parts it does not ship. Some are named in the README, and the playbooks and scripts show the rest.

  • Cursor features: subagents, cloud subagents, Custom Modes, /loop and the built-in create-skill
  • The cursor-team-kit plugin: /deslop before commits, control-cli and control-ui for live checks
  • MCP servers, which the why skill maps to seven evidence categories, reporting missing ones as gaps
  • The GitHub CLI or the Origin CLI, Bun for the watcher and orchestration scripts, and Graphite's gt for Orchestrate
  • A verification skill for your app, which /create-verification-skill generates and proves once
  • A Cursor plan with the models you assign, and Teams or Enterprise for a team marketplace

Two of these raise access questions. The why investigators run in agent mode, because read-only mode strips MCP access, and the skill only notes that they should not write. Cloud subagents take their MCP servers from your team's configuration in Cursor, so read-only credentials on those servers are the control that holds.

Forks need one more check. The autopilot playbooks re-read themselves from trunk every hour and the plan playbook runs a checker script, both through paths from the cursor/plugins layout. Those paths resolve in your repository only if the plugin sits at the same place.

Paths the playbooks assume, from the cursor/plugins layout
git show origin/main:pstack/skills/poteto-mode/playbooks/autopilot-full.md
git show origin/main:pstack/skills/poteto-mode/playbooks/autopilot-stack.md
node pstack/skills/poteto-mode/scripts/check-plan.mjs <plan.md>

Adopting pstack in a team

We treat pstack like any dependency that changes how code reaches your main branch. That comes down to five steps.

  1. 1

    Pin a revision in your own fork

    Fork the repository, as the README invites, and name an owner for local changes. Cursor's team marketplaces on Teams and Enterprise plans can import your fork, update when the tracked branch moves and mark the plugin Default Off, Default On or Required. Engineers who install it from there get a new revision only when you move that branch.

  2. 2

    Read the playbooks with your reviewers

    Note where each playbook acts without asking and decide what to leave out. Autopilot-full merges on an agent's verdict, for example, and poteto-mode posts to team chat and updates tickets without asking.

  3. 3

    Wire in your commands and approval points

    Swap generic checks for your build, test and CI commands, generate a verification skill and add stops where your process needs a person. Change conventions that differ from yours, such as ready-for-review pull requests and Conventional Commits titles.

  4. 4

    Choose models and limits from your own results

    Run a scoped set of changes your team has already merged under two or three candidate role rules and panel sizes. Compare correctness, review effort, retries and cost per accepted change, then give everyone the same rule file.

  5. 5

    Rerun the same tasks before every upgrade

    Diff the new revision, including the setup-pstack defaults, and rerun the task set before moving the tracked branch. Where defaults changed, delete stale role lines or rerun /setup-pstack.

What pstack does not enforce

Playbooks, principles and the role rule are text in the model's context. Cursor's documentation says applied rules sit at the start of that context and that AI guidance should not be your only security control. pstack's own principle on shared state agrees: instructions and conventions are not concurrency control.

The limits sit outside the plugin, in the agent's credentials, branch protection, required approvals and required CI checks. Autopilot-full says its root agent never grants or bypasses an approval your Git host enforces, which helps only if your Git host enforces one. Babysit says it never merges, and if the agent's token can merge without review, that sentence is all that stands in the way.

So check the controls before the playbooks: narrow token permissions, required reviews and status checks on protected branches, and no deploy credentials in agent environments. Whether pstack improves your team's results is something only a measurement on your own tasks can show.

How pstack compares with other packaged workflows

Superpowers, from Jesse Vincent and Prime Radiant, moves from brainstorming and a written plan to subagent-driven execution with test-driven development and code review. The compound engineering plugin runs brainstorm, plan, work, simplify and review, then records lessons in docs/solutions. GitHub's Spec Kit goes from a constitution through specification, plan and tasks to implementation. All three are MIT-licensed, and the first two list Cursor among their supported hosts.

pstack picks one of 23 playbooks by kind of task and has no planning skill. In the README, Lauren Tan writes of not believing in planning, calls code the best spec and notes that Cursor's plan mode works well with pstack. Its weight sits on evidence from the running app and on review across several model families, and it depends on Cursor's subagents, Custom Modes and /loop.

How we help teams adopt pstack

We read the pstack revision you intend to use, list what it expects and go through the playbooks with your reviewers. We wire in your commands and approval points, leave out what your controls do not support, and set model assignments and parallel limits from measured results and cost on your own tasks, which then serve as your upgrade check.

pstack workshops run as team sessions on your own repositories, remotely or on-site in Germany and Poland, in English, German or Polish. We do not run public courses or issue certificates.

Questions

Does pstack work outside Cursor?

Partly. Its skills use the Agent Skills format, so other tools can read the files. Its help skill notes that most workflow skills spawn Cursor subagents with per-role models, and that Custom Modes and /loop are Cursor features, so those parts may not work elsewhere.

Can pstack merge pull requests on its own?

Some playbooks can. Babysit stops at merge-ready, Shipping merges only when asked to land or ship, and Autopilot-full lets owner agents merge after a clean verification under a full-autonomy grant. Whether any of that is possible depends on the agent's credentials, branch protection and required approvals, which pstack does not control.

How much will pstack add to our Cursor costs?

It depends on panel sizes, worker counts, reasoning budgets and models, so we do not quote a figure. Each subagent uses its own tokens and bills at the list price of its model. We measure cost per accepted change on your tasks before and after adoption.

Do we have to use all 23 playbooks?

No. According to its help skill, installing pstack changes nothing until someone invokes a skill, and poteto-mode runs only the playbook that matches the task. In your own fork you can remove playbooks your team does not want.

Do you run pstack workshops or training?

Yes, as team sessions on your own repositories, remotely or on-site in Germany and Poland, in English, German or Polish. There are no public courses or certificates.

Talk to an engineer

A 30-minute call with one of our engineers about your coding agent setup, what it costs and where it can improve. No access to your systems and no data shared.

Not using coding agents yet? Use the same form and tell us what you are planning.

Our team has built software and production AI for trivago, SAP, Tonies, EWE and tecRacer.

The button opens a draft in your email program. You send it yourself. Privacy policy (draft)