Engineering consulting

OpenAI Codex consulting for repositories and teams

For teams putting OpenAI Codex to work on their own code. We get the repository into a state the agent can build and test, settle sandbox, approval and account policy with the people who own them, and work through real backlog tasks with your engineers. Scope follows your problem.

Discuss your setup

Getting the repository ready

An agent that can't run the build or the tests leaves the whole verification job to the reviewer. So the first pass is unglamorous: we check what actually works from a clean checkout inside the sandbox, including the parts that normally depend on a developer's machine, such as private registries, local services, credentials and generated code.

How much of this exists varies by repository. We start where the gap between a green local run and a green agent run is widest.

  • Build, test and lint commands that run from a clean checkout, recorded in AGENTS.md
  • Instructions covering layout, conventions and the definition of done, kept short enough that they still get followed
  • A test target fast and deterministic enough to run on every iteration, with slow or flaky suites named as such
  • Paths the agent leaves alone: generated code, vendored dependencies, migrations, infrastructure

Sandbox, approvals and where the code runs

In a local setup the CLI works against a checkout on the engineer's machine, under their account and inside the sandbox and approval policy you configure. The cloud agent works on a clone in an environment OpenAI hosts, with its own setup steps, secrets and network rules. The IDE extension and the desktop app can be connected differently, so where they run is established per setup, not assumed. A policy written for one doesn't carry over to the other, so we scope each surface you plan to use on its own.

Where prompts and repository context go depends on the provider that is actually configured: OpenAI's models, or another provider or a locally served model where the CLI is set up for one. A local runtime is not by itself a guarantee that nothing leaves the machine. What is sent, under which account and with what retention depends on the working surface, your configuration and your agreement with the provider. We document the flow as configured for each surface, and your data and contract owners check it against that agreement and the provider's current documentation.

  • Sandbox mode and approval policy per repository: what runs unprompted, what asks first, what is off
  • Network access from the sandbox and, where it is on, the registries and services the build needs
  • Account type, model access and who administers them
  • Shared settings in versioned configuration instead of per-laptop defaults
  • Repositories and data in scope, agreed before the pilot

From interactive use to unattended runs

We pilot interactively: a small group, a handful of representative tasks, your normal review and required checks. That shows which kinds of task in this repository come back mergeable and which cost more in review than they save in typing.

Non-interactive runs, in CI or as cloud tasks, are a separate decision with a different failure mode: nobody is watching when a run goes wrong. We propose them for narrow, well-specified work where checks outside the agent can catch a bad change, and where pipeline code, not the agent, decides what gets pushed or merged. Credentials available to such a run are scoped to that job.

  • Interactive first, on tasks sized so a reviewer can read the result in one sitting
  • Unattended runs only behind tests and review the agent can't modify
  • Review burden tracked alongside throughput: rework, reverts and review findings by team, task and model, not by individual engineer

Team enablement and handover

Enablement happens in your repositories, on backlog items. The sessions concentrate on the judgement calls that take practice: how to cut a task so the diff stays reviewable, what context to supply and what belongs in AGENTS.md, when a run has drifted far enough that restarting is cheaper than steering, and how to review a change you didn't write. Teams that also use the cloud agent practise delegation separately, because reviewing a pull request from a run nobody watched is a different job from reviewing a diff you steered. The team finishes by agreeing conventions for instruction files and review, so the setup doesn't depend on who configured it.

Your team keeps the instruction files, the sandbox and approval configuration, the runbooks and any measured comparisons from the pilot. The reasoning behind each choice is written down so your engineers can change the setup without us.

Questions

We already use Claude Code or Cursor. Does adding Codex mean maintaining a second set of instructions?

Partly. AGENTS.md is read by Codex and several other tools, so build commands and conventions can live in one place. Where a tool uses its own file, that file can usually point at the shared one. Permissions, sandboxing and tool access are configured per tool and don't translate one to one. We keep the shared part shared and document the tool-specific part. If you want to know which tool does better on your tasks, repository evaluations give you the method.

What does Codex cost to run for a team?

It depends on the arrangement you are actually on: usage included in a ChatGPT plan, credits on top of it, token-based billing on an Enterprise workspace, usage on your own API key, or a mix. We take the split from your invoices and usage data and relate it to accepted changes, as described under LLM cost optimization. The arrangement also shapes the workflow: a team that runs out of its included usage mid-afternoon works differently from one billed by consumption.

Can we book Codex training on its own?

Yes, if the agent can already build and test the repository reliably. We check that first, because sessions in a repository where the agent can't verify its own work mostly teach workarounds.

Related services

Harness engineering for coding agents

Most of how a coding agent behaves in your repository is decided by its harness: what it is told, what it can run and reach, and what checks its output before a person sees it. We engineer that software runtime with your team so the agent works with your build, tests and review process.

Discuss your setup

Not using coding agents yet, or want to discuss how your team uses them now? Tell us briefly about your team, repositories, tools and the problem you're working on.

Open email draft

Opens a draft in your email app. You send it; the website can't confirm delivery.

Free usage review

Send one month of usage data and get a one-page spend breakdown, at no charge.

Free usage review