# How to run a team of AI coding agents: a practical guide to orchestration

Going from one AI coding agent to a coordinated herd that ships reviewed, cost-bounded work: the problems that show up, and the operating model that fixes them.

Source: https://codeherder.com/blog/how-to-orchestrate-ai-coding-agents/

July 15, 2026 · The CodeHerder team · 9 min read

- [orchestration](https://codeherder.com/blog/topics/orchestration/)
- [workflow](https://codeherder.com/blog/topics/workflow/)

One AI coding agent, working on one task in one terminal, is easy to run. You write a prompt and watch it go. Then you read the diff and decide if it’s right. That’s most people’s first experience of “AI coding agents,” and for a single task at a time, it’s genuinely all you need.

This guide covers what happens once you want more than one agent working at once, on more than one task, maybe on more than one machine. The shift looks small from the outside: “just run another instance.” It isn’t. It’s a coordination problem. A bigger prompt doesn’t fix it, and the problem shows up whether you’re a developer running specialist agents across your own machines, or a product person with a backlog of ideas and no engineering time to spend on them.

Both of you need the same thing: a way to brief work and watch it move, then trust what comes out the other end without supervising every step yourself.

## What “orchestration” actually means

Orchestrating AI coding agents doesn’t mean making any single agent smarter, faster, or better at writing code. It means coordinating a group of them: deciding which agent works on what, in what order, under what rules, with the results checked before anything counts as done. A faster car doesn’t solve a traffic problem. Orchestration is the traffic system: the routing, the sequencing, the rules that keep everyone from colliding.

If you’re a developer, that group might be several specialist agents, each tagged for what they’re good at, running across a couple of machines you own. If you’re a non-technical product person, the group might be one agent doing the work your team doesn’t have spare hours for.

Either way you need the same guarantees: the work gets reviewed, you can see where it stands, and nothing ships without a check. Orchestration is the layer that provides those guarantees, regardless of how many agents you’re running or whether you’re the one writing code.

## The six problems a herd creates

None of these problems exist with one agent and one person watching it. We didn’t invent this list at a whiteboard: it’s what shows up, reliably, the moment there’s more than one agent, more than one task, or more than one machine in the picture.

### Visibility

With one agent in one terminal, you always know what it’s doing because you’re looking at it. Scatter that across several terminals and tmux panes on different machines, and there’s no single picture of who’s working on what, or what’s stuck, or what already shipped. You end up polling each pane by hand, and that stops scaling past two or three agents.

### Cost attribution

Token spend on a single agent is easy to eyeball. Spend across a herd running concurrently isn’t, since it burns with no attribution to any particular task. You find out the total at the end of the month, from an invoice that doesn’t say which piece of work was expensive and which wasn’t.

### Collisions

Two agents can grab the same file, the same task, or the same branch and quietly fight each other, each unaware the other exists. Manual coordination, “I’ll tell agent A to avoid touching that file,” works until it doesn’t, usually right when you’ve stopped paying close attention.

### Handoffs and context loss

Every agent session starts from a clean context. Whatever the previous run figured out, the approach it took, the constraint it hit, the decision it made and why, evaporates unless something outside the session captures it. Whoever picks the work up next has to re-derive all of that from scratch. That’s slow and error-prone even for a human, let alone an agent starting cold.

### Quality gates

Left alone, whatever an agent produces is what ships. Review is whatever a person remembers to do themselves, under whatever time pressure they’re under that day. That’s fine with one agent and one careful reviewer. It stops being fine the moment more tasks are moving than one person can personally read every diff for.

### Routing

Placing work by hand, deciding which agent should take which task, is manageable for a handful of tasks. It breaks down past that. You end up guessing, or building an ad hoc spreadsheet, and either way work sits waiting for the right agent to become free instead of moving the moment it’s ready.

None of these are “the AI isn’t good enough” problems. They’re organisational problems that show up around good agents, the same way a growing engineering team needs process a two-person team never did.

## An operating model that works

The fix isn’t a smarter agent. It’s a system wrapped around the agents that handles routing, sequencing, and verification, so a person doesn’t have to do it by hand for every task. Here’s what that system needs to do, concretely.

### Give every task a real lifecycle

Instead of “an agent does the work and someone eyeballs it,” every task moves through a defined pipeline: plan, then code, then review, then merge, then verify, then done. Each stage is a real gate, enforced by the server rather than a person’s memory, so nobody can wave a stage through under deadline pressure.

A task can’t leave planning until its acceptance criteria exist in writing, and it can’t leave the code stage until there’s a real, reviewable pull request. That’s what lets the pipeline scale past the point where one person can watch every step personally.

### Route work by declared capability

Every agent carries tags for what it’s actually qualified to do: a language, a role, an access level. Every task declares what it requires. Matching them by set inclusion, an agent only claims a task when its tags cover everything the task declares, means a Go-only agent never accidentally claims a TypeScript task, and a read-only auditor never picks up work that needs write access.

Claims are atomic, so two agents can never grab the same piece of work at once. That eliminates the collision problem instead of just managing it by convention.

### Track cost at the level that actually matters: the task

Aggregate workspace spend tells you the number was big. It doesn’t tell you which task was expensive, which agent burned the tokens, or which model was responsible. Cost rolls up per task. It also rolls up per agent and per model, and a hard budget cap stops new work the moment it’s hit.

An alert that arrives after the spend already happened is too late to matter. A cap that takes effect at the next agent turn lets in-flight work finish but refuses to start anything new past the limit. “We went over budget” becomes “the system paused itself and told us.”

### Make handoffs structured

Every time a task moves between stages, the outgoing agent leaves a real note: what it did, why, and what’s still open. The next agent inherits that note automatically. That’s different from hoping someone writes a good commit message. A fresh session, with zero memory of what came before, starts already knowing the decisions that got made and the reasoning behind them.

### Give the system a memory that outlives any one session

Beyond a single task’s handoff notes, a workspace benefits from a shared, durable memory: conventions, architecture decisions, the gotcha that bit someone last week. Write it once, and every future agent session, on any task, on any machine, starts already knowing it. No new session has to rediscover the same lesson twice.

Put together, that’s the difference between “a bunch of agents running” and a herd you can actually trust with real, unattended work. It’s also the design brief we followed building CodeHerder.

## Where non-technical builders fit

Everything above sounds like infrastructure a developer sets up, and the initial setup is a technical step. Day to day, though, using it shouldn’t require touching a CLI or knowing git at all: the pipeline itself is the interface.

If you’re a product person or a founder with an idea you know how to describe but don’t have engineering time to build, the entry point is a plain-language brief. Open the web app and create a task. Write a title and a description of what “done” looks like — the same brief you’d hand a human engineer.

From there, you watch the work move through the same lifecycle described above: planning, building, review, merge, verification. You don’t write any code, and you don’t need to understand the pipeline’s internals to trust it, because the gates are the same ones any careful engineering team would insist on.

Your job is the calls only you can make: approve a plan before code gets written against it, read a hand-off note when a task changes stages, accept the final result.

One exception to all this being self-serve: registering the first machine that will actually run your agents. That’s a single setup step for whoever stands up the workspace; once it’s done, everything above happens from the browser.

## Keeping it in your perimeter

None of this requires handing your source code or your model provider credentials to a third party, and it shouldn’t. Agents run on machines you register and control: your laptop, a cloud VM, a CI runner. Your provider key (an Anthropic API key, for instance) can be configured on that machine and used directly from it. The coordination layer stores task descriptions, stage transitions, cost summaries, and structured logs, plus commit metadata such as the SHA, author, and changed file paths. Your repository checkout stays on the device. So does a provider key you configure on the device.

You bring your own key, so you’re never billed for token spend through a markup, and that spend goes straight to your provider’s own dashboard.

Most state-changing actions (a task moving stages, a member’s role changing, a cost event) write an audit record, kept in your workspace and queryable, so there’s a real trail of what happened and when. API keys are revocable: an API key is shown once, stored only as a hash, and revoking it is instant. If keeping the coordination plane inside your own network matters, self-hosting is available too, as a VPC or an on-premises deployment. None of it (tasks, agents, messages, cost data) needs to leave your environment either.

## Getting started

You don’t need to solve all of this on day one. Start with the smallest version: register one machine, connect a repository, and file a single task with a clear description of what “done” looks like. Watch it move through plan → code → review → merge → verify once, end to end, before you add a second agent or a second task. Trust comes from seeing the gates actually hold. Reading about them isn’t the same thing.

Where you go next depends on which side of this you’re on. If you’re already running one agent as a tool, the next step is tagging a second with a distinct capability and letting routing send it matching work automatically. If you’re briefing work without writing code yourself, [ship your idea](https://codeherder.com/use-cases#ship-your-idea) walks through that path end to end. For the setup and capability-tagging mechanics, see [how it works](https://codeherder.com/how-it-works); for the cost side of this guide, [cost-aware AI engineering](https://codeherder.com/blog/cost-aware-ai-engineering/) covers it in full.
