Multi-agent orchestration: the mistakes that cost us weeks
Agents that say done without evidence, retry loops, file ownership clashes, and our rules: two failures means stop, measure output not activity.
JPECOM Team5 min read
Running one coding agent is easy. Running several at once looks like a multiplication of productivity and turns out to be a multiplication of failure modes. We lost weeks, not hours, to a handful of mistakes that look obvious in hindsight. This post is a list of them, in the order they hurt, along with the rules we adopted afterwards. None of it is exotic. All of it is the kind of thing you only believe after paying for it once.
Mistake 1: believing "done"
The first and most expensive mistake was treating an agent's report as a fact. An agent says the task is complete, the summary is confident and well written, and the work is marked finished. Later you discover that the tests were never run, that the file was edited in the wrong place, or that the change exists only in a temporary directory that was cleaned up.
Agents are not lying in any deliberate sense. They are producing the most plausible description of a successful outcome, and a plausible description is cheap to write. The fix is to make "done" something a machine can check:
- Every task carries a check command that exits zero only if the work is really finished.
- The orchestrator runs the command itself. It does not ask the agent whether the command passed.
- A task without a check command is a task we are not willing to start.
We wrote more about this in the post on contracts and verification. The short version is: evidence first, narrative second.
Mistake 2: retry loops that look like progress
The second mistake was letting a failed step be retried automatically. It feels sensible. Transient errors exist, and a retry is cheap. But an agent facing a deterministic failure will happily produce attempt after attempt, each slightly different in wording and identical in effect. The log grows, the quota drains, and the dashboard shows a busy worker.
Our rule, which we now apply everywhere, is blunt: two failures without progress means stop.
- The first failure is information. Read the log, and change something.
- If the second attempt fails in the same way, or the output is effectively identical, stop. Do not run a third try of the same thing.
- A stopped task escalates: a stronger model or a person looks at it with the logs, and the approach is changed before anything runs again.
- There is also a hard ceiling after which the task is parked until a human reopens it.
The key phrase is "without progress". A retry that changes the approach and gets further is fine. A retry that changes nothing is a loop. We record what changed between attempts, because writing it down is what exposes the cases where the answer is nothing.
Mistake 3: two agents, one file
The third mistake was letting parallel agents share a working tree. Two agents editing the same file produce interleaved changes that neither understands. Even when they touch different files, one agent's half-finished edit can break the other's tests, which the second agent then tries to fix, producing a third set of changes nobody asked for.
What worked:
- Ownership by path. Each task declares the files or directories it may write. Anything else is read-only for that task.
- Isolated working copies. Each agent gets its own checkout on its own branch. Changes come back through review and merge, not through shared mutation.
- Declare new files in advance. An agent allowed to create files in one folder cannot silently sprawl into another.
- One writer per area at a time. If two tasks need the same area, they run in sequence.
Isolation also limits the damage of a bad run. If an agent deletes the wrong thing, it deletes it in a copy.
Mistake 4: measuring activity
For a long time our dashboards showed which agents were running. They were running. We felt reassured. Then we noticed that a worker could be "active" for a full day and produce nothing: no commits, no files, no results, just a process that stayed alive and logs that scrolled.
Activity is cheap to fake and easy to mistake for work. We moved to measuring output:
- A commit, a merged change, a delivered file, or a completed task with a passing check.
- A daily expectation per worker. A worker with no output in a day is treated as an incident, not as a quiet day.
- Heartbeats and process lists are for diagnosing, never for reporting progress.
An agent that is alive and unproductive is a bigger problem than one that is dead, because nobody is looking at it.
Mistake 5: no one owns the cleanup
Agents spawn processes, tools and temporary files. When a task ends badly, those can outlive it. We had stuck background processes consuming memory long after their tasks were forgotten. Now every launched process has an owner who is responsible for ending it, and a timeout that ends it if the owner does not. Long jobs write a short progress note on a schedule, and a missing note counts as a dead job.
Mistake 6: pipelines nobody could read
Finally, we built orchestration that was too clever to debug. Tasks that spawned tasks, with state hidden in several places. When something went wrong, it took longer to understand the pipeline than to fix the bug. Simple wins: a queue of task files, a status for each, a log for each attempt. If you can read the whole state with a file browser, you can debug it at two in the morning.
What we would tell a team starting today
Start with one agent and a check command. Add a second only when the first is boring. Give each agent its own copy of the code. Cap retries before you need to. Measure what was produced, not what was running. Tooling such as Tower can enforce these rules for you, but the rules themselves are free and work with a shell script and discipline.
Takeaways
- Never accept "done" without a machine-run check.
- Two failures without progress means stop and change the approach.
- Give every task a declared set of files it may write, in an isolated working copy.
- Measure output, not activity; silence for a day is an incident.
- Every process needs an owner and a timeout.
- Keep the orchestration simple enough to read in a file browser.
Related posts
Prompt → contract → verification: how we make agent work checkable
Write the done-criterion first, decide what a machine can check versus what needs a human, and keep evidence receipts so agent work can be trusted.
4 min read
Why we run coding agents on subscriptions, not API keys
Predictable cost, quota as a budget, and the cases where pay-per-token API access is still the right call for coding agents.
5 min read
AI Update, October 2026: Notable Shifts for People Who Build Products
Trends we have been seeing: CLI agents are maturing, subscriptions come in more tiers, local models are worth a look, and verification tooling is moving to the center.
4 min read