Agent CLIs compared from daily use: Claude Code, Codex, Gemini CLI, OpenCode
An opinionated, practice-based look at four terminal coding agents: what each does well, where each tends to break, and how we split work between them.
JPECOM Team4 min read
People keep asking which coding agent in the terminal is "the best". After months of using four of them side by side, our honest answer is that the question is slightly wrong. They differ in temperament more than in raw ability, and the useful question is which one you hand which kind of job. Everything below is opinion based on our own use. It contains no benchmarks, and your results will vary with your codebase, your prompts and the version you are running.
How we use them
We use these tools to write and change real code, run tests, and prepare changes for review. We do not run formal comparisons. We notice what makes us reach for one tool instead of another, and what makes us swear at one. Tools also change quickly, so treat anything here as a snapshot of habits rather than a verdict.
Claude Code
Opinion: the best at long, multi-step work where the plan matters.
- Good at: reading a codebase, forming a plan, and carrying it through many steps while keeping track of the original goal. It is good at explaining its reasoning, and its handling of project instruction files is mature.
- Where it breaks: it can be over-eager to do more than you asked, and long sessions drift if you do not reset context. It can also declare success a little early, which is why we never accept "done" without a check.
- We reach for it when: the task needs judgement, design decisions, or review of someone else's change.
Codex
Opinion: an efficient, focused implementer.
- Good at: well-specified, bounded jobs. Give it a clear contract, a list of files it may touch and a test command, and it tends to deliver quickly and stay in its lane.
- Where it breaks: open-ended scopes. When the brief is vague or the file set is unbounded, it can make sweeping edits. Isolating its work in a separate working copy is a habit we strongly recommend, so a bad run costs you a deleted folder rather than a lost repository.
- We reach for it when: the spec is tight and the work is mostly typing: new modules, tests, mechanical refactors.
Gemini CLI
Opinion: useful for breadth and for large inputs, less predictable for fine editing.
- Good at: taking in a lot of material at once, summarising it, and answering questions about large amounts of text or code. Having a generous context window to work with is convenient for exploration.
- Where it breaks: in our hands it has been less consistent at precise, surgical edits and at following long tool-using sequences to the end. We have also seen confident statements about details that turned out to be wrong, so anything factual that touches a customer needs verification.
- We reach for it when: we need a second opinion, a broad survey of unfamiliar material, or a quick read of a big document.
OpenCode
Opinion: the flexible option for people who want to choose their own model.
- Good at: being an open, configurable terminal agent that can talk to many providers. If you want to swap models, try a local one, or avoid being tied to one vendor, that flexibility is the point.
- Where it breaks: flexibility pushes work onto you. Quality depends heavily on which model you plug in and how you configure it, and you spend more time on setup and tuning than with a tool that ships with a single tuned stack.
- We reach for it when: we want to experiment with models, work against a local model, or keep an exit route from any single provider.
What breaks in all of them
The failures that matter are shared, and they are worth more than the differences:
- Claiming completion without evidence. Every one of these tools will sometimes say a task is done when a test was never run, or ran and failed. Require a command that proves it.
- Retry loops. Given a failing step, an agent may keep making small variations of the same attempt. Cap the attempts.
- Scope creep. An agent asked to fix one function may "helpfully" restructure a module. Tell it which files are in bounds.
- Context bloat. Quality tends to fall as a session grows. Fresh sessions with a short brief beat marathon conversations.
- Confident wrong facts. Fine for code you can test, dangerous for claims about the outside world.
How we split the work
Our practical routing looks like this:
- Planning, design and review go to the tool with the strongest reasoning.
- Bounded implementation with a clear contract goes to the fastest implementer, in an isolated working copy.
- Large reading and summarising tasks go to whichever tool handles the biggest input comfortably.
- Experiments with other models go to the configurable one.
To keep that routing manageable we describe tasks in a tool-neutral way: a goal, a contract, allowed files, and a check command. Any of the four can pick up such a brief. Orchestration tooling, such as our own Tower, exists to run that kind of routing, but the idea works fine with nothing more than a folder of task files.
Takeaways
- Treat agent CLIs as colleagues with different temperaments, not as a ranking.
- Use the strongest reasoner for planning and review, a focused implementer for tight specs.
- Isolate risky runs in a separate working copy.
- Shared failure modes (false "done", loops, scope creep) matter more than brand differences.
- Everything here is opinion; test the tools on your own work before committing.
Related posts
AI Update, October 2026: Notable Shifts for People Who Build Products
Trends we have been seeing: CLI agents are maturing, subscriptions come in more tiers, local models are worth a look, and verification tooling is moving to the center.
4 min read
Local-first AI tooling: what stays on your machine and what doesn't
A plain description of data custody when you use AI agents: what lives locally, what providers see, why bring-your-own CLI helps, and the honest limits.
5 min read
Prompt → contract → verification: how we make agent work checkable
Write the done-criterion first, decide what a machine can check versus what needs a human, and keep evidence receipts so agent work can be trusted.
4 min read