Skip to content
← All posts

Agent CLIs compared from daily use: Claude Code, Codex, Gemini CLI, OpenCode

An opinionated, practice-based look at four terminal coding agents: what each does well, where each tends to break, and how we split work between them.

JPECOM Team4 min read

People keep asking which coding agent in the terminal is "the best". After months of using four of them side by side, our honest answer is that the question is slightly wrong. They differ in temperament more than in raw ability, and the useful question is which one you hand which kind of job. Everything below is opinion based on our own use. It contains no benchmarks, and your results will vary with your codebase, your prompts and the version you are running.

How we use them

We use these tools to write and change real code, run tests, and prepare changes for review. We do not run formal comparisons. We notice what makes us reach for one tool instead of another, and what makes us swear at one. Tools also change quickly, so treat anything here as a snapshot of habits rather than a verdict.

Claude Code

Opinion: the best at long, multi-step work where the plan matters.

  • Good at: reading a codebase, forming a plan, and carrying it through many steps while keeping track of the original goal. It is good at explaining its reasoning, and its handling of project instruction files is mature.
  • Where it breaks: it can be over-eager to do more than you asked, and long sessions drift if you do not reset context. It can also declare success a little early, which is why we never accept "done" without a check.
  • We reach for it when: the task needs judgement, design decisions, or review of someone else's change.

Codex

Opinion: an efficient, focused implementer.

  • Good at: well-specified, bounded jobs. Give it a clear contract, a list of files it may touch and a test command, and it tends to deliver quickly and stay in its lane.
  • Where it breaks: open-ended scopes. When the brief is vague or the file set is unbounded, it can make sweeping edits. Isolating its work in a separate working copy is a habit we strongly recommend, so a bad run costs you a deleted folder rather than a lost repository.
  • We reach for it when: the spec is tight and the work is mostly typing: new modules, tests, mechanical refactors.

Gemini CLI

Opinion: useful for breadth and for large inputs, less predictable for fine editing.

  • Good at: taking in a lot of material at once, summarising it, and answering questions about large amounts of text or code. Having a generous context window to work with is convenient for exploration.
  • Where it breaks: in our hands it has been less consistent at precise, surgical edits and at following long tool-using sequences to the end. We have also seen confident statements about details that turned out to be wrong, so anything factual that touches a customer needs verification.
  • We reach for it when: we need a second opinion, a broad survey of unfamiliar material, or a quick read of a big document.

OpenCode

Opinion: the flexible option for people who want to choose their own model.

  • Good at: being an open, configurable terminal agent that can talk to many providers. If you want to swap models, try a local one, or avoid being tied to one vendor, that flexibility is the point.
  • Where it breaks: flexibility pushes work onto you. Quality depends heavily on which model you plug in and how you configure it, and you spend more time on setup and tuning than with a tool that ships with a single tuned stack.
  • We reach for it when: we want to experiment with models, work against a local model, or keep an exit route from any single provider.

What breaks in all of them

The failures that matter are shared, and they are worth more than the differences:

  1. Claiming completion without evidence. Every one of these tools will sometimes say a task is done when a test was never run, or ran and failed. Require a command that proves it.
  2. Retry loops. Given a failing step, an agent may keep making small variations of the same attempt. Cap the attempts.
  3. Scope creep. An agent asked to fix one function may "helpfully" restructure a module. Tell it which files are in bounds.
  4. Context bloat. Quality tends to fall as a session grows. Fresh sessions with a short brief beat marathon conversations.
  5. Confident wrong facts. Fine for code you can test, dangerous for claims about the outside world.

How we split the work

Our practical routing looks like this:

  • Planning, design and review go to the tool with the strongest reasoning.
  • Bounded implementation with a clear contract goes to the fastest implementer, in an isolated working copy.
  • Large reading and summarising tasks go to whichever tool handles the biggest input comfortably.
  • Experiments with other models go to the configurable one.

To keep that routing manageable we describe tasks in a tool-neutral way: a goal, a contract, allowed files, and a check command. Any of the four can pick up such a brief. Orchestration tooling, such as our own Tower, exists to run that kind of routing, but the idea works fine with nothing more than a folder of task files.

Takeaways

  • Treat agent CLIs as colleagues with different temperaments, not as a ranking.
  • Use the strongest reasoner for planning and review, a focused implementer for tight specs.
  • Isolate risky runs in a separate working copy.
  • Shared failure modes (false "done", loops, scope creep) matter more than brand differences.
  • Everything here is opinion; test the tools on your own work before committing.

Related posts