Skip to content
← All posts

Why we run coding agents on subscriptions, not API keys

Predictable cost, quota as a budget, and the cases where pay-per-token API access is still the right call for coding agents.

JPECOM Team5 min read

The first month we ran coding agents seriously, the surprise was not the quality of the output. It was the bill. An agent that loops, re-reads a large repository on every turn, or retries a failing test fifty times does not feel expensive while it runs, and then it arrives as a number on an invoice. This post explains why our default is a flat subscription instead of an API key, how we treat quota as a budget, and where we think the API is still the better tool.

The problem with metered cost and autonomous loops

Pay-per-token pricing is a fine model when a human sits between every request. A person asks a question, reads the answer, decides what to do next. The human is the rate limiter.

An agent removes that rate limiter. It reads files, runs commands, reads the output, and decides to try again, all without waiting for anyone. Three things then compound:

  • Context re-reading. Every turn carries the conversation so far, plus tool definitions, plus instruction files. The same tokens are billed again and again.
  • Retries. A flaky test or an ambiguous instruction can produce a long chain of near-identical attempts.
  • Parallelism. Once you run more than one agent, cost scales with the number of agents, not with the amount of useful work.

None of these are bugs. They are what an agent is. The question is whether you want your cost to be an open-ended function of how well-behaved your prompts are on a given night.

What a subscription changes

A flat subscription turns an unbounded variable into a bounded one. You still can run out, but you run out of allowance, not out of money you did not plan to spend. That difference changes how a team behaves:

  • Experimentation is cheap in feeling and in fact. Nobody hesitates to try a prompt because "it might cost a lot".
  • The worst case is a pause, not a surprise charge. A paused agent is an annoyance; a surprise charge is a finance conversation.
  • Budgeting becomes a planning problem you can solve in advance, instead of a monitoring problem you solve after the fact.

This is our experience, and it depends on the plans your provider offers and the terms attached to them. Check the current terms for your own tool before building a workflow on top of one; plans and limits change.

Quota is a budget: treat it like one

Once cost is bounded, the scarce resource becomes quota. We manage it the way you would manage any small budget:

  1. Know what the allowance looks like. Most plans have a short rolling window and a longer one. Learn both, because the long one is the one that bites at the end of the week.
  2. Spend on judgement, not on typing. The expensive, high-quality model is for planning, review and decisions. Mechanical work such as renaming, formatting or boilerplate can go to cheaper models or to a different tool.
  3. Keep context small. Short instruction files, narrow task scopes, and fresh sessions instead of one endless conversation. This is the single biggest lever we have found, and it costs nothing.
  4. Put a ceiling on loops. Every automated retry loop gets a maximum number of attempts. If it has not worked after a small number of tries, a human or a stronger model looks at it. More on this in our post on orchestration mistakes.
  5. Look at the meter regularly. Not obsessively, but on a schedule. A weekly glance is enough to notice that one project is eating the allowance.

The mindset shift is simple: a token you did not need to spend is a token available for the task that matters.

When the API is the right answer

We are not against API keys. They are the correct choice in several situations:

  • Products that serve other people. If your application calls a model on behalf of customers, you need an interface designed for that, with usage you can meter and attribute. A personal subscription is meant for personal use of a tool, not as the backend of a service.
  • Automation without a person's login. Scheduled jobs, CI pipelines and servers need credentials that belong to a system, not to someone's account.
  • Bursty or very low volume use. If you use a model a few times a month, a monthly plan is waste and pay-per-use is cheaper.
  • Strict data and compliance requirements. Some organisations need contractual terms, regional processing or audit features that only the API or enterprise offerings provide.
  • Fine-grained control. Choosing the exact model, setting parameters, and caching prompts deliberately are API-level capabilities.

A useful rule of thumb: if a human is the one working with the agent, a subscription is usually the better fit; if a program is the one calling the model, you probably want an API.

A note on the risk of depending on one plan

Building a workflow around a subscription means depending on terms you do not control. Limits can change, tiers can be renamed, and a feature you rely on can move to a more expensive plan. We mitigate that in two ways: we keep our task definitions tool-agnostic so work can move between agent CLIs, and we keep a small fallback path to a different provider for the day a plan changes underneath us. The portability is worth more than any single discount.

Takeaways

  • Autonomous agents remove the human rate limiter, so metered cost can grow faster than useful work.
  • A flat subscription bounds the worst case and makes experimentation feel safe, though terms vary and can change.
  • Manage quota like a budget: small context, capped loops, expensive models for judgement only.
  • Use an API when software, not a person, is the caller, or when you need compliance, control or very low volume.
  • Keep your workflow portable so a change of plan is an inconvenience, not an outage.

Related posts