Local-first AI tooling: what stays on your machine and what doesn't
A plain description of data custody when you use AI agents: what lives locally, what providers see, why bring-your-own CLI helps, and the honest limits.
JPECOM Team5 min read
"Local-first" is an appealing phrase and an easy one to abuse. It can suggest that nothing leaves your computer, which is almost never true when a cloud model is involved. This post tries to be precise: which parts of an AI-assisted workflow can stay on your own machine, which parts cannot, and what that means for how you design tools and handle data. We are describing a design approach and our own understanding. For any real compliance question, read your provider's current terms rather than a blog post.
What local-first means to us
Local-first, in our sense, is about custody and control, not about isolation:
- Your files, history, settings and task records live on your machine, in formats you can read.
- The tool works without a vendor's server being in the middle of your own data.
- You decide what is sent to a model provider, and you can see what it is.
- If a service disappears, you keep your work.
Notice what this does not say. It does not say the model runs locally. It says the system of record is yours.
What stays on your machine
In a typical agent workflow, these can and should remain local:
- The working tree. Source code, documents and generated outputs sit in folders you own, under version control you run.
- Task definitions and logs. The briefs you write, the record of attempts, the evidence receipts. Plain files make them searchable and portable.
- Credentials. Keys and tokens for your accounts stay in your environment or a local secret store, not in a hosted dashboard you do not control.
- Orchestration state. Which task is running, which failed, which is waiting. A local queue is easier to inspect than a remote one.
- Your configuration. Instruction files, routing rules and preferences.
Local custody also makes experiments safe. You can copy a folder, try something, and delete the copy.
What does not stay local
Here is the honest part. When an agent uses a hosted model, the following leave your machine:
- The prompt, including whatever instructions and context the tool assembles.
- File contents the agent reads, because to reason about code the model must receive it.
- Command output and error messages that are fed back into the conversation.
- Metadata the provider needs to operate the service, such as your account and request timing.
This means the question "does my data leave the machine?" has a precise answer: the parts the agent reads and the output it sees do. Whatever the agent can open can end up in a request. That is why the most important privacy control is deciding what the agent can read in the first place.
What providers do with it
We cannot answer this for you, and we are cautious of anyone who claims a single answer. Retention, use for training, human review and regional processing differ by provider, plan and settings, and they change over time. What we recommend:
- Read the terms for your specific plan, not the marketing page. Consumer, team and API offerings can differ.
- Look for the data controls in your settings and set them deliberately instead of accepting defaults.
- Assume anything sent could be stored for some period, and decide on that basis whether it should be sent.
- Treat secrets as never-send. No keys, passwords or personal customer data in prompts or in files the agent can read.
Bring your own CLI
One design choice that helps with custody is bring your own CLI: the orchestration tool does not embed or proxy a model. It launches the agent command-line tools you already installed and signed in to yourself. The consequences are practical:
- Your login and your plan stay yours. The orchestrator never holds your account credentials.
- You can swap tools without rewriting your workflow.
- What is sent to a provider is whatever your own CLI sends, under your own agreement with that provider, rather than a second hop through another service.
- You can audit the tool's behaviour with the same methods you use for any local program.
This is how we think about Tower, our own orchestration tool: a local program that coordinates CLIs you control. The same pattern is available with a handful of scripts if you prefer.
Practical habits
- Scope what agents can read. Run them in the project folder, not your home directory. Keep sensitive material outside it.
- Use ignore rules for secrets, environment files and exports of customer data.
- Redact before you paste. If you need an agent to look at a log or a record, remove names, contacts and tokens first.
- Separate workspaces. Work involving confidential material should have its own folder and, where possible, its own configuration.
- Keep records locally, so you can answer the question "what did we send, and when?" without asking a vendor.
The honest limits
Local-first is a direction, not a guarantee. A few limits worth stating plainly:
- If you use a cloud model, the cloud sees what you send. Local custody of the rest does not change that.
- Local models can keep prompts on your hardware, but they need capable hardware and are not equal to the largest hosted models for every task. Whether that trade is worth it depends on your data and your tasks.
- A local tool is only as safe as the machine it runs on. Disk encryption, updates and access control still matter.
- Convenience pulls toward syncing everything to a hosted service. Each convenience is a decision about custody; make it knowingly.
Takeaways
- Local-first means your system of record is yours, not that nothing leaves the machine.
- Prompts, files the agent reads, and tool output are sent to hosted models.
- The strongest control is limiting what the agent can read; never expose secrets.
- Bring-your-own CLI keeps accounts and agreements in your hands and your workflow portable.
- Read your provider's current terms for your specific plan before relying on any assumption.
Related posts
AI Update, October 2026: Notable Shifts for People Who Build Products
Trends we have been seeing: CLI agents are maturing, subscriptions come in more tiers, local models are worth a look, and verification tooling is moving to the center.
4 min read
Agent CLIs compared from daily use: Claude Code, Codex, Gemini CLI, OpenCode
An opinionated, practice-based look at four terminal coding agents: what each does well, where each tends to break, and how we split work between them.
4 min read
Prompt → contract → verification: how we make agent work checkable
Write the done-criterion first, decide what a machine can check versus what needs a human, and keep evidence receipts so agent work can be trusted.
4 min read