The pitch is simple enough. One person, four subscriptions, four different companies holding the conversation history. The alternative is a folder of open-source programs that talk to a model running on the same machine you are typing on.
The four tools
Jan is a desktop front end for local and remote models. It speaks the OpenAI-compatible API, so it can point at Ollama on the same laptop or at a rented GPU elsewhere. MCP server support and web search are what moved local chat from a toy into something usable for daily work.
Vane is the project formerly known as Perplexica. The rebrand came with domain filters, file uploads, media search and custom widgets. It queries the web through a search API and lets a local model read the results, which is the Perplexity pattern without handing your questions to a company.
OpenCode is a terminal coding agent shaped like Claude Code and Codex. It reads a repository, proposes edits and runs commands. Open Design followed the same route for design work.
Open Notebook takes documents, notes and sources and lets a model reason over them, close to the notebook features inside the large assistants.
The subscription arithmetic
The standard price for one of these services is $20 a month, whether that is ChatGPT Plus, Claude Pro, Perplexity Pro or Google AI Pro. Four of them cost $80 a month, or $960 a year, roughly €70 a month before the VAT differences between member states. Team seats are a different order of money; figures quoted for business plans reach $600 a month.
| Item | Per month | Per year |
|---|---|---|
| Four individual subscriptions | $80 | $960 |
| The same in euros, before VAT | ~€70 | ~€840 |
| Local stack, software | $0 | $0 |
| Local stack, electricity estimate | ~€0.40 | ~€4.60 |
What running it costs
The software is free. The compute is not. A reasonable estimate for a laptop on integrated graphics: a local session adding about 25 W for two hours a day comes to roughly 18 kWh a year. At €0.25 per kWh that is under €5. At German household prices near €0.35 per kWh the total stays below €7. These are assumptions, not measurements from a lab, but the order of magnitude is not in dispute.
Set that against €840 a year and the electricity disappears into rounding. The real costs are the hardware and your time. A four-billion-parameter model such as Gemma 4 E4B runs acceptably on integrated graphics. Anything larger wants dedicated VRAM, and 16 GB is a practical floor for comfortable work.
What changed since the summer
Vane is the visible part of a wider shift. Lightweight local models in the four-billion-parameter range now follow strict negative instructions, the kind that used to trip small models, and they do it on phones. XDA reports that a user running Qwen 3.5 on an iPhone 16 handled about 90% of daily prompts on the device without touching the cloud. Dedicated replacements for paid workflow features, OpenCode and Open Design among them, appeared over the same months.
Data, GDPR and the AI Act
Local inference changes the paperwork. If prompts and documents never leave the device, there is no processor to name in a data processing agreement and no transfer to document. For a small EU company that is a genuine reduction in compliance work, not a marketing line.
It does not change the transparency rules. Article 50 of the AI Act has applied since 2 August 2026 and covers what you publish: synthetic media needs machine-readable marking, and chatbots must tell users they are talking to a machine. A voice clip generated on your own laptop and posted publicly is still covered by that.
The same date brought enforcement powers for general-purpose AI models, with the European AI Office and national authorities able to evaluate providers and levy fines. That side of the regulation targets model providers rather than someone running open weights at home. All four tools download in the EU without regional restrictions. Output quality in Czech, German or French depends on the model, not the app; Mistral Medium 3.5 is one European option at $1.50 per million input tokens and $7.50 per million output tokens.
Where the cloud still wins
Frontier API pricing is flat right now. GPT-6.1 Sol, released on 29 September, costs $2 per million input tokens and $10 per million output tokens. Claude Sonnet 5.5 sits at the same rates. Gemini 4 Argon launched on 30 September with an introductory $2/$10 that moves to $4/$20 later. DeepSeek-V4.1-Flash undercuts all of them at $0.30 peak and $0.15 off-peak per million input tokens.
Paying per token is the sensible complement to a local setup. Long agentic coding runs, very large contexts and heavy tool use are where hosted models earn their price, according to the vendors' own positioning and the pricing gap. The open-source replacements also arrive without support contracts. If an update breaks OpenCode, there is no ticket queue to join and no uptime commitment to point at.
A setup that holds up
Ollama plus Jan for writing, translation and summarising. Vane when a search should not be logged by a provider. One paid API key for the tasks the local models cannot finish. We test local models on an RTX 5060 Ti 16 GB in AI Arena, and the hardware ceiling is what decides which of these tools is practical on a given machine. The XDA write-up that prompted this piece describes the same split: small models for routine work, a cloud key for the exceptions.
Do I need a dedicated graphics card?
No, but the ceiling is low. Models around four billion parameters run on integrated graphics and recent phones. Anything from roughly 20 billion parameters upward needs dedicated VRAM, and 16 GB is a practical floor for comfortable work.
Can I mix local and cloud models in one app?
Yes. Jan and similar front ends switch between a local Ollama endpoint and a hosted API by changing the model selection. That makes it straightforward to route confidential documents to the local model and keep a paid key for everything else.
Does running everything locally remove my AI Act obligations?
No. The obligations attach to the role you play, not to where the computation happens. If you publish synthetic media in the EU, the Article 50 marking requirements apply whether the model ran on your laptop or on someone else's server.