Skip to main content

One weekend of AI coding can cost $340. Flat $20 plans cannot hold

Ilustrační obrázek
The top tier of frontier coding models has converged on $2 per million input tokens and $10 per million output. GPT-6.1 Sol and Claude Sonnet 5.5 list there, as does Gemini 4 Argon at its introductory rate. It is not the whole market — Grok 4.7 sits at $2/$6 and DeepSeek runs far cheaper — but it is where flagship agent products route by default. That is cheap for a chat window. It is expensive for an agent that reads a repository, edits a file, runs the tests, reads the failure and starts over. Cursor and Replit have already moved away from unlimited flat plans. The arithmetic below shows why the caps appeared.

Three months of repricing

The shift is not subtle. Cursor and Replit have both moved from unlimited flat-rate tiers toward credit pools, hard caps and metered overage fees, as startupfortune.com describes in its analysis of agent subscription pricing. The same pattern shows up across serious agent products: a base allowance of credits, then a per-token or per-request price once the allowance runs out.

The reason sits in the model bills. On 29 and 30 September, OpenAI shipped GPT-6.1 Sol and Google released Gemini 4 Argon. Both list at $2.00 per million input tokens and $10.00 per million output tokens, matching Claude Sonnet 5.5 two days earlier. Gemini 4 Argon's figure is an introductory rate, with $4.00 and $20.00 as the standard price.

What a million tokens costs today

ModelInput, USD per 1MOutput, USD per 1M
DeepSeek-V4.1-Flash$0.15 off-peak / $0.30 peak$0.60 off-peak / $1.20 peak
Meta Muse Spark 1.3$1.25$4.25
GLM-5.3 (Z.ai)$1.40$4.40
Mistral Medium 3.5$1.50$7.50
Grok 4.7$2.00$6.00
GPT-6.1 Sol$2.00$10.00
Claude Sonnet 5.5$2.00$10.00
Gemini 4 Argon, introductory$2.00$10.00
Gemini 4 Argon, standard$4.00$20.00

Input prices span $0.15 to $4.00 per million tokens, a range of more than 25x. Output prices stretch further. A vendor that lets you pick the model can move its own cost by an order of magnitude without changing the feature list. The figures come from OpenAI, Anthropic and Google pricing pages. Convert to EUR at the rate on the day you sign, because a $20 seat lands near €18 before VAT at recent rates.

The arithmetic of a $20 seat

Take frontier pricing, since that is what most agent products route to by default. A single agentic coding task that pulls in 100,000 tokens of context and writes back 15,000 tokens costs $0.20 in input and $0.15 in output, so $0.35. A developer who triggers twenty of those a day for 22 working days spends about $154 a month. The subscription is $20.

That is before retries. Agents fail and rerun, and each rerun pays for the full context again.

The weekend that cost $340

The source case is blunt: one engineer's agent burned $340 worth of API tokens across a single weekend. Repeat that pattern every weekend and you approach $1,400 a month for one seat.

Now price a customer pool. A thousand subscribers at $20 bring in $20,000 a month. If 50 of them behave like that heavy user, the vendor pays roughly $68,000 in model bills and still carries support, hosting, salaries and payment fees. The 950 light users cannot close the gap. Light use might cost $2 or $3 a month, so total model spend lands near $70,000 against $20,000 of revenue.

The tail of the usage curve

Traditional SaaS has near-zero marginal cost per extra seat. An AI agent does not. Every reasoning step, tool call and long context window is a direct purchase from a model provider. Usage distributions for coding agents have a long tail: most users sit in a narrow band and a small group runs far above it. The spread is visible in the numbers above: a light user might spend $2–$3 a month, while one heavy weekend can burn $340 in API tokens.

Flat pricing works when the tail is thin. Here it is not.

What European buyers should check

VAT and the real EUR price. A $20 seat is not a €20 seat. Digital subscriptions carry national VAT across the EU, currently from around 17% to 27% depending on the country. Businesses usually recover it, private customers do not. Add it before comparing plans.

Where the inference runs. Coding agents send source code to a model endpoint. If that endpoint sits outside the EEA, the transfer needs a legal basis, typically a data processing agreement with standard contractual clauses. Ask which provider serves your region and whether an EU-hosted route exists. Model choice changes jurisdiction as well as cost: cheaper routes through DeepSeek, Z.ai or Mistral alter the bill and the legal analysis at the same time.

What the AI Act already requires. Since 2 August 2026, Article 50 transparency obligations apply. Users must be told when they are interacting with an AI system, and synthetic media needs machine-readable marking. GPAI providers face binding supervision and statutory fines from the European AI Office and national authorities under the same date. A coding agent that acts on your behalf is worth checking against those disclosure duties.

Cost controls in the contract. Per-seat caps, spend alerts and the option to pin a cheaper model for routine tasks matter more than the headline price. A plan with no cap is a plan where you carry the vendor's variance.

Self-hosting as a hedge

Open weights remain an option. Llama 4 Scout and Maverick were released under free open weights in April 2025, and quantised builds run on consumer hardware. We keep a 16 GB RTX 5060 Ti in AI Arena for this kind of testing, and a local coding assistant is realistic for completion and small refactors while staying off a metered API. Whether it fits your workflow is a separate question from whether the API bill is convenient.

Does a credit pool plan end up costing more than a flat plan?

Not automatically. Light users often pay less, because unused allowance is no longer bundled into a fixed fee. The people who pay more are the ones who previously consumed far above average. Your own monthly token count is the number that decides it, not the sticker price.

Can I avoid metered billing by running the agent locally?

Partly. Local models remove per-token costs but add hardware and electricity, and capability drops on complex multi-file changes. Many teams split the work: local models for completion and code search, a paid API for the hard tasks.

Are current prices stable enough to budget around?

No. Gemini 4 Argon's introductory $2/$10 becomes $4/$20 at the standard rate, and off-peak discounts such as DeepSeek's $0.15/$0.60 depend on when you run the job. Build the budget on the standard rate rather than the launch rate.

One number worth carrying: at $10 per million output tokens, a $20 monthly plan buys two million output tokens. Anything past that is either subsidised by other subscribers or written into the terms as a cap.

Discussion

No comments yet — be the first to share your thoughts.
X

Don't miss out!

Subscribe for the latest news and updates.