Skip to main content

Agentic AI clocks in: 180% more code, only 30% more releases

Ilustrační obrázek
Agentic AI has stopped being a demo and started showing up on the timesheet. Platforms now push work forward while employees are asleep, with tool access, persistent memory and the ability to act on email, tickets and shared repositories. The first hard numbers are in — and they are not the ones the slide decks promised. McKinsey recorded coding activity up 180% while shipped releases grew just 30%, and roughly a third of companies got less productive after rolling agents out. Meanwhile, since 2 August 2026 the EU AI Act's binding enforcement powers apply. Here is what the measurements actually say, what agents cost per million tokens in euros, and where oversight stops being a PowerPoint slide.

What "clocking in" actually means this week

UC Today's weekly roundup of the three most-asked productivity automation questions lands on a shift that has been building for months: the assistant era is ending. A chatbot answers a question. An agent holds a task, decides the next step, calls a tool, and continues after the human closes the laptop.

The concrete integrations are now shipping. Meta's Muse handles voice, visual and email workflows with open weights under Apache 2.0. Microsoft is steering Copilot toward autonomous background task execution rather than prompt-and-wait. xAI pushed Grok Bot for Enterprise into both Cursor and Grok Enterprise accounts, which effectively means an agent can be handed a repository and a deadline. In our own production stack we already treat long-running scripted agents as infrastructure, not as a feature — and that framing is exactly what most European companies have not yet made.

The 180/30 problem nobody budgeted for

McKinsey's Technology Trends Outlook contains the number that should be pinned above every engineering manager's monitor: AI tools increased software coding activity by 180%, but shipped releases grew only 30%. The gap is not a rounding error. It is 150 percentage points of generated work that has not cleared verification, integration, testing or deployment.

The distribution inside that average is even less comfortable. About 80% of engineers using AI tools see roughly a 3% productivity acceleration. The top 20% see 55%. And around 30% of companies reported an outright productivity drop after teams adopted agentic tooling.

My read, from running automated pipelines: the bottleneck moved. It used to be writing the code. It is now reviewing it, and reviewing plausible machine output is cognitively harder than writing the first draft yourself. An agent that produces 3× the volume does not produce 3× the trust. When it acts autonomously — sends the customer email, edits the CRM record, merges the branch — you have shifted the review step from "later" to "after the damage".

Shadow AI is the real line item

The Solace / IDC enterprise survey adds the operations side. Among decision-makers at companies with more than $1 billion in revenue, 80% are already investing in or running AI agents. 90% have increased focus on real-time data architecture to support them. But 40% name connecting agents in real time to reliable enterprise data as their primary production obstacle, and 46% say data quality and consistency is the top technical issue.

Translation: agents are easy to buy and hard to feed. The same survey found that companies it classifies as real-time data leaders achieve an average 23% annual gain in speed to market and risk reduction — which is, in practice, the value of not having your agent act on a stale record.

What an agent costs per million tokens — in euros

Agentic loops burn tokens in a way chat never did. A single task can involve dozens of model calls, tool schemas and re-read context. At an assumed rate of roughly $1.08 per euro (check the live rate before you budget), here is how the current API field compares:

Model Input $/M Output $/M Input €/M Output €/M
GLM-5.3-Flash (MIT, open weights)0.150.50~0.14~0.46
DeepSeek-V4.1-Flash (MIT)—1.20 peak / 0.60 off-peak—~1.11 / ~0.56
Gemini 3.8 Flash1.507.50~1.39~6.94
Claude Sonnet 5.52.0010.00~1.85~9.26
Grok 4.72.006.00~1.85~5.56

The spread between the cheapest open-weight option and the most expensive frontier model here is roughly 20× on output tokens. For a background agent that polls a queue every 30 seconds and summarises state changes, that difference decides whether the project survives its first finance review. Our benchmark rig exists exactly for this class of question — if you want the raw local throughput numbers, they live in AI Arena.

The European angle: since 2 August, the paper trail is mandatory

This is where European deployments now diverge sharply from US ones. The EU AI Act's binding enforcement took effect on 2 August 2026. The European AI Office and national authorities now hold real enforcement powers, Article 50 transparency disclosures apply, and general-purpose AI providers face strict compliance obligations under penalty.

In practice, for anyone running an agent that interacts with people or generates content, that means: the automated nature of the interaction has to be disclosed, outputs have to be marked where required, and you must be able to show who was accountable for what the agent did. Add the employee AI-literacy obligation and "we didn't know the agent did that" stops being an excuse and starts being a finding.

The sovereignty picture is also more boring than the 2024 rhetoric. European labs have largely stopped chasing proprietary frontier general-purpose models and pivoted to sovereign enterprise software, open weights, or hosting third-party foundation models. Meta's Muse under Apache 2.0, DeepSeek under MIT, GLM-5.3-Flash under MIT and Mistral's open-weight line mean a European company can now run a capable agent stack on its own hardware, inside its own jurisdiction, with no cross-border data transfer to justify in a DPIA.

Oversight is now the product

The honest conclusion from the data is that the hard part is no longer deployment. It is governance. Agent logs that are actually readable, a human checkpoint at the irreversible action, a data layer that does not serve a three-day-old record, and a named person who owns the outcome. Companies that treat those as engineering requirements rather than compliance overhead will be the ones converting the 180% into something closer to 180%.

Everyone else will keep generating code faster than they can ship it — and paying frontier prices to do it.

Do AI agents need their own AI Act compliance file?

Not a separate file, but a documented one. Under the enforcement regime effective 2 August 2026, if an agent interacts with people or generates content you must disclose the automated nature, mark synthetic output where required, and keep logs sufficient to show human oversight and accountability. Most companies fold this into their existing GDPR records of processing, which is cheaper than a parallel system.

Is it cheaper to self-host an agent model or pay per token?

It depends entirely on call volume. An open-weight model like GLM-5.3-Flash at roughly €0.46 per million output tokens is hard to beat on price, but you pay for GPU, electricity, engineering time and updates. Per-token API pricing wins below a few million output tokens per month; above that, self-hosting usually wins — and gives you the data-residency story for free.

Why do some teams get slower after adopting agents?

Because review capacity did not scale with generation capacity. McKinsey's finding of coding activity up 180% against shipped releases up 30% points at verification, testing and integration as the real constraint. If you cannot review the output faster, adding agents adds queue, not throughput.

Discussion

No comments yet — be the first to share your thoughts.
X

Don't miss out!

Subscribe for the latest news and updates.