Skip to main content

Only 7% of Firms Scale AI — jaam's M. Platform Bets on Context

Ilustrační obrázek
On 29 September 2026, jaam automation introduced M. by jaam — an “agentic work management” platform whose central claim is not a smarter model, but a Context Engine that knows your processes, your process history and your decision patterns. The diagnosis is plausible. The invoice, the audit trail and the EU AI Act are where it now gets tested. The launch date and product description come from the FF News launch report.

Almost every enterprise AI announcement this year makes a version of the same argument: the models are good enough, the missing piece is context. jaam automation is making that argument with its own product rather than someone else’s. According to the launch report from FF News, M. by jaam coordinates complex workflows across people, enterprise systems and AI agents, with agents gradually assuming what the company calls “incremental responsibility” as confidence in the underlying context grows.

What “agentic work management” means in practice

Strip the vocabulary away and there are two components. The first is orchestration: an agentic layer that can plan a multi-step task, call tools, read from and write to systems of record, and hand off to a human when it should. The second — and this is the part jaam is selling hardest — is the Context Engine, a unified store of organisational knowledge and process history that gives agents something to reason about beyond the current prompt.

That second part is where the real engineering difficulty sits. An LLM that does not know your approval thresholds, your supplier quirks, your historical exception handling and your organisational chart will produce confident nonsense at scale. Giving it that knowledge means ingesting a lot of material that is often personal data, often unstructured, and often permissioned differently for every employee. “Context engine” is a marketing term; the underlying question is whether it respects your existing access model, your retention rules and your audit requirements.

The practical enterprise question is therefore less whether an agent can produce a convincing answer than whether it can operate inside a controlled, multi-step process. That means clear process ownership, measurable ROI, reliable data plumbing and a defined boundary between autonomous actions and human approval.

The gap the marketing numbers describe

The launch material frames the problem as a gap between broad AI adoption, individual productivity improvements and measurable enterprise-level financial impact. Those claims should be read as the company’s survey-based framing rather than as a universal market measurement: the supplied launch coverage does not identify a survey name, publisher and publication date for the specific percentages, so those figures are not repeated here.

The underlying distinction remains important. Individual productivity gains are local: a person finishes a draft faster. EBIT impact requires a process to be redesigned end to end, with fewer handoffs, faster cycle times and measurable exceptions. A tool that makes each of ten people 20% faster, while leaving the process between them untouched, often changes nothing at the P&L line. It just moves the bottleneck.

That is also why adoption figures need normal survey scepticism. “Using AI” can mean anything from a licensed Copilot seat to a production agent, while “impact” may be self-reported rather than measured against a controlled baseline. The useful test is not the headline percentage but whether a specific process produces a repeatable improvement after inference costs, exceptions and human review are included.

Our cost math: what an agentic workflow actually costs to run

Before anyone buys an agentic platform, it is worth knowing what the inference underneath costs, because that is the part that scales with volume rather than with seats. Take a realistic mid-complexity agent task: 40 sequential LLM calls, roughly 4,000 input tokens per call and 500 output tokens per call. That is 160,000 input tokens and 20,000 output tokens per completed task — my arithmetic, not a vendor benchmark, and deliberately conservative on the input side.

The following are model list/API prices, quoted in US dollars per 1 million tokens and linked to the relevant provider pricing pages. They are used only for the arithmetic below; availability, model names, regional terms and enterprise discounts can change.

ModelInput / output per 1MCost per taskPer 1,000 tasks
GPT-4o mini$0.15 / $0.60~$0.036~$36
Claude 3.5 Haiku$0.80 / $4.00~$0.21~$208
Claude 3.5 Sonnet$3.00 / $15.00~$0.78~$780

For this illustration I use the explicit conversion 1 USD = €0.92 — equivalent to 1 EUR = $1.087 — rather than a range. The table therefore represents roughly €0.033 to €0.72 per completed task, depending on routing. It is not a promise of an invoice total.

Two details matter more than the headline spread. First, real agents re-send accumulating context on every step, so the input side is often two to three times my estimate — pushing the most expensive runs higher. Second, prompt caching can change the picture substantially, but cache pricing and eligibility are provider- and model-specific. The current provider pricing pages linked above should be checked before using any cache assumption in a business case; no separate cache-price estimate is included here.

If a platform re-sends the same 100,000-token organisational context on every call without caching it, you are paying for the same bytes dozens of times per task. That is precisely why the Context Engine claim is a cost claim as much as a capability claim. A large shared context is only affordable when it is cached, versioned and permission-aware.

Any agentic platform worth evaluating should let you route different steps to different models — a low-cost model for classification and extraction, a frontier model for the one step that genuinely needs judgement. On our own AI Arena rig we treat tokens/sec and time-to-first-token as production metrics, not curiosities: for an agent, latency compounds across 40 calls.

These calculations cover only the stated input and output tokens. They exclude retries, tool calls, embeddings, storage, platform fees, taxes and human review. They also exclude engineering, monitoring, security, integration and support costs. In other words, they are a narrow inference-cost illustration, not a total cost of ownership.

The European angle: agents that act are regulated differently

Since August 2025, general-purpose AI obligations under the EU AI Act have applied according to the Act’s transitional timetable. Article 50 transparency obligations apply from 2 August 2026, but they are use-case-specific. They do not mean that every agent or every output must simply be marked in the same way. Depending on the interaction or content involved, the relevant duty may concern informing people that they are interacting with an AI system or making certain synthetic or manipulated content identifiable.

Higher-risk classification is also use-case-specific. An agent used in areas such as employment, credit or certain critical-infrastructure contexts may fall within the high-risk rules, but an enterprise workflow agent does not automatically enter the full high-risk regime merely because it is autonomous or uses a foundation model. The applicable obligations, and the dates on which they apply, depend on the system’s intended purpose, classification and the Act’s transitional rules. Where the high-risk regime does apply, requirements can include risk management, logging, human oversight and technical documentation.

Read that against a Context Engine that ingests organisational knowledge and process history. Under GDPR you are processing personal data at scale, you need a lawful basis, and systematic monitoring or large-scale profiling can trigger a data protection impact assessment. Ask any vendor for the sub-processor list, the hosting region, the data processing agreement and — most importantly — a clear answer to whether agents can write to systems of record and how those writes are logged. Read-only assistance is a much easier conversation with your DPO than an agent that can approve a purchase order.

jaam’s launch material does not spell out per-region hosting or EU-specific terms, so treat that as the first procurement question rather than an assumption. The same applies to pricing: no public EUR rate card was part of the announcement, and enterprise agentic platforms are still sold per engagement far more often than per token.

What I would do before signing

Pick one process with a measurable baseline — invoice exception handling, contract intake, order-to-cash reconciliation. Define success at process level, not seat level: cycle time, exception rate, cost per completed instance. Instrument inference cost per instance from day one, including the review time humans still spend. Then decide whether the orchestration layer you are paying for is doing something your own engineers could not assemble from an API, a queue and a cache.

Is agentic work management just RPA with an LLM bolted on?

No, and the difference cuts both ways. Classic RPA replays a fixed script against a stable UI. An agent chooses its own steps, which is what lets it handle variation — and also what makes it non-deterministic, more expensive per run and harder to certify. You trade predictability for flexibility, which is why logging, permission boundaries and human checkpoints matter more here, not less.

How do I measure whether an AI agent actually helped?

Measure per completed process instance, not per user. Add inference cost, orchestration cost and the human review time still required, then compare against your pre-pilot baseline for cycle time and exception handling. A pilot that looks great in a demo but adds 15 minutes of supervision per case is a net loss.

Can open-weight models run agentic workflows cheaply enough to self-host?

Possibly, if your context is moderate. Open-weight options such as Llama 4 under the community licence, or Mistral’s Devstral line, remove per-token API fees but move the cost into GPUs — and agents with long context need memory for the KV cache, which grows with every step. Self-hosting wins on data residency and predictable volume; hosted APIs usually win on cost at low volume.

The cost table above is an illustrative arithmetic example, not a vendor benchmark or a total cost-of-ownership estimate. The real decision depends on routing, retries, tool use, data controls, platform charges and the amount of human review a workflow still needs.

Discussion

No comments yet — be the first to share your thoughts.
X

Don't miss out!

Subscribe for the latest news and updates.