Skip to main content

Skylark Labs' CFAM lets robots learn on-site — Europe's certification assumes they don't

Ilustrační obrázek
Skylark Labs, working with researchers from Carnegie Mellon University and UC Berkeley, has published an architecture that lets a robot keep learning after it has been shipped — no cloud round-trip, no full retraining. The measured numbers are believable: action success up from 74.0% to 87.9%, forgetting down to −0.5 percentage points against −11.4 for standard LoRA fine-tuning. The awkward part for European buyers is not the maths. It is that EU certification machinery is built around a defined assessed version, not around a robot that keeps adapting after delivery.

What CFAM actually changes

The system is called Continual Field-Adaptive Models (CFAM), detailed in the paper Continual Field-Adaptive Models (CFAMs) for Post-Deployment Physical AI, published this month and covered by Robotics & Automation News on 27 September. The design borrows a well-known idea from neuroscience: complementary learning systems, where one fast, episodic memory records what just happened and a slower system consolidates it into long-term behaviour.

In CFAM, that fast layer is a store of Competence Capsules — compact records of successful real-world task executions. The slow layer is a set of Geometric Residual Transforms (GRT), which adjust movement parameters rather than rewriting the whole control policy. The practical consequence: a deployed machine can absorb a new bin height, a different lighting condition or a slightly sticky gripper without anyone shipping the weights back to a GPU cluster.

That last point is the one worth stressing. "Learning after deployment" is not a new slogan — everyone from NVIDIA's Isaac GR00T stack to Google DeepMind's robotics line has been chasing it. What is new here is the combination of on-device adaptation and an explicit benchmark for how much old skill gets destroyed in the process.

The numbers, and what they mean

Action success rate is the headline figure, but forgetting is the more interesting one. In the paper's evaluation, success rises from 74.0% for the pre-adaptation baseline to 87.9% with CFAM. Those two success figures are directly comparable with each other. The forgetting comparison is separate: CFAM reports backward transfer of just −0.5 percentage points, versus −11.4 points for a LoRA-style fine-tune. The 74.0% figure is not necessarily a LoRA result; it is the paper's baseline before CFAM adaptation.

MetricPre-adaptation baselineCFAMLoRA-style fine-tune
Action success rate74.0%87.9%—
Backward transfer (forgetting)—−0.5 pp−11.4 pp

Read the second row carefully: −0.5 pp is, for practical purposes, "your old skills stayed where they were." That is a different claim from "our model is smarter." It is a claim about stability, which is exactly what factory and logistics operators care about when an arm has to sort 40 SKUs in the same shift.

The evaluation ran on an internal dataset of more than 2.6 million trajectories across five physical robot platforms. That is a substantial corpus — and also the main caveat. It is vendor-collected data, evaluated by the vendor's own team. There is no third-party replication yet, no public leaderboard, and I have not seen an independent lab reproduce the backward-transfer figure. Impressive demo; unknown production behaviour under conditions Skylark did not choose.

The cloud cost, back-of-envelope

Here is a calculation worth doing, because "on-device learning" is often sold as a cost story and the cost story is usually wrong. Assume an average trajectory is around 300 tokens of state and action data. At 2.6 million trajectories that is roughly 780 million tokens for a single full pass over the field dataset.

This next part is a hypothetical token-cost illustration, not a current price quote. Depending on the provider, rate tier and peak/off-peak pricing, such a pass could sit anywhere from roughly $100 to $1,500 based on the range of current public model prices. At an assumed rate of 1.08 USD/EUR, that is roughly €93 to €1,390. In other words: raw compute is not the reason CFAM matters.

What actually costs money is everything around the training call: getting shop-floor data off the site, round-trip latency when the arm has to pick the right part now, and the governance of sending camera feeds to a third party. On-device adaptation wins on latency and data governance, not on dollars per token. That is the honest framing.

The open question I would put to Skylark is the compute and power footprint of a capsule update on real edge hardware. Public material does not break that down, and it is the number that decides whether a €30,000 cobot can afford the feature — the same class of question I chase on our own AI Arena rig, where VRAM and tokens/sec decide what is deployable long before benchmark charts do.

The European problem: certification is built around a defined version

The EU AI Act's obligations are staged. The relevant milestone here is 2 August 2026, when the regulation's rules for Annex III high-risk AI systems and the transparency duties for certain systems under Article 50 began to apply under Regulation (EU) 2024/1689 — an application date for specific obligations, not a single "enforcement" date. National market surveillance authorities oversee compliance in the member states. Separately, the Machinery Regulation (EU) 2023/1230 applies from 20 January 2027; that product safety regulation was drafted with self-modifying safety functions in mind.

CFAM does not automatically create a substantial modification. Under the AI Act, a modification is substantial only if it affects the system's compliance with the regulation or changes its intended purpose. Whether a field adaptation crosses that line depends on the robot's classification, its safety role and the risk impact of the specific change. A German or Czech warehouse operator buying a self-learning arm therefore cannot treat the CE mark as a one-time receipt: each update must be checked against the original technical file and, where the threshold is met, a fresh conformity assessment may be required.

Expect notified bodies to ask a blunt question: can you pin the behaviour to a version, log every change, and switch learning off entirely? If the answer is no, a European procurer may have to treat the system as a research demonstrator rather than a compliant product — subject to specific legal analysis. This article is not legal advice, and the assessment will always be use-case-specific.

Then there is GDPR. Storing capsules locally is not an exemption. If those capsules contain frames, audio or movement traces that identify workers on a shop floor, they are personal data. Local storage reduces the transfer problem; it does not create a lawful basis, and repurposing them to improve a vendor's general model runs straight into purpose limitation. The Commission's own framework pages are worth reading before any pilot procurement.

Where it is actually shipping

Commercially, Skylark's momentum is American and defence-shaped: robotic arms with combined visual and tactile sensing demonstrated at TechCrunch Disrupt in July 2026, a partnership with Hexadyne in August 2026 to scale self-learning platforms into U.S. and allied markets (PACRIM, GCC), a $35 million five-year licensing deal with ideaForge, and a multi-million-dollar road infrastructure monitoring contract in March 2026. Those details are reported by Robotics & Automation News; none of the published material identifies an EU entity, EU pilot, EUR pricing or listed hardware requirements.

So for a European integrator, the realistic answer is: interesting architecture, not yet a purchasable product for the EU on the basis of published information. What you can do today is put the requirement in your next tender — version pinning, change logging, a documented learning envelope, and the ability to freeze behaviour for audit.

Can a European company deploy a robot that keeps learning after delivery today?

It depends on the system's classification and the nature of the update. Not every field adaptation is a substantial modification; it becomes one if it affects compliance with the AI Act or changes the system's intended purpose. High-risk systems, safety components and Annex III cases need a case-specific assessment. Practically, ask the vendor for a frozen build you can revert to, and seek qualified legal counsel before deployment.

Does CFAM replace fine-tuning and cloud training entirely?

No. It handles small field adaptations — new bin heights, altered lighting, a worn gripper. Base competence still comes from large-scale pre-training elsewhere. CFAM is a maintenance mechanism, not a model-training replacement.

Is CFAM open source or available in the EU?

There is no indication of open weights, an SDK or an EU availability programme in the published material. The published partnerships point to U.S. and allied defence and infrastructure markets. Until hardware requirements and EU pricing exist, European teams can only study the paper.

Discussion

No comments yet — be the first to share your thoughts.
X

Don't miss out!

Subscribe for the latest news and updates.