What is actually in the box
Kolibri-1 is a mixture-of-experts model. It holds 78.1 billion parameters in total but activates roughly 3.46 billion of them per token. That gap matters for anyone planning capacity: memory has to hold all 78.1 billion weights, while compute per token stays close to a small dense model. Aleph Alpha scaled it from an internal prototype called Kolibri Origin, which had 30 billion parameters and 3.27 billion active, using an automated training pipeline the company calls Model Factory. According to The Decoder, the whole path from prototype to release took about three months.
The context window is 262,144 tokens natively, with validation extended to 1,048,576 tokens. A custom UniBPE tokenizer with a 128,000-token vocabulary is claimed to cut German text fragmentation by up to 15%, which is a quiet but real argument for anyone running German-language document pipelines. Training used about 24 trillion tokens: 20 trillion in pre-training, 3.44 trillion in mid-training and 201 billion for the long-context extension. German data made up roughly 23.4% of the mix. The run used 768 NVIDIA B200 GPUs over 21 days.
The arithmetic behind those figures is checkable. 768 GPUs running for 21 days is 387,072 GPU-hours; Aleph Alpha quotes 392,000, which fits extra short runs or evaluation passes. Total energy consumption was around 950 MWh.
Hardware and the real cost
The weights ship in FP8 format and take about 78 GB of VRAM. In practice that means one H200 or B200, or two A100 80 GB or H100 cards, plus the KV cache on top if you push the context window far past the native 262k. Quantising to 4-bit brings the weights down to roughly 40 GB by simple arithmetic, which still does not fit on a single 24 GB card.
On our own AI Arena rig, an RTX 5060 Ti with 16 GB, Kolibri does not fit at any quantisation level that leaves room for a usable context. That is not a criticism of the model; it is the boundary between open weights aimed at enterprises and open weights aimed at enthusiasts, and Kolibri sits firmly on the enterprise side.
Two cost estimates are worth writing down, both clearly labelled as estimates rather than vendor figures. At an industrial electricity rate of roughly €100 per MWh, 950 MWh of training power comes to about €95,000. At a rough $2 per GPU-hour for B200-class rental, 392,000 GPU-hours is about $780,000. The second number dominates, and neither includes staff, data work, evaluation or the hardware purchase itself. There is no software licence line at all under Apache 2.0.
Benchmarks, as reported by the vendor
Aleph Alpha lists AIME 2025 at 96.9, GPQA Diamond at 84.3 in English and 81.3 in German, HumanEval+ at 92.7, LiveCodeBench v6 at 85.9, SWE-Bench Verified at 66.4 and TerminalBench 2.1 at 27.7. On the RULER long-context benchmark it scores 63.2% at the 1M-token setting. These are self-reported numbers; no independent replication has been published yet, and the drop from strong coding results to 27.7 on terminal tasks is the kind of detail worth watching in third-party runs.
The most interesting figure is not a benchmark at all. The model refused to answer 44% of queries that lacked sufficient context or evidence. For a general chatbot that refusal rate would be a problem. For a system reading internal policy documents, contracts or public-sector files, abstaining instead of guessing is the feature that gets it through a procurement review.
How it compares on price
Kolibri has no per-token price because Aleph Alpha has not published a hosted endpoint. The comparison against commercial APIs is therefore about total cost of ownership rather than a rate card. Current list prices are billed in USD, with no EUR pricing published by the vendors below:
| Model | Input / 1M tokens | Output / 1M tokens |
|---|---|---|
| DeepSeek-V4.1-Flash (off-peak) | $0.15 | $0.60 |
| GPT-6.1 Sol | $0.40 | $2.00 |
| GLM-5.3 | $1.40 | $4.40 |
| Mistral Medium 3.5 | $1.50 | $7.50 |
| Grok 4.7 (under 200K context) | $2.00 | $6.00 |
| Claude Sonnet 5.5 | $2.00 | $10.00 |
Meta's Muse Glimmer, released on 10 August 2026, is the other Apache 2.0 open-weight option in that group. The difference is deployment: Muse Glimmer also runs on commercial clouds, while Kolibri's pitch is specifically on-premises and offline execution with European compute behind the training run.
Regulation and the Cohere merger
The EU's Digital Omnibus, passed in June 2026, pushed high-risk AI compliance deadlines to December 2027, or August 2028 for embedded products. Transparency rules under Article 50 and general-purpose AI governance took effect in August 2026 instead of being delayed. A deployer running open weights locally still carries those transparency duties; downloading a model does not move the obligation to someone else.
On 16 September 2026, Aleph Alpha announced a definitive agreement to combine with Cohere, pending regulatory approval. Weights already published under Apache 2.0 cannot be withdrawn retroactively, so the version released this week stays available regardless of how that review ends. What is not yet confirmed is the commercial roadmap, hosted pricing or long-term support model for the merged entity.
What a developer can do this week
Pull the weights, quantise them, and test them on rented H200 capacity before committing to hardware. Teams with existing German-language document workloads should measure tokenizer efficiency against their own corpora, since the 15% fragmentation claim is the kind of number that varies by text type. Anyone without datacentre-class GPUs will be waiting for a hosted endpoint that does not currently exist.
Can Kolibri run on a single consumer GPU?
No. The FP8 weights are about 78 GB, and 4-bit quantisation lands near 40 GB before the KV cache. That needs two 24 GB cards at minimum, or a 48 GB workstation card, and long contexts push requirements higher.
Is there a Kolibri API with per-token pricing?
As of the release, Aleph Alpha has published no hosted API price for Kolibri. The only route is running the weights yourself or through a third-party host.
Could the Cohere merger change the licence?
The version published under Apache 2.0 stays under Apache 2.0. A merged company can change the terms of future releases, but not the rights already granted for existing weights.