Aleph Alpha puts a 1M-context model on your own servers

Aleph Alpha's Kolibri has 78.1 billion parameters but uses only 3.46 billion of them for each token, and you can download every weight under Apache 2.0. The German lab built this English-German Mixture-of-Experts model, with a context window of up to one million tokens, for governments and regulated industries that want to run it on their own servers.
According to a post on Threads, Kolibri is probably the first general-purpose model from the EU, apart from Mistral, to reach this level of performance. The post includes this description of the model:
"The model supports an explicit reasoning mode and tool calling. It is optimized for long-context and inference efficiency."
At a glance
- Aleph Alpha released Kolibri on October 3, 2026, German Unity Day, with full weights on Hugging Face so customers can run it on-premises instead of sending internal data to an outside inference service.
- Six of 384 experts handle each token, and only 10 of 50 layers use full attention, while the other 40 look back through a 512-token sliding window to contain inference costs.
- The benchmarks are Aleph Alpha's own, run on its harnesses at the highest available reasoning setting where applicable, and long-context adaptation stopped at 256,000 tokens even though serving settings go to 1,048,576.
If you have not been following Aleph Alpha: according to Skypage, the company was founded in 2019 in Heidelberg and is often called Europe's answer to OpenAI, with a mission centered on sovereign AI. According to Aleph Alpha, it first built a training pipeline and tested it by building Kolibri Origin, a 30B total, 3B active model with a much shorter 65k token context window. Kolibri then went through the same pipeline.
Kolibri trained on nearly 24 trillion tokens across 768 B200 GPUs
Kolibri builds on Kolibri Origin and on Aleph Alpha's automated Model Factory. Training ran on 768 B200 GPUs. Pre-training covered 20 trillion tokens over 21 days, and mid-training plus long-context adaptation brought the total to nearly 24 trillion tokens. Aleph Alpha built the model in Germany and trained it in Germany and Finland.
German makes up 21.3% of pre-training tokens. The bilingual vocabulary has 128,000 entries, and the tokenizer is designed to keep German compound words intact instead of chopping them into fragments. According to Aleph Alpha, the pipeline made hundreds of ablation experiments possible, and pre-training kept running without anyone stepping in when hardware failed or a data connection dropped.
On launch day, Konark Modi wrote that his team had hosted Kolibri-1 and made it free for anyone to try for a limited period. He called the release on German Unity Day especially fitting.
Kolibri scored 96.9 on AIME 2025 on Aleph Alpha's own harnesses
In the company's benchmarks, Kolibri posted an English overall score of 75.5 and a German overall score of 70.8. It scored 96.9 on AIME 2025, 85.9 on LiveCodeBench v6 and 61.4 on BFCL v4 overall. Aleph Alpha ran all of these itself on its own harnesses, using the highest available reasoning setting where applicable.
Aleph Alpha says Kolibri sits on the Pareto frontier for quality versus serving cost in English and German. By its account, none of the compared models gives more quality at the same serving cost, or the same quality for less. It also says Kolibri can match models with up to four times as many active parameters on math, code, grounding, agentic and long-context tasks.
The company says it specialized the model for German, math, coding, long-context work and agentic tasks. Its sector evaluation suites cover public administration, automotive, semiconductors, industrial technology and aerospace, and none of them use customer data. Kolibri has four reasoning settings (none, low, medium and high) plus tool calling, and it is deployed through Aleph Alpha's inference package and a Kolibri-specific vLLM plugin.
Kolibri avoided a wrong answer on 44% of AA-Omniscience items
Grounding is at the center of the release. Kolibri was trained on abstention examples and with Aleph Alpha's Merlin-Arthur procedure, an adversarial training method that teaches the model to hold back an answer when the evidence is missing. Kolibri avoided a wrong answer on 44% of AA-Omniscience items, against 14.8% for Kolibri Origin, and reached 0.23 on the company's M/A grounding score.
According to Artificial Analysis, the AA-Omniscience Index rewards correct answers, penalizes hallucinations and does not penalize refusals. It runs from -100 to 100, where 0 means as many correct answers as incorrect ones. Its separate Hallucination Rate divides incorrect answers by all non-correct responses, including partial answers and questions not attempted.
Why do only six of 384 experts work on each token?
A Mixture-of-Experts model splits much of its capacity into many sub-networks, called experts, and runs only a few at a time. According to NVIDIA, a learned gating network, the router, picks a small subset of experts for each token after the self-attention blocks, and the routing repeats across successive MoE layers. It works like a large office where a receptionist puts each call through to the few specialists suited to it.
In Kolibri, that means six of 384 experts per token, so only 3.46 billion of 78.1 billion parameters do the work. NVIDIA notes the trade-offs. Training needs careful load balancing so every expert gets comparable exposure, and at inference, spreading experts across GPUs puts strain on memory bandwidth.
Attention gets similar treatment. In full attention, every token can look at every earlier token, but Kolibri uses it in only 10 of its 50 layers. The other 40 see just the last 512 tokens.
Every quality figure here comes from Aleph Alpha's own harnesses. The inputs also do not say whether the 44% is a share of all AA-Omniscience items or of the non-correct responses that Artificial Analysis uses for its Hallucination Rate. In our view, the bigger practical gap is context: adaptation reached 256,000 tokens, so the full 1,048,576-token window goes beyond the length the model was trained on, and no quality figure at that length is given.
The 1,048,576-token window stays untested
The inputs give no quality figure for Kolibri at the full 1,048,576 tokens, and no date has been given for a long-context evaluation at that length. Independent results on AIME 2025, LiveCodeBench v6 and BFCL v4 are not reported in the inputs either. Konark Modi's free hosting of Kolibri-1 runs only for a limited period, and the end date is not stated.
Related stories
- Six models, 0.9B to 375B, ship with training logs
- Ideogram 4.5 prices images from 0.8 to 22 cents at 2K
- Xiaomi's new MiMo models carry an unverified top-6 claim
- Cohere claims a WMT26 lead with a non-reasoning translator
- 552B MoE opens DeepSeek's new architecture family
- GLM-5.3 weights land with a hyperscaler clause
Comments
No comments yet. Be the first.
Join the conversation
Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.
