The tech industry talks incessantly about open source, yet few distinguish between releasing model weights and releasing the underlying training data.
On German Unity Day, October 3, 2026, European AI firm Aleph Alpha unveiled Kolibri, a massive 78-billion-parameter model. While they open-sourced the model weights under a permissive license, they did not release a single kilobyte of raw training data.
The release immediately ignited fierce debate across developer communities. One camp insisted that withholding data violates the fundamental tenets of open source, while pragmatists countered that the industry has long accepted open weights as the de facto standard.
A Shift in Meaning: Delivering the Product Without the Raw Ingredients
Giving users a free meal is fundamentally different from handing over the ancestral recipe and supplier roster. For years, tech giants have used the umbrella term “open source” to gloss over their proprietary lock on core pre-training datasets.
Kolibri charts a pragmatic middle ground. Aleph Alpha uploaded its Mixture-of-Experts (MoE) architecture—featuring 3 billion active parameters—to Hugging Face under the Apache 2.0 license. But instead of uploading the raw pre-training corpus, they published an exhaustive, reproducible blueprint detailing how the dataset was constructed from scratch.
Engineers involved in pre-training joined online discussions to address fine-grained questions about data cleaning and filtering. Handing over physical hard drives is merely symbolic; true transparency lies in demonstrating total command over the entire data distillation pipeline.
Figure: White hummingbird logo and wordmark against a green gradient background. Source: Aleph Alpha
Rejecting Translationese: Reclaiming Control with 21% Native Language
Training a model of this scale required ingesting 20 trillion tokens. In the core pre-training corpus, native German accounted for 21.3%, while translated text was restricted to just 6%.
That deliberate ratio reflects a strategic defense. Over-relying on translated corpora inevitably imprints Anglo-American cultural biases onto the model. Aleph Alpha clearly documented its methodology across exact deduplication, fuzzy deduplication, substring deduplication, and heuristic filtering.
Throughout the corpus expansion from 7.5 trillion to 20 trillion tokens, the distillation of quality classifiers was laid bare. Any team that can publish an end-to-end data curation blueprint earns the right to claim sovereign independence.
Compute Cannot Buy Restraint: Teaching the Machine to Say “I Don’t Know”
In benchmark evaluations, Kolibri achieved a score of 96.0 in the AIME 2026 mathematics competition, matching rival models with four times its active parameter count.
Aleph Alpha implemented the Merlin-Arthur protocol during training, specifically coaching the model to acknowledge the boundaries of its knowledge. When confronted with out-of-distribution queries or unanswerable contexts, the model directly responds that it does not know, refusing to produce confident hallucinations.
Preventing hallucinations requires an acute perception of knowledge boundaries—an engineering discipline that cannot be achieved simply by throwing more GPUs at the problem.
Figure: Details of the Kolibri-1 model card on Hugging Face. Source: Hugging Face / Aleph Alpha
Replacing the Ingredients with the Recipe
Copyright liabilities remain an unresolved ledger. Facing scrutiny from German authors and publishers, Aleph Alpha has yet to disclose specific compensation amounts, leaving legal questions lingering in the background.
Even so, by documenting the data generation process in reproducible detail, Kolibri brings the industry’s gray zones into sharp relief. Handing over storage drives is no longer the sole benchmark of openness.
When an AI lab provides full disclosure of its recipes, filtering heuristics, and classifier distillation pipelines, technological sovereignty remains firmly in its own hands.
Reference Links:
- Aleph Alpha Launches Kolibri on German Unity Day
- HN Discussion (Open-Weight vs. Open-Data)