Open Weights Without the Raw Data: Aleph Alpha Lays Bare Its 78B Model Recipe

Open Weights Without the Raw Data: Aleph Alpha Lays Bare Its 78B Model Recipe

Artificial IntelligenceLarge Language ModelsOpen Source Ecosystem

Sources:Aleph Alpha 技术报告 + HN 讨论

The tech industry talks incessantly about open source, yet few distinguish between releasing model weights and releasing the underlying training data.

On German Unity Day, October 3, 2026, European AI firm Aleph Alpha unveiled Kolibri, a massive 78-billion-parameter model. While they open-sourced the model weights under a permissive license, they did not release a single kilobyte of raw training data.

The release immediately ignited fierce debate across developer communities. One camp insisted that withholding data violates the fundamental tenets of open source, while pragmatists countered that the industry has long accepted open weights as the de facto standard.

A Shift in Meaning: Delivering the Product Without the Raw Ingredients

Giving users a free meal is fundamentally different from handing over the ancestral recipe and supplier roster. For years, tech giants have used the umbrella term “open source” to gloss over their proprietary lock on core pre-training datasets.

Kolibri charts a pragmatic middle ground. Aleph Alpha uploaded its Mixture-of-Experts (MoE) architecture—featuring 3 billion active parameters—to Hugging Face under the Apache 2.0 license. But instead of uploading the raw pre-training corpus, they published an exhaustive, reproducible blueprint detailing how the dataset was constructed from scratch.

Engineers involved in pre-training joined online discussions to address fine-grained questions about data cleaning and filtering. Handing over physical hard drives is merely symbolic; true transparency lies in demonstrating total command over the entire data distillation pipeline.

Kolibri official visual Figure: White hummingbird logo and wordmark against a green gradient background. Source: Aleph Alpha

Rejecting Translationese: Reclaiming Control with 21% Native Language

Training a model of this scale required ingesting 20 trillion tokens. In the core pre-training corpus, native German accounted for 21.3%, while translated text was restricted to just 6%.

That deliberate ratio reflects a strategic defense. Over-relying on translated corpora inevitably imprints Anglo-American cultural biases onto the model. Aleph Alpha clearly documented its methodology across exact deduplication, fuzzy deduplication, substring deduplication, and heuristic filtering.

Throughout the corpus expansion from 7.5 trillion to 20 trillion tokens, the distillation of quality classifiers was laid bare. Any team that can publish an end-to-end data curation blueprint earns the right to claim sovereign independence.

Compute Cannot Buy Restraint: Teaching the Machine to Say “I Don’t Know”

In benchmark evaluations, Kolibri achieved a score of 96.0 in the AIME 2026 mathematics competition, matching rival models with four times its active parameter count.

Aleph Alpha implemented the Merlin-Arthur protocol during training, specifically coaching the model to acknowledge the boundaries of its knowledge. When confronted with out-of-distribution queries or unanswerable contexts, the model directly responds that it does not know, refusing to produce confident hallucinations.

Preventing hallucinations requires an acute perception of knowledge boundaries—an engineering discipline that cannot be achieved simply by throwing more GPUs at the problem.

Kolibri-1 model card Figure: Details of the Kolibri-1 model card on Hugging Face. Source: Hugging Face / Aleph Alpha

Replacing the Ingredients with the Recipe

Copyright liabilities remain an unresolved ledger. Facing scrutiny from German authors and publishers, Aleph Alpha has yet to disclose specific compensation amounts, leaving legal questions lingering in the background.

Even so, by documenting the data generation process in reproducible detail, Kolibri brings the industry’s gray zones into sharp relief. Handing over storage drives is no longer the sole benchmark of openness.

When an AI lab provides full disclosure of its recipes, filtering heuristics, and classifier distillation pipelines, technological sovereignty remains firmly in its own hands.

Reference Links:

  • Aleph Alpha Launches Kolibri on German Unity Day
  • HN Discussion (Open-Weight vs. Open-Data)