Tuesday, October 6, 2026
spot_img

What we can learn from the state of AI by how Aleph Alpha’s Kolibri was developed

Aleph Alpha, a Canadian-European lab based in Germany that recently merged with Canada’s Cohere, released Kolibri on October 3, 2026. It’s a German-English language model with downloadable weights under the Apache 2.0 licence. The company is positioning it for organizations that want to run and adapt AI themselves, particularly in public administration and industry. Its stated priorities include German-language performance, documented development and the ability to decline questions (epidemic stop) when available information can’t support an accurate answer. [1] what’s fascinating about the model is not only its performance how its development in German affects, not only sovereignty but model capabilities in ways that we have explored in our research, including the fact that it adds an epistemic stop that doesn’t exist in English trained models.

The release illustrates how AI capability is now moving between companies and countries. Aleph Alpha built its own model, while using other models to help produce its training material. Its technical report identifies GLM-5.2, GLM-5.3 and Qwen3.8-27B as the principal generators of its supervised fine-tuning data. GLM is developed by Z.ai; Qwen is Alibaba’s model family. [3]

Understanding that sequence explains both the new German model and part of the wider proliferation of AI. A model can be a finished product, a starting point for another model, or a tool used to train one. These roles allow capability to spread through several routes at once.

The release messaging emphasises sovereignty and practical deployment. Some commentary has treated its use of Chinese-generated training examples as contradicting the claim that it was trained from scratch. Those claims describe different stages. Starting with new weights doesn’t require every training example to be human-written. The origin of the examples is relevant, but it is a separate question from whether the model inherited another model’s weights.

Model Structure Comes First

First, Aleph Alpha chose the model’s structure and how it would represent text.

Kolibri uses a mixture-of-experts architecture: different inputs activate different subsets of its learned parameters. It has about 78 billion parameters in total and activates approximately 3.46 billion per token. A token is a unit of text processed by the model, which may be a word, part of a word or punctuation. [2]

Selective activation reduces the calculation required for each token. This doesn’t mean that the entire model occupies the memory of a three-billion-parameter model, but affects where a model can run and how much serving it costs.

The tokenizer determines how text, images and code etc are divided into those units. Aleph Alpha developed one suited to German word structure. [2] Language efficiency starts here: different ways of dividing the same sentence can require different numbers of tokens, affecting processing cost and how tokens fit into an input.

How Pre-Training Works

Second, it trained the initial model.

The parameters began as random values. Aleph Alpha then pretrained Kolibri on 20 trillion tokens, using a bilingual corpus that also included code. [2]

During this stage, the model repeatedly predicts the next token in a sequence. The training process compares its prediction with the actual continuation and adjusts its parameters to reduce the error. Repeating this across many examples develops representations of language, factual associations, code and other patterns present in the data.

This is where “trained from scratch” applies. Aleph Alpha didn’t begin this process by loading Qwen’s completed weights and renaming the result. Its own parameters were learned through training.

Extended Training and Length Optimization

Third, it continued training and extended the input length.

The company reports another 3.44 trillion tokens of mid-training, followed by 201 billion tokens for long-context training. Sequence lengths increased from 16,384 tokens in pretraining to 65,536 and then 262,144. [2]

These stages let developers change the balance of training material and teach a model to handle longer sequences. They are additional weight updates, rather than just changes to the number displayed in a product spec.

Kolibri’s advertised one-million-token capacity goes beyond its native training length. Aleph Alpha says it tested that extension but recommends at most 262,144 tokens for complex work and serving efficiency. [1] Being able to digest a document and reliably use all of it are separate capabilities.

Fine Tuning Design with GLM and Qwen

Fourth, other models helped create demonstrations of behaviour Kolibri should learn.

This is where GLM and Qwen enter the process. A teacher model receives a task and generates an example response. Depending on the task, that response might include an explanation, code, a sequence of searches or tool calls. The resulting material becomes candidate training data for the new model.

The teacher’s parameters are not copied into the student. The student learns from the teacher’s outputs. This is one form of transferring capability through synthetic data, sometimes referred to as distillation.

For Kolibri, the report describes generated and regenerated examples, filtering for benchmark overlap, reasoning-effort labels and experiments with different mixtures of datasets. [3] A “teacher’s” output is an input to a development process, versus an automatic guarantee of quality.

This is the stage at which expertise from domain experts have a substantial role. They can decide which tasks need examples, which answers are acceptable, what evidence is required and which behaviours should be excluded. A model capable of generating millions of responses still needs a useful spec of what those responses should teach.

Fifth, Aleph Alpha used those demonstrations for supervised fine-tuning.

Supervised fine-tuning changes the model’s weights using examples of desired responses. A training example effectively says: given this question, conversation or document, this is the kind of continuation to produce.

Aleph Alpha trained candidate mixtures and averaged the weights of two selected runs before reinforcement learning. [3] This is why the data recipe is part of the model’s design. Which examples are included, how frequently they appear and how they are evaluated can change the resulting behaviour.

The GLM and Qwen contributions persist without those models needing to answer every later Kolibri query. The training examples have influenced Kolibri’s own parameters. Running the finished model locally does not, by itself, necessitate the sending of user prompts to the companies that supplied the teachers, an accusation that was made at one point of Chinese models that were quickly gaining proficiency.

Fascinating lessons from German language training


The German-language experiments show why the training mixture is consequential.

In research published before Kolibri’s release, Aleph Alpha compared 13 fine-tuning mixtures on the same model, varying the German training data. Its baseline had no German examples in this fine-tuning stage, not no German exposure throughout training. It could solve German questions while generating its intermediate reasoning in English.

Adding German reasoning examples changed that behaviour, but initially reduced accuracy. On German AIME mathematics questions, scores fell from 70.2 to a low of 48.3. At that low point, approximately a quarter of German reasoning sequences repeated themselves until they exhausted the available context, producing no final answer.

Then performance rose again. Increasing the German mathematics examples brought the score back to 67.3 and reduced looping to around 15%. The recovery depended on examples relevant to the task; increasing German data generally did not reliably improve performance. These were fine-tuning experiments, so the looping rates should not be presented as measurements of the finished Kolibri release. (aleph-alpha.com)

The scoring also reveals a practical distinction. When the baseline had to produce an answer that was both correct and German, its score was only 42.3. The German-trained variants consistently answered those mathematics questions in German. They improved language consistency even while some struggled to complete their reasoning. (linkedin.com)

For model development, this separates several capabilities that a fluent response can make look interchangeable: understanding a question, working through it, answering in the required language and finishing successfully. Improving one can initially disrupt another. The recovery shows that relevant training examples can repair much of that disruption.

This also gives “knowing when to stop” a concrete meaning. Having an epidemic stop dramatically, reduces the number of hallucinations because the model is not forced to come up with an answer. Completing a reasoning sequence and recognizing insufficient evidence are separate behaviours. A repetition loop demonstrates a failure to finish; it doesn’t establish whether the model “recognises” the limits of its knowledge. Reliable systems need both termination and appropriate abstention.

The cultural question sits alongside this. Translating an English example preserves much of its original context: its assumptions, institutions and conversational conventions. Developing a model for German users therefore involves choosing what situations and knowledge its training examples represent, as well as which language they use. The experiment demonstrates a language-and-task dependency; the broader implication is that local expertise remains necessary when adapting models for actual communities and workplaces.

Sixth, it trained through feedback on attempted tasks.

Kolibri’s post-training also uses reinforcement learning, with more than 1.2 million internally curated tasks covering areas including reasoning, retrieval, coding and tool use. [3]

In general, this stage lets a model try to complete tasks and receive scores according to defined criteria. Training then adjusts the model to make rewarded behaviour more likely. The reward design is consequential: producing an answer, producing a correct answer and producing an answer supported by the supplied evidence are different objectives.

Aleph Alpha says it specifically trains and evaluates abstention. [1] The intended behaviour is to decline when the documents do not contain an answer. Whether that works consistently must be established in the actual deployment. A fluent statement of uncertainty is useful only when the decision to abstain is appropriately calibrated.

Finally, the weights become available for other organisations to deploy and adapt.

Open weights give a developer access to the numerical parameters that define the trained model. Subject to the licence and available hardware, the developer can run it, alter it and build a product around it. Access to weights does not necessarily include the complete training dataset or everything needed to reproduce the original training.

This creates several routes to new AI products. A team can adapt an existing model’s weights. It can train a new model with examples generated by existing models. Or it can combine multiple models inside one application, assigning each a different job.

The result can be more products with distinct languages, costs and purposes, even when some of their capabilities share a common origin. Model proliferation and shared foundations can increase together.

Qwen’s distribution helps explain the scale of that opportunity. Alibaba reported more than three billion global downloads in August; Hugging Face’s own summer report recorded approximately 2.06 billion downloads across Qwen repositories on its Hub during the first seven months of 2026.

Hugging Face also counted 151,448 Qwen-derived models on its platform. These include adaptations and conversions, rather than 151,448 independently invented architectures. Downloads likewise do not equal unique users or operational deployments. Nevertheless, the report documents substantial activity building on the model family. [5]

There are two distinct effects here. Downloadable weights enable developers to work directly with an existing model. Generated examples enable a model’s outputs to influence another model, including one with independently trained weights. Synthetic-data generation can also happen through an API; Kolibri’s use of teacher models alone does not establish that open-weight access was essential to every such step.

Together, these routes make distributed AI a practical development model. Expertise can be applied closer to a particular organisation, language or task. An application can use a locally operated model for routine processing and call a larger model when needed. A specialised supplier can sell a completed workflow rather than requiring every customer to become a model developer.

Distribution doesn’t eliminate dependencies. Hardware, training data, software libraries and teacher models still have origins. It does change which dependencies apply during everyday use and which decisions a deploying organisation can control. That is a more useful way to examine sovereignty than treating a model’s nation of origin as a complete picture.

Implications for labour are widespread

Daniel Kahneman’s work on judgment provides a useful starting point. He studied systematic errors and unwanted variation in human decisions. In a 2018 discussion with Tyler Cowen, he argued that we would reach a point when human involvement would not necessarily improve a system once its performance had been completely validated. He also emphasised structuring decisions and examining evidence before relying on intuition. These were arguments about judgment and automation, but they are useful ways to look at the employment implications, especially code engineering, of today’s language models. [6] he describes models becoming so effective at certain functions that they will be improvements upon human judgement, something which has already largely manifested because of pattern recognition capabilities

He uses the word “noise” to describe the inefficient parts of human labour and unwanted variability in judgment. Administrative busywork is a different problem, although an organisation can suffer from both. AI might reduce document handling while still producing inconsistent judgment. Each benefit should be measured separately.

There is already direct evidence that existing language-based systems can change work. Erik Brynjolfsson, Danielle Li and Lindsey Raymond studied the introduction of an AI assistant among 5,172 customer-support agents. Productivity, measured as issues resolved per hour, increased by 15% on average. Less experienced workers gained more, while the most experienced workers saw small speed gains and small quality declines. [7]

That finding establishes a vision of how labour markets will be affected, quite different from “all jobs will disappear” and more like what we are seeing manifest already: a system can materially change performance in an existing occupation by assisting and coaching and being delegated to in part without being capable of performing every part of it. The next economic change can come from distributing useful capabilities, integrating them into workflows, improving workflow, elevating human involvement and reducing cost.

A service business contains of these workflows: gathering information, checking documents, preparing estimates, updating records, scheduling appointments and following up with customers. An example is a small repair company. Its owner’s specialist work may be diagnosing and fixing equipment. Reducing the time spent preparing quotes, checking work and chasing paperwork could make the business easier to operate. This is an illustration, not evidence that any model can already handle an entire process reliably.

Kolibri offers a concrete example of how this development process works: independently trained weights, teaching examples from other models, specialised training and a release that others can operate. The broader opportunity is to make useful capabilities available to more people and let them apply their own expertise. We already have evidence of task-level gains. Turning those gains into better working lives is work that can begin now.

Sources

1. Aleph Alpha: Kolibri Has Landed.

2. Kolibri model card and training details.

3. Kolibri technical report, especially supervised fine-tuning and reinforcement-learning sections. Stage numbering in this article is an explanatory grouping of the reported process.

4. Alibaba’s reported three-billion-download milestone, The Business Times/Bloomberg.

5. Hugging Face: State of Open Models, Summer 2026.

6. Daniel Kahneman on Cutting Through the Noise, Conversations with Tyler, December 19, 2018.

7. Brynjolfsson, Li and Raymond: Generative AI at Work, revised November 2024; published in the Quarterly Journal of Economics in 2025.

8. Meta: V-JEPA 2 and physical reasoning.

Featured

Jennifer Evans
Jennifer Evanshttps://patternpulse.ai
Principal, patternpulse.ai, and cofounder, Tech Reset Canada. AI policy, research and analysis. Entrepreneur since 2002, marketer since 1998, machine learning since 2009. Based in Toronto and Southeast Asia.