Story 02 October 4, 2026 issue
Daily AI-generated issue
Aleph Alpha's open Kolibri uses 11.2% fewer German tokens than GPT-5's tokenizer
Top News · 410 HN points
Need a strong German and English model you can run on your own hardware? Aleph Alpha launched Kolibri 1 as an open-weight model trained from scratch in Germany and Finland. In Aleph Alpha's evaluation, it scored 75.5 in English and 70.8 in German overall, ahead of Qwen3.5 35B-A3B at 74.7 and 69.8.
Here's what changed:
- Its UniBPE tokenizer uses 11.2% fewer tokens on German text than GPT-5's tokenizer, in Aleph Alpha's measurement.
- On AIME 2025 it hit 96.9 in English and 87.5 in German, against 89.6 and 84.4 for the closest model with about 3B active parameters.
- It is an MoE with 78.1B total and 3.46B active parameters per token.
- Context is 262,144 tokens natively, validated up to 1,048,576 tokens.
- On Omniscience, it admitted uncertainty or gave a partial answer 44% of the time.
One catch: coding and multi-turn tool calling trail rivals. It scored 66.4 on SWE-bench Verified versus 73.8 for Qwen3.6 35B-A3B.
Try it: pip install 'aleph-alpha-inference>=1', then vllm serve Aleph-Alpha/Kolibri-1 (Apache 2.0 weights on Hugging Face, about 78 GB of GPU memory).
This issue is researched and written by AI models, and every fact is checked against its cited source. No human edits it before it is sent.