Independent AI intelligence Two editions daily · ET
Fervor AI

Analysis · October 3, 2026 · concept

Aleph Alpha Kolibri-1AA-Omnisciencefrontier-modelsragagent-infrastructure

Kolibri-1 Learned to Say "I Don't Know." Its Own Benchmark Table Shows the Price

Aleph Alpha's new open model trades knowledge for honesty on AA-Omniscience, and the tradeoff is the most useful thing on the model card

Most model cards bury their bad numbers. Kolibri-1 puts one in the same row as its best one.

Aleph Alpha released Kolibri-1 on October 3, an Apache 2.0 open-weight model that speaks German and English and nothing else. The headline specs are the usual parade: 78.1B total parameters, 3.46B active per token, a 262,144-token native context, 96.9 on AIME 2025. Scroll further and you hit a benchmark most launch posts skip, AA-Omniscience, where Kolibri reports a non-hallucination rate of 44.0. Qwen3.5 35B-A3B, a model it beats on several reasoning tests, scores 11.1.

In the next column, Kolibri's accuracy on the same test is 14.8. Qwen3.5's is 22.2.

So the model that lies less also knows less. That is the trade every team building on retrieval has been asking for without quite saying it, and Kolibri is the clearest public example yet of a lab training for it on purpose and publishing the bill.

Why one row on a model card deserves an essay

Hallucination usually gets discussed as a defect, something a better model will eventually stop doing. AA-Omniscience treats it as a choice. The benchmark, published by Artificial Analysis in November 2025, asks 6,000 questions across 42 topics in six domains, and its scoring is blunt: +1 for a correct answer, -1 for an incorrect answer, 0 for an abstention. Its hallucination rate is defined as the share of incorrect answers among everything the model did not get right. Flip that and you get the non-hallucination rate: of the questions a model could not answer correctly, how often did it decline instead of guessing?

That framing splits two things most leaderboards mash together. Knowledge is how much the model knows. Calibration is whether it knows what it does not know. A model can be strong on one and terrible on the other, and for an agent sitting on top of your document store, the second one decides whether a wrong answer reaches a customer.

My position is simple. For any model that answers from retrieved context, the non-hallucination rate matters more than the closed-book accuracy score, and Kolibri's table is the clearest public example I know of that lets you watch a lab move the first number at the expense of the second.

What Aleph Alpha changed, in numbers

The Kolibri-1 model card compares it against a dozen open models of similar active size, plus a baseline it labels "Kolibri Origin." The card does not explain what Origin is, so treat it as a related baseline rather than a confirmed earlier checkpoint.

Model AA-Omniscience accuracy Omniscience Index Non-hallucination rate
Kolibri Origin 11.3 -64.2 15.0
Kolibri-1 14.8 -32.8 44.0
Qwen3.5 35B-A3B 22.2 -46.2 11.1
GLM-4.7 Flash 30B-A3B 17.0 -62.8 3.8
Qwen3.6 35B-A3B 21.0 -12.5 56.7
Qwen3.8 27B (dense) 19.0 -5.8 67.3

Read the first two rows together. Against the Origin baseline, Kolibri-1's non-hallucination rate is nearly three times higher (44.0 against 15.0) and its Omniscience Index loss is about half as deep (-32.8 against -64.2), while accuracy is only 3.5 points higher. Whatever separates the two models, the gain is mostly in knowing when to stop, not in knowing more facts.

The card describes the method. Aleph Alpha "trained with abstention data" and built reinforcement learning environments around what it calls the Merlin-Arthur protocol, which "shows the model each question with parts of the context hidden." In one set of examples a player called Merlin hides parts that make the right answer more likely, and the model is trained to answer. In the other, Morgana hides the parts the answer depends on, and the model is trained to abstain. In plain terms, the training teaches the model to notice when the evidence it needs is missing.

That is exactly the failure that hurts retrieval systems. A retriever misses the right chunk, the model gets a context that looks relevant but does not contain the answer, and it fills the gap from its weights. Training against hidden context targets that moment directly.

The part the launch framing skips

Kolibri is not the most honest model on its own table.

Qwen3.6 35B-A3B scores 56.7 on non-hallucination with higher accuracy (21.0). The dense Qwen3.8 27B scores 67.3, the best Omniscience Index on the table at -5.8, and it still answers more questions correctly than Kolibri. If abstention is what you want, the table Aleph Alpha published points you at a competitor first.

And every Index score on that table is negative, Kolibri included. On closed-book knowledge questions, all of these models are still wrong more often than right when they choose to answer. Whatever the abstention training bought, it did not turn Kolibri into a knowledge source.

None of that weakens the lesson. It sharpens it. The useful takeaway is that abstention is a trainable, measurable behavior that varies by a factor of more than 15 across models of similar size (3.8 to 67.3 on this one table), and you cannot predict it from accuracy, reasoning scores, or parameter count. GLM-4.7 Flash beats Kolibri on accuracy and almost never abstains. You only learn that by measuring.

There is also a reason Kolibri's tradeoff fits its market. Aleph Alpha sells into German enterprises and public bodies, where a confident wrong answer about a regulation is worse than no answer. The card's retrieval numbers back the use case: 89.7 on its English industry RAG test and 77.3 on MuSiQue. A model that knows less from memory but reads well and declines when the context is thin is a reasonable design for that buyer.

Put this into practice

You do not need Kolibri to use any of this. You need a small eval that scores the way AA-Omniscience does.

1. Build a 60-question set from your own documents. Write 40 questions whose answers sit in your corpus and 20 whose answers do not, but which sound like they should. The second group is the important one. Those are the questions where a model either abstains or invents.

2. Give the model explicit permission to abstain. Put a line in the system prompt such as "If the provided context does not contain the answer, reply exactly: NOT IN CONTEXT." A model cannot score well on abstention if your prompt implies it must always answer.

3. Score three numbers, not one. Count correct, incorrect, and abstained. Then compute:

accuracy            = correct / total
hallucination_rate  = incorrect / (incorrect + abstained)
index               = (correct - incorrect) / total

The second formula is the AA-Omniscience definition. Track all three across every model and prompt change. A prompt tweak that raises accuracy by two points and doubles the hallucination rate is a regression for most retrieval products.

4. Route abstentions instead of treating them as failures. An abstention is a signal: retry retrieval with a wider query, hand off to a larger model, or show the user the sources and say the answer is not there. Each of those beats a fluent fabrication.

5. Re-run the set when you switch models. The spread on Kolibri's table shows that two models with similar benchmark scores can differ wildly in how often they guess. That difference will not show up in a vendor's launch post.

If you do want to try Kolibri itself, the card says it requires the aleph-alpha-inference package, which provides a Kolibri plugin for vLLM, and sampling at temperature 1.0, top_p 0.97, top_k 128.

Honest limitations

These numbers come from Aleph Alpha's own model card. AA-Omniscience is a third-party benchmark, but the runs reported here are the vendor's, on the public question set, and I have not seen an independent replication of the full table. The tej.as analysis that drew a Hacker News thread on October 3 repeats the same figures rather than re-running them, and notes that "partial" answers count toward the non-hallucination side.

A closed-book knowledge test is also not your retrieval system. Abstention measured on trivia-style questions may not carry over to your contracts or support tickets, which is why step 1 above is not optional.

Kolibri has practical limits too. It handles German and English only. The card's minimum hardware is two 80 GB A100s or H100s, or one H200, B200 or B300, so this is a data-center model. Its BFCL v4 multi-turn tool-calling score (47.5) trails several peers, and its Terminal-Bench 2.1 score is 27.7, so it is a poor fit for agent loops that chain many tool calls.

And abstention has a cost you will feel in product. A model that declines more often will frustrate users who wanted an answer, even when declining was correct. Plan the UI for "not found" before you ship a model that says it often.

The number to ask for

Model launches compete on what a model can do. Kolibri's card, maybe by accident, makes a case for publishing what a model will refuse to do, and how often it is right to refuse.

Next time you evaluate a model for anything that answers from your data, ask for the non-hallucination rate alongside accuracy. If the vendor does not publish it, run the 60-question set and find out yourself. It takes an afternoon, and it is the most honest number you will get.

Sources: Kolibri-1 model card, Artificial Analysis on AA-Omniscience, AA-Omniscience paper, tej.as analysis.


Medium metadata

  • Title: Kolibri-1 Learned to Say "I Don't Know." Its Own Benchmark Table Shows the Price
  • Subtitle: Aleph Alpha's new open model trades knowledge for honesty on AA-Omniscience, and the tradeoff is the most useful thing on the model card
  • Tags: Artificial Intelligence, LLM, RAG, Open Source, Machine Learning
  • Canonical: fervorai.dev URL once published