How do you know an AI isn't making things up?
Ask an AI something, and it answers. It also answers when it does not know. That is what makes language models both useful and dangerous: the answer sounds just as confident whether it is right or made up.
The usual response to that problem is a promise: our AI does not make things up.
We give a different answer. Trail is our knowledge engine: the company's extra brain, which both the team and an AI draw their knowledge from. Trail is built to measure whether it makes things up. Every night.
A promise cannot be checked. A measurement can.
You cannot tell from an answer whether it is right. You can only hold it up against something you know is right. That is why all quality control in Trail starts in the same place: with an answer key.
We build test sets from the brain's own data. Questions where we know in advance which answer and which source are the right ones. Then we can count how often Trail gets it right.
But a measurement can lie too. It can produce nice numbers because it is too easy to pass. So there is always a control running alongside it: the same questions with deliberately wrong answers. That control scores 2–3 %. It is the proof that the measurement can actually tell right from wrong. Without it, a good number would mean nothing.
The first thing the measurement found
Trail searches in two ways at once. It searches by words: does what you asked about appear in the text? And it searches by meaning: is the text about what you asked, even in other words? That combination is called hybrid search, and it sounds obvious that two ways are better than one.
They were not. The first measurement showed that hybrid search was slightly worse than plain word search at putting the right answer on top.
We did not guess why. We adjusted how much say each of the two gets, giving word search more weight (meaning got 0.15 of the vote). Then we measured again:
- 59 % right answer among the top five with hybrid, against 53 % with word search alone, on the fair test.
- 75 % against 71 % in a customer's brain.
The price is about half a second extra per search. That choice was made on purpose: a little slower, but better answers. The point is not the number itself. The point is that we would have believed the opposite if we had not measured.
Brain Health: four checks every night
That measurement has become a fixed routine we call Brain Health. Every night, Trail checks a brain in four ways:
- Answers without sources. How many answers were given without pointing to where they came from?
- Trap questions. Does it know the answer when it should? And does it say "I don't know" when it does not?
- Claim checks. Is what the answer says actually in the sources? And is what a Neuron contains actually in the source it was made from?
- Search quality. Does it still find the right thing, measured against the answer key?
Here too there is a built-in control: every night it shows that the judge itself can tell right from wrong. A judge that approves everything is no judge.
The result is shown on the brain's Health page, with numbers and the trend over time. If a brain gets worse, you can see it — and see when it started.
It requires the answers to come back
You cannot measure what you do not see. So it is a condition that everything using Trail reports its answers back.
- The text is kept for 30 days and then deleted automatically. The text is stored encrypted, and Trail's developers have no direct access to read it.
- The numbers are kept forever, so the trend over time can be followed.
- Data is stored in the EU, in Stockholm.
- There is a spending cap per brain per day, so a nightly check never becomes a surprise on the bill.
The most honest measurement
While building the control, we also measured something we had not expected to learn from. We asked two strong language models, Sonnet and Opus, to say which topics a source is about. They agreed only 35 % of the time.
That is the most important lesson in the whole piece of work. If two strong models agree on only a third, quality cannot be measured by asking a model whether an answer seems good. It has to be measured against a concrete answer key. Anything else is a feeling.
One of the most important assets we have
Quality control is not an add-on to Trail. It is an extremely important asset in our development of and research into Trail. A Trail brain holds a customer's entire knowledge, so when we check the quality of our technology, we are also checking the quality of our customers' brains.
It is also what lets us improve without guessing. When we change something in Trail, we can measure whether it got better — as the search example above showed. And if an improvement makes something else worse, we see it straight away.
Why it matters to you
When an AI answers on behalf of your business — to customers, to staff or to yourselves — the question is not whether it can make things up. They all can. The question is whether anyone notices when it does.
In Trail, it gets noticed. At night, with numbers, against an answer key.
Talk it through with Christian if you want to know what it means for your own knowledge.
Trail is the company's extra brain: a knowledge engine that gathers everything the company knows into one base that both the team and an AI can search and answer from. With sources, and with quality measured every night.