Your own model is cheaper than you think
There's a bill no one talks about openly when building with AI. Every time a document needs to be read, understood, and added to a knowledge base, it costs money. Not much per document, but many documents. And it's routine work billed at the price of the most expensive model on the market.
At broberg.ai, we asked a simple question: does it really have to be the smartest model in the world doing this work?
The question became concrete in Trail, one of our flagship products.
The bill is at the doorstep
First, a clarification, because many people think the opposite. A knowledge base in Trail doesn’t cost money every month because the knowledge needs to be maintained. A Neuron, once compiled, stays there. It only costs again if new facts make it outdated and need to be entered.
The bill is at the doorstep. In Trail, sources aren’t chopped up and indexed. They’re read by a language model, which writes a lasting graph of Neurons on your behalf. There’s no vector index to rebuild or embedding drift to monitor. But that reading does cost. For a single source, it’s pocket change. For a company uploading thousands of reports, emails, PDFs, and meeting minutes into its brain, it’s no longer pocket change.
That’s the bill Trail Scout removes.
Curation is the answer key
What makes it possible is something Trail already does. The model suggests, humans decide, and every change can be verified and rolled back.
All those decisions are effectively an answer key. Every approval, correction, and rejection is a human saying what was right. And an answer key is exactly what you need to teach a small model to do the job.
The method is called distillation. A large model is the teacher, a small model is the student, and the student learns to perform that one task just as well as the teacher. Not writing poems or solving math olympiads. Just reading a source and compiling it correctly into Neurons and connections.
Our own student model is called Scout, because it goes ahead on the path and clears the terrain before the rest arrive.
Quality lies in the framework
The obvious objection is that a small model must be worse. It is, if you let it write freely. But a large part of quality lies in the framework around the model.
Two examples from our work. The model must point to the exact passage in the source behind each Neuron, and a machine check afterward verifies that the passage actually exists there. If it can’t be found, the Neuron is removed. And the model gets a list of existing Neurons so it can only connect to something that exists. It can’t invent a connection.
None of this requires a large model. It requires proper system design. And it makes a small, specialized model surprisingly credible in its own task.
From one project to a platform
This is where it gets interesting as a product. We’re not building Scout as one model, but as the first of many.
A company’s knowledge isn’t like anyone else’s. A consulting engineer, a physiotherapy clinic, and a shipping company have vastly different concepts, document types, and relationships. The general frontier model knows them all a little. A model trained on the customer’s own curated knowledge knows them well.
Trail Scout therefore becomes a platform for both training and inference of customer-specific models. The same approach we use for our own model, applied to the customer’s own brain:
- The base model is always open source. No customer gets locked into a vendor, and the license is known from the start.
- The customer’s curated knowledge is the training material. It’s already there, as a byproduct of normal use.
- The model runs where the customer wants it. On their own machines or in the EU.
- The price per ingest drops drastically because the frontier model no longer has to read every single document.
The frontier model doesn’t disappear. It becomes the emergency exit. When the small model is unsure, or something looks wrong, the task goes to the large model. How often that happens is the metric we track, and it should fall month by month as new curation improves the model.
We take our own medicine
We’re in the middle of it right now. The plan is in place, the training material is being built, and the first models aren’t trained yet. So we don’t yet know exactly how close we can get to the frontier model on our own task. That’s why we measure it properly from the start, against an answer key that’s never been used for training, instead of guessing.
But the direction is hard to avoid. The work the most expensive models in the world do today is largely routine work on narrow tasks. And that’s exactly what you can teach a small model to do well, cheaply, and close to your own data.
As always, we build it for ourselves first. Scout has to run on our own brains before it’s offered to anyone else’s.