I've been paying a frontier model to tell me that "my card payment failed three times this morning" is a billing question. Every time. At frontier prices.
Daft, right? Most of those calls are easy. So I stuck a small model on my own gaming GPU in front of Claude, let it answer the ones it was sure about, and only sent the rest up the chain. On a thousand banking support messages, nearly three-quarters never left the card (an RTX 3080 Ti, bought for far less serious reasons).
And the GPU turned out to be the boring part. The model that did best was the dullest one in the test, a 22.7M-parameter classifier, and on the banking test it beat every purpose-built decision model I threw at it. Including the hosted one. I'll get to that.
First, the catch that nearly sank the whole idea.
The small print, up front because it matters. This was measured on 27 and 28 September 2026. One dataset (Banking77), one card, one night. The £ figures are estimates from published list prices, not invoices. Claude's answers came from an interactive Claude Code session working through batched answer sheets, not from the API. Every number links to the committed report it came from.
A confidence score that looks certain and means nothing
The plan only works if the local model's confidence means something. If it says 0.9, it has to be right nine times out of ten. Otherwise you're keeping answers because they sound sure, which is how you end up explaining yourself to a customer.
So how sure was laya-en, the open decision model, out of the box? Very. How often was it right? 37.2% of the time, with an expected calibration error (ECE) of 0.502. Its stated confidence sat about 50 points away from its hit rate. It'd tell you it was nearly certain, then be wrong two times in three.
You can't gate on that. Nobody can.
The fix is calibration, and it's less magic than it sounds. It changes what the number means and leaves the model alone. Fitting a small calibrator on a separate 1,000-item split took laya-en's ECE from 0.502 to 0.056, and Von's from 0.185 to 0.021. Accuracy didn't move an inch.
Calibrated isn't the same as good, mind. Raw laya-en still couldn't get its error under 5% at ANY threshold. A well-calibrated model that's wrong most of the time is just honest about being useless.
So I fine-tuned it on the same card
9,003 training items, same GPU, 1,518 seconds. About 25 minutes.
With a threshold of 0.95, picked on the calibration split and judged on the held-out split, the fine-tuned Laya kept 73.0% of decisions local and sent the rest to Claude Opus 5.5. Blended accuracy was 93.3%, against 94.2% for Claude alone, and the estimated bill fell from £3,898 to £1,054 per million decisions.
One point of accuracy for about a quarter of the bill. I'll take that.
Where does £3,898 come from? Each Banking77 prompt is about 1,259 input tokens once you list all 77 intents, which puts Claude Opus 5.5 at £3,898 per million decisions, or £1,949 through the Batch API. Tokens are estimated from characters, not counted. And I haven't costed the GPU, because I already owned it.
Then the boring model walked in
I put a fine-tuned all-MiniLM-L6-v2 through exactly the same measurement, calibration, threshold and cascade code. No special treatment. It's the sort of small sentence encoder people were training years before ChatGPT turned up, and I added it as a baseline, the thing to measure the clever models against.
It won. It scored 91.5% on the held-out set with a calibrated ECE of 0.024, against 87.3% for the fine-tuned Laya. In the cascade it kept 94.7% local at the same 94.2% blended accuracy as Claude alone, for an estimated £207 per million.
Here's the whole thing side by side (from the report):
| First stage | Kept local | Blended accuracy | Est. £ per million |
|---|---|---|---|
| Claude Opus 5.5 alone | 0% | 94.2% | £3,898 |
| Jev, hosted by TypeSafe | 51.3% | 93.6% | £1,952 |
| Fine-tuned Laya | 73.0% | 93.3% | £1,054 |
| Fine-tuned MiniLM | 94.7% | 94.2% | £207 |
(MiniLM isn't served through Tau, so that £207 leaves out local latency and energy.)
So is the decision model rubbish? No. Is MiniLM magic? Also no. On a fixed task with thousands of labelled examples, a small trained classifier is STILL the thing to beat. It's been true for years and it's still true now. Decision models earn their keep on questions you haven't trained for, or when you're asking lots of questions about one piece of text. Banking77 is neither of those. I'd rather a tool told me that than hid it.
What about the hosted one?
TypeSafe's hosted Jev speaks the same /v1/systemone contract, so I pointed the same code at it, on the same items (the responses say jev-1.13.0). The whole thing cost an estimated $0.14 for Banking77 and $0.04 for a second dataset, at the published $0.042 per million input tokens. Cheap as chips.
On Banking77 it's a solid generalist. It scored 79.2% held-out out of the box, ahead of laya-en's 37.2% and Von's 77.1%, with an ECE of 0.093 raw and 0.029 calibrated. Fine-tuning beat it, and MiniLM beat everything.
The second dataset, synthetic support tickets, is where it earned its place. Scored against Claude, it agreed on 52.5% of items out of the box, ahead of every local model. At a 20% target it kept 39.8% of decisions, with the served answers agreeing with Claude on 89.3%, for an estimated £614 per million against £997. Nothing local came close.
Two things to know before you gate on it. Jev rounds its probabilities to 2 decimal places. And a hosted endpoint can't load your calibrator, so its calibrated figures are the calibrator applied offline, on my side of the call. You'd have to do the same.
Is it fast enough?
On the 3080 Ti with CUDA, laya-en answers one question in 18.59 ms, and ten questions about the same text in 114.11 ms. On the i9-11900K's CPU, the single question takes 529.21 ms. Those are FP32 medians, no FP16 tricks. Hosted Jev had a median of 223 ms, but that's over the network from my desk, so it isn't like for like.
What went wrong (quite a lot, actually)
- I fitted the calibrators on rounded numbers. The Runtime rounded probabilities to 4 decimal places, the calibrators were fitted on that, and then applied to the unrounded values. With 77 options most probabilities round to zero, so yes, it mattered. My earlier figures (laya-en 0.065, Von 0.039) came from that bug. Refitted at full precision they're the 0.056 and 0.021 above. I'd also claimed two parts of the code agreed to within 2.4e-4, having only checked inputs that didn't saturate. Oops. The decisions log has the fix.
- The ticket labels are close to noise. Claude agreed with them on 23.8% of items, against a 41.0% majority baseline. A constant guess did better than the frontier model! MiniLM happily "learned" the noise to 55.6%.
- On urgency, nothing local stands in for Claude. Scored against Claude, no local decision model kept more than 0.1% of decisions.
- Three FP32 models on one 12 GB card slowed the last one down. The fine-tuned Laya, measured last, had a median of 15,219 ms per request in its raw phase, against 243 ms in a fresh Runtime. VRAM spill is the likely cause, but I haven't proved it. Accuracy isn't affected, and the published latency comes from the clean run.
- My first MiniLM run scored 80.8% because I stopped it after five epochs with the loss still falling. That flatters whatever you compare it against, so I retrained it properly. Which is how it ended up winning.
- Claude disagreed with the Banking77 labels on 5.8% of items. Some are Claude's mistakes and some are the dataset's.
- Calibration didn't always help. The fine-tuned Laya only went from 0.071 to 0.060, because fine-tuning had already fixed its temperature. laya-en on the tickets missed its target, 0.273 to 0.155, because my selection rule picked isotonic over a temperature fit that scored 0.016 against 0.216 on the calibration split. I set that rule before seeing any held-out result, so I left it alone. Von on the tickets barely moved, 0.045 to 0.042, but it was already close.
Try it on your own decisions
Everything above came out of Tau, which is two things, both Apache-2.0:
-
The Runtime is a .NET server that answers the
/v1/systemonedecision contract locally. It runs the open Laya and Von decision models through ONNX Runtime on CUDA, DirectML or a plain CPU. You send it some text and a set of questions, and it sends back an answer with a probability for each option. -
The Workbench is a command-line tool called
tau. It asks one question of any/v1/systemoneendpoint, local or hosted: can I trust this model's confidence enough to gate on it, and what does that save? The last stage writes an HTML report with the misses left in.
The repo's README walks through fetching a model, exporting it to ONNX and grabbing the ONNX Runtime natives. After that, starting the Runtime on an NVIDIA card is one line:
dotnet run --project src/Tau.Runtime -c Release -- --urls http://localhost:8088 --Tau:Provider=cuda
Leave out --Tau:Provider=cuda and it runs on the CPU. On Windows, --Tau:Provider=directml runs on any DirectX 12 GPU. If the provider you asked for won't load, it refuses to start (I'd much rather that than a silent fall back to the CPU).
Then ask it something:
curl -s http://localhost:8088/v1/systemone \
-H "Content-Type: application/json" \
-d '{
"model": "jev-latest",
"state": "My card payment failed three times this morning and I need it sorted today.",
"questions": {
"category": {
"type": "choice",
"instructions": "Which support category fits this message?",
"criteria": {
"billing": "Billing or payments issue",
"technical": "App or website technical issue",
"account": "Account access or security issue"
}
},
"urgent": { "type": "noul", "instructions": "Does the customer need an answer today?" }
}
}'
jev-latest is an alias that lets the Runtime pick a model, and English text goes to laya-en. You get back a choice with a probability for each option, plus a noul score for the yes/no question. GET /v1/models lists what's loaded.
From C#, the Tau.Client package wraps the same call and maps the answer onto an enum:
using Tau.Client;
using var client = new SystemOneClient(new Uri("http://localhost:8088/"));
var decision = await client.DecideAsync<Category>(
state: "My card payment failed three times this morning and I need it sorted today.",
instructions: "Which support category fits this message?");
Console.WriteLine($"{decision.Value} p={decision.Probabilities[decision.Value]:F2}");
enum Category { Billing, Technical, Account }
One trap I nearly fell straight into. Laya's contract returns a confidence field for choice questions, and it's derived from the entropy of the whole distribution. It is NOT a calibrated probability. Calibrate and threshold on the probability of the chosen answer instead, which is why that example prints Probabilities[decision.Value].
To measure your own decision, you describe it in a decision.yaml (the question, your labelled data, the local models, the frontier model and a target error), then run the stages:
tau label examples/banking77/decision.yaml # ingest cached frontier answers
tau measure examples/banking77/decision.yaml # raw accuracy and ECE, calibration and held-out splits
tau calibrate examples/banking77/decision.yaml # fit temperature and isotonic calibrators
# restart the Runtime with --Tau:CalibratorsDirectory=examples/banking77/calibrators, then:
tau measure examples/banking77/decision.yaml --phase calibrated
tau threshold examples/banking77/decision.yaml # pick the threshold for the target error
tau cascade examples/banking77/decision.yaml # simulate local-first and price it
tau report examples/banking77/decision.yaml # write report.json and report.html
tau run does the lot in order and skips any stage whose output is current. It never calls a paid API unless your spec lists a hosted endpoint under external:, and then only under the budget you set there. The key is read from an environment variable at run time and never written anywhere. If frontier answers are missing, tau label exports batches to be answered and exits with code 2.
Or rerun my Banking77 example end to end, report and all:
./scripts/examples.ps1 -Example banking77
It checks the data and model packages first and prints the exact command for anything missing. Then read the Banking77 report and the tickets report. They were harder on me than I've been here.
The repo is at https://github.com/Fortitude-Group/tau, and the full write-up is on the Fortitude Omnis R&D page. Go on, prove me wrong on your own data.

Top comments (0)