Testing a model on your own work
Leaderboards tell you a model is good at leaderboards. Forty real cases from your own backlog, graded honestly, answer the question you actually have.
A new model lands, the benchmark chart is up and to the right, and somebody asks whether we should switch. The chart cannot answer that. It measures a distribution of tasks that is not yours, scored in a way that is not how you would score it, on problems the model may well have seen during training.
The replacement is unglamorous and takes about a day: a few dozen real cases from your own work, a grader you trust, and the discipline to re-run it. This is the same move as building your own oracle when there is no test suite, pointed at the model instead of the renderer.
Why the public numbers don’t transfer
Three separate problems, and they compound.
The task isn’t yours. A coding benchmark measures self-contained problems with a test harness attached. Your job might be “read this alert storm and say which service broke”, which nobody benchmarks. Strength on one says surprisingly little about the other.
Contamination. Public test sets are on the public internet, and models are trained on the public internet. Nobody fully solves this. A high score can mean capability, or it can mean the answers were in the training data, and from the outside you cannot tell the two apart.
Small differences aren’t differences. Two points between models on a leaderboard is well inside the noise of the measurement, and it is certainly inside the noise of your workload. People reorganise their stack over gaps that would not survive a re-run.
None of this makes benchmarks worthless. They are a coarse filter: they will tell you which handful of models are in the right league. They will not tell you which one to put behind your ticket triage.
Forty cases, out of the work you already have
Go through last month’s tickets, log excerpts, documents, support threads, whatever the model will actually be pointed at, and pull out 30 to 50 of them. Not invented examples. Real inputs, with their real mess intact.
For each one, write down what a good answer looks like. Not the exact wording, the outcome: which service it should name, which fields it should extract, that it should refuse, that it should say it doesn’t know.
The mix matters more than the count:
- The typical case, several times over. This is most of your volume.
- The nasty ones. The ticket with three problems in it, the log line with a red herring, the document with a contradiction. These are where models separate.
- The ones with no good answer. The question your data cannot answer. You are checking that it says so instead of inventing something.
- The ones it must not touch. Anything where the right move is to escalate to a human.
Keep the whole thing in git next to the prompt. The prompt, the model name, the tool definitions and the case file are one versioned artifact, because changing any one of them changes the results.
Grading, in three tiers
Use the cheapest grader that actually works, and reach for a fancier one only when you must.
- Machine-checkable. Does the JSON parse and match the schema? Does the suggested command run? Is the extracted ID the right ID? This is free, deterministic and boring, and you should push as much of the eval into this tier as you can, even if it means asking the model for structured output you then check mechanically.
- A human with a rubric. Fifty prose answers takes an hour to read. Write three yes/no criteria per case instead of a 1-10 score, because scores drift between sessions and yes/no does not. Do this at least once yourself before you automate it away: it is where you learn what your model is actually bad at.
- A model as judge. Necessary at volume, and it comes with known biases: judges prefer longer answers, prefer answers that look like their own, and in pairwise comparisons are influenced by which candidate came first. Make it grade one answer at a time against a reference and an explicit rubric, swap the order when you do compare, and spot-check twenty of its verdicts against your own. If the judge disagrees with you on a fifth of them, its numbers are decoration.
What forty cases can and cannot tell you
This is the part that gets skipped, so here it is with numbers on it. Pass rate is a proportion, and proportions measured on small samples are noisy. At a pass rate around 80% on 40 cases, one standard error is about six points.
So do not use a small set to crown a winner. Use it for the two things it is genuinely good at:
Catching regressions that matter. If a prompt change or a model swap takes you from 34/40 to 19/40, that is not noise, that is a broken system, and you found it before your users did.
Telling you the shape of the failures. This is the real prize and it needs no statistics at all. Sort the failures into buckets and read them. Twelve failures all involving dates is a fixable problem. Twelve failures spread evenly across everything is a model that does not suit the job.
While you are in there, record more than pass rate: latency at the median and the tail, cost per case, and, most importantly, how the wrong answers were wrong. An answer that is obviously wrong costs a reviewer two seconds. An answer that is plausibly wrong costs an incident. Those should never share a column.
Make it a gate, not an event
Run the set on a schedule and on every change to the prompt, the model, the tool schemas or the retrieval layer. Pin temperature to zero to cut variance, while knowing that identical output is not guaranteed even then, since batching and floating-point order shift underneath you on the server side.
Then let it earn its keep. When a provider deprecates your model with six weeks’ notice, or a new one arrives that costs a third as much, the question “is this safe to switch to” turns from a week of nervous anecdote into an afternoon: run the file, read the failures, decide. A day of work, cashed in at every upgrade for as long as the system lives. On that arithmetic there is no excuse for not having one.