When a team picks a model for a task, the first number they ask for is accuracy. It is the wrong first number.
Accuracy tells you how often the model is right. It does not tell you what a right answer costs, how often the model guesses when it should have stopped, or whether a cheaper model would have done the same job.
Cost per correct answer
Take the same task, invoice matching, and run it across several models on a month of real cases. Score each one on three things:
- How often it matched correctly.
- How often it flagged an invoice it was unsure about, instead of guessing.
- What each correct match cost, in tokens and in the human time spent on the ones it got wrong.
Divide the total cost by the number of correct answers. That is the number to compare. In our benchmarks, models that sit within two points of each other on accuracy can be far apart on this number, because one of them guesses and the other flags.
Why flagging matters more than accuracy
A wrong match that goes through quietly costs more than a flag. Someone finds it later, or nobody does.
A model that says "I'm not sure, a person should look" on the hard two percent is doing the job right.
That holds even if a benchmark counts it as a miss. A controller would rather clear forty flags a month than find one silent error in the ledger.
What to do with the number
Once you have cost per correct answer per task, model choice stops being a debate. Route each task to the model that wins on that task. When a cheaper model catches up, move, without rebuilding the agent.
Our cost-per-correct-answer paper has the full method and the results across the models we test. Read it here.
