Confident at 0.91, right by coin toss
A Carnegie Mellon team put a cheap decision-only judge against sixteen others, including frontier models. It costs $0.044 per 1,000 judgments and answers in 0.15 seconds, against $12.182 and 1.89 seconds: 277 times cheaper.
In favour- Where the verdict can be read off the text it lands within about three points of the frontier model, and above 0.9 confidence the two are nearly tied, 94.7 against 94.4 per cent on 3,744 items.
- The cascade works and was pre-registered: accept when confident, escalate when not, 0.9 points more accurate than the frontier judge at 41 per cent of its fee.
- On reference-free prose it is near chance and sure of itself: 53.5 per cent agreement with the labels at a mean maximum probability of 0.91, with an error-detection AUROC of 0.498, which is no signal at all. Two frontier judges did the same at 0.94 and 0.96.
- Style fools it. When the worse answer is the better written one it trails by 9 to 23 points. Set thresholds on your own labels using the lower confidence bound, not the point estimate.
Our takeThe paper's own sentence is the keeper: "An inexpensive judge does not have to be right everywhere; it has to know where it is not." Note also that the work discloses public funding and vendor API credits, and tested one version of one product. Confidence is a number the model reports, not evidence that it is right.
← All Focus posts