Trends · 2026-09-18 · 1:42

Schema-safe is not truth-safe: a model can be confidently, validly wrong

A new class of "typed decision" models makes an invalid answer impossible: the output is guaranteed to fit the schema you declared, with a calibrated confidence attached. That kills the usual LLM-plus-regex-plus-retry loop for classification and routing. It does not kill the other failure. The demo itself shows a decision at 0.96 confidence that is simply wrong: perfectly formed, and false.

In favour
  • Guaranteed-schema output with calibrated confidence is clean engineering for routing, triage and classification, where the shape of the answer matters as much as the answer.
  • Confidence you can threshold on is better than a free-text "I think", provided you measure the calibration yourself.
Worth watching
  • Type safety protects the schema, not the truth; a threshold on confidence alone does not catch the confident error.
  • The vendor's headline speed-up (nearly two hundred times) against an independent measurement of roughly three times is the usual gap between a slide and a benchmark.
Our takeTake the idea, not the dependency: schema-constrained decoding exists for open models you can run yourself (Outlines, GBNF grammars, XGrammar). Same guarantee on the shape, and the model stays on your hardware.

Source: Ο Jev για τις type-safe αποφάσεις ΤΝ

← All Focus posts