Programming language models instead of prompting them
DSPy treats prompts as parameters to be compiled, not hand-tuned text. You declare inputs and outputs, and optimizers tune the prompt (and examples) against a metric and a dataset.
In favour
Backend-agnostic: runs against local models via Ollama or vLLM, so optimization stays on your own hardware.
Turns prompt engineering from craft into something measurable and repeatable.
Worth watching
Not a quick win: you need a metric, an evaluation set, and some patience — it is a framework, not a chat box.
Value shows up on recurring, measurable tasks, not one-off prompts.
Our takeFor anyone who prefers measuring to guessing, this is the tool that formalizes it — best aimed at one task that already runs on a brittle hand-written prompt.