Look at any production system that calls a language model and find the prompt. It will be in a string literal, or a text file that nobody has opened in three months. It has no tests. Nobody can tell you whether the edit made last Tuesday improved anything, because nothing was measured before or after. It was changed because somebody was annoyed with an output on a Thursday afternoon. This is the only component in the stack we would tolerate that from, and we tolerate it because prompts look like prose rather than like code.

What "programming instead of prompting" actually means

There is a family of tools, the best known of which is DSPy from the Stanford NLP group, built on a simple inversion. You do not write the prompt. You declare what goes in and what comes out, you supply a metric and a set of examples, and an optimiser searches for the instruction text and the demonstrations that maximise the metric.

The prompt stops being the thing you author and becomes the thing the system produces. It is closer to a compiler output than to a document: you maintain the specification and the tests, and the artefact is generated. If that sounds like a small change of emphasis, consider what it forces you to have. You cannot optimise without a metric. You cannot have a metric without deciding what good looks like. You cannot evaluate without a dataset. Most teams have none of the three, which is precisely why their prompts are unmeasured.

The three preconditions, stated honestly

This is not a quick win, and the projects we have seen fail did so by skipping straight to installation.

A metric that correlates with what you care about. For extraction and classification this is easy: exact match, field-level accuracy, an F-score. For generation it is hard, and a bad metric is worse than none, because an optimiser will maximise exactly what you asked for. If your metric rewards length, you will get length. This is the same failure we described in agent security as reward hacking, arriving here by invitation rather than by accident.

An evaluation set. In practice fifty to two hundred labelled examples, split so that you keep some back. Building it is a day of unglamorous work with someone who knows the domain, and it is the single most valuable day in the whole exercise, because the evaluation set outlives the framework, the model and probably the project.

A task that recurs. Optimisation is worth it where the same shape of work runs hundreds or thousands of times. For a prompt you will use twice, writing it by hand is correct and everything in this article is overhead.

Where it pays

The pattern is consistent: structured, repeated, measurable. Extraction from documents into fields. Classification and routing, where a support ticket or an email is sorted into one of a known set. Normalising messy input into a schema. Anything that produces a record rather than a paragraph.

It pays badly for one-off drafting, for open-ended writing, and for tasks where the definition of success lives in a person's judgement and changes with context. If you cannot write down what a good output is, an optimiser cannot find one.

There is also a quieter benefit worth naming. Once the prompt is generated from a specification, swapping the underlying model stops being a rewrite. You re-run the optimiser against the new model and compare metrics. Teams that hand-tune prompts are, without intending to, building a switching cost into their own stack.

The part you should steal even if you never install anything

Most of the value here is not in the framework. It is in four habits, and you can adopt all of them this month with no new dependency.

  1. Put the prompt in version control as its own file, not embedded in application code, with a comment saying what it is for and who owns it.
  2. Write down the metric, even informally: what makes an output right, and what the unacceptable failure is. One paragraph.
  3. Build the evaluation set. Thirty real inputs with known-good outputs beats a perfect framework with none.
  4. Run it before and after every change. This converts prompt editing from an opinion into an experiment, and it is the whole game.

A team that does those four things has captured most of the benefit. A team that installs the framework without them has acquired a dependency and a false sense of rigour.

The sovereignty angle

Because these frameworks are backend-agnostic, the optimisation can run against a model on your own hardware through a local server rather than a commercial endpoint. That matters for two reasons.

The obvious one is confidentiality: optimisation means sending your evaluation set, which is by construction a distilled sample of your real work, through the model many times. Doing that against a third-party endpoint is a larger disclosure than any individual prompt you would have thought twice about. The less obvious one is ownership. What you get at the end is a text artefact and a set of demonstrations that live in your repository. That is a thing you own, can diff, can review, and can carry to a different model. It is the opposite of the arrangement where your accumulated prompt craft is spread across a vendor's console.

We set out what it takes to run the local side of this in The Private AI Assistant.

Three traps

Overfitting. An optimiser with a small evaluation set will find the prompt that wins on those examples. Hold out a test set you never optimise against, and treat the gap between the two scores as your honesty measure.

Opacity. Optimised prompts are frequently strange: oddly phrased instructions, examples chosen for reasons nobody can articulate. You end up with a better prompt that you understand less well than the one you wrote. That is a real cost when something goes wrong at three in the morning, and it is an argument for keeping the specification and the metric legible even when the output is not.

Solving the wrong half. If the model is failing because your retrieval is returning the wrong documents, no amount of prompt optimisation will fix it. Measure first, and find out which component is actually losing you points, before optimising the one that is easiest to optimise.

Where to start

Pick the single prompt in your organisation that runs most often and breaks most embarrassingly. Spend one day building a hundred examples with known-good answers and writing down the metric. Measure the current prompt against it, and write the number on the wall.

That number is the point. Whether you then reach for a framework or keep editing by hand matters far less than the fact that you can now tell the difference. The prompt has become code, not because of the tooling, but because it finally has a test.


This article grew out of a Focus note on DSPy from September 2026. Related reading: Verification Overhead on why an evaluation set is what makes checking cheap, and The Private AI Assistant.