What Is Prompt Optimization and How Do You Improve Prompts Systematically?
Prompt optimization is the practice of improving a prompt through repeated measurement rather than intuition. Instead of rewriting a prompt and hoping it works better, you define what "better" means, run the current prompt and its variants against a fixed set of test cases, and keep the version that scores highest. This matters most when a prompt is already in production, when multiple people edit it, or when small wording changes have caused silent regressions. It is less necessary for one-off prompts you will never reuse.
Prompt engineering vs. prompt optimization
These terms get used interchangeably, but they describe different activities.
| Prompt engineering | Prompt optimization | |
|---|---|---|
| Goal | Get a good result once | Get the best result reliably |
| Method | Rewrite, test by eye | Compare variants against a scored eval set |
| Evidence | Your judgment on one output | Task accuracy, format compliance, latency, cost |
| Output | A prompt that seems to work | A prompt version you can defend and roll back |
| Fails when | The task is high-stakes or high-volume | You have no baseline or no test cases |
Prompt engineering produces a candidate. Prompt optimization decides whether that candidate is actually better.
The optimization loop
The loop has four stages, and you repeat it until improvements stop mattering.
1. Define success criteria
Before touching the prompt, write down what a correct output looks like. Be specific enough that two people would agree on whether a given output passes. Useful criteria include:
- Task accuracy — did the model produce the right answer, classification, or extraction?
- Format compliance — did it follow the required structure (JSON, a fixed number of bullets, a specific field set)?
- Latency — how long does a run take end to end?
- Cost — how many tokens does the prompt and completion consume?
Pick the one or two criteria that actually matter for your use case. A support-triage prompt cares about classification accuracy; a summarization prompt may care more about format and length.
2. Build a small eval set
You do not need thousands of examples to start. Twenty to fifty representative inputs, each with an expected output or a pass/fail rule, will expose most prompt problems. Include:
- Common cases the prompt handles today
- Edge cases you have seen fail
- At least a few adversarial or ambiguous inputs
Keep this set fixed while you compare variants. If you change the test cases between runs, you cannot tell whether the prompt improved or the test got easier.
3. Compare variants
Change one thing at a time so you know what caused the difference. Run each variant against the same eval set and record the scores. A variant only wins if it improves your primary criterion without unacceptable regression on the others — a prompt that gains 5% accuracy but doubles latency may not be worth shipping.
4. Promote the best version
Once a variant wins, make it the production prompt and record what changed and why. This version history is what lets you roll back when a later edit turns out to be worse.
Techniques that usually move the score
These are the changes most likely to show up as measurable gains:
- Clarify instructions. Replace vague verbs with explicit ones. "Summarize this" becomes "Summarize this in exactly three sentences, each under 20 words."
- Add few-shot examples. Two or three well-chosen input/output pairs often fix format and tone problems faster than more instruction text.
- Specify an output schema. If you need structured output, state the exact fields and types rather than describing the shape in prose.
- Decompose the task. Split a prompt that asks for extraction, reasoning, and formatting into separate steps. Each step is easier to evaluate and debug.
- Constrain the scope. Tell the model what to do when the input does not fit — for example, return a specific error value instead of guessing.
Each of these is a hypothesis. Test it against the eval set rather than assuming it helps.
Where observability feeds back in
Evals tell you how a prompt performs on cases you chose. Observability tells you what happens on cases you did not choose. Run traces from production show the actual inputs, the prompt version used, the model's output, and how long each step took. When a prompt starts failing in production, traces show you the input that broke it.
That failing input becomes a new eval case. This is the feedback loop: production traces surface real failures, those failures join the eval set, and the next optimization round is measured against them. Without this loop, your eval set drifts away from reality and your scores stop predicting production behavior.
Platforms that combine routing, observability, and automated evals — Respan describes itself as an AI router with built-in observability and automated evals — are built around exactly this cycle, so that traces and eval results live in the same place as the prompt versions they describe.
Common pitfalls
- Overfitting to a handful of examples. If you tune the prompt until it passes your five test cases, it will fail on the sixth. Keep some cases out of the tuning set and use them only for final checks.
- Silent regressions. A wording change that fixes one case often breaks another. Always re-run the full eval set after any edit, not just the case you were fixing.
- No baseline. If you never scored the original prompt, you cannot prove the new one is better. Score first, then change.
- Optimizing the wrong metric. A prompt that maximizes accuracy but triples cost per call may be a net loss. Decide which trade-offs are acceptable before you start.
- Changing the eval set mid-comparison. This invalidates every comparison you made before the change.
A practical starting sequence
- Pick one prompt that is already in use and matters.
- Write down its success criteria and collect 20–50 test inputs with expected outputs.
- Score the current prompt. This is your baseline.
- Make one change, re-score, and keep the change only if the primary metric improves without breaking the others.
- Repeat until gains flatten.
- Ship the best version, record the change, and add production failures to the eval set as they appear.
The core discipline is simple: never accept a prompt change you cannot measure. Everything else — few-shot examples, schemas, decomposition — is a tool for generating candidates worth measuring.