Testing and Iterating Your Prompts
Why Testing Matters
A prompt that works perfectly on the examples you tested during development may fail unpredictably in production. Real-world inputs are messy, diverse, and sometimes deliberately adversarial. Testing your prompts systematically before deploying them -- and monitoring them after deployment -- is what separates professional AI engineering from amateur experimentation.
Prompt testing is not optional. It is as essential to AI development as unit testing is to software development.
The Core Problem: LLM Non-Determinism
LLMs are not deterministic by default. The same prompt can produce slightly different outputs each time it runs because of a parameter called temperature. At temperature 0, the model always chooses the most likely next token. At higher temperatures (0.7 to 1.0), the model introduces randomness, making outputs more creative but less consistent.
This means:
- Testing your prompt once is not enough
- You need to test across multiple runs to see the range of outputs
- Edge cases matter as much as the typical case
Building a Test Set
A test set is a collection of representative inputs with known expected outputs that you use to evaluate your prompt's performance.
Step 1: Collect representative inputs
Gather 20 to 50 real or realistic examples of the inputs your system will process. For a customer email classifier, this means collecting 20 to 50 real customer emails across different categories.
Step 2: Label the expected outputs
For each input, define what the correct output should be. For classification tasks, this is the correct category. For extraction tasks, this is the correctly extracted data. For generation tasks, this is harder -- you may define criteria rather than exact outputs.
Step 3: Include edge cases
Deliberately include:
- Ambiguous inputs that could belong to multiple categories
- Very short inputs with little information
- Very long inputs that might exceed context limits
- Unusual formatting (all caps, no punctuation, multiple languages)
- Inputs that are completely outside your expected domain
Evaluating Prompt Performance
For Classification Tasks
Calculate accuracy: the percentage of inputs correctly classified.
Accuracy = (Correct classifications / Total inputs) x 100
A good starting threshold for production is 90% or above, though the right threshold depends on your use case. A medical diagnosis tool requires higher accuracy than a content tagging system.
For Extraction Tasks
Measure precision and recall:
- Precision: Of the items the model extracted, what percentage were correct?
- Recall: Of all the items that should have been extracted, what percentage did the model find?
For Generation Tasks
Evaluation is qualitative. Define a rubric with criteria and score outputs against them:
- Does it include all required information?
- Is it the right length?
- Is the tone appropriate?
- Does it contain any factual errors?
Common Prompt Failure Modes
Understanding why prompts fail helps you fix them faster:
1. Format failures
The model produces output in the wrong format -- extra text around a JSON object, missing fields, inconsistent capitalisation.
Fix: Add "Return ONLY the JSON. No other text." or use explicit delimiters: "Wrap your answer in <answer> tags."
2. Instruction following failures
The model ignores one or more of your instructions.
Fix: Move critical instructions to the beginning or end of the prompt. Repeat key constraints. Use numbered lists for multi-step instructions.
3. Hallucination failures
The model invents information that was not in the input -- names, dates, statistics, or facts.
Fix: Add explicit instructions: "Only use information from the provided text. If information is not present, state that clearly." Use grounding techniques (provide source documents).
4. Ambiguity failures
The model misinterprets ambiguous inputs in ways you did not anticipate.
Fix: Add examples that demonstrate how to handle ambiguous cases. Add a fallback category or default behaviour.
5. Context length failures
Very long inputs get truncated or the model loses track of early context.
Fix: Restructure prompts to put the most important context at the beginning and end (the "primacy and recency" effect). Break long documents into chunks.
The Iteration Cycle
Prompt engineering follows a tight iteration loop:
- Write the initial prompt
- Test against your test set
- Analyse failures -- identify patterns in what fails
- Hypothesise -- why is it failing? What change might fix it?
- Modify the prompt
- Retest -- did the change improve accuracy? Did it break something that was working?
- Repeat until performance meets your threshold
Critical rule: When you change a prompt to fix one type of failure, always retest everything. Changes that fix one problem can introduce new problems elsewhere.
Prompt Versioning
Treat prompts like code. Version control your prompts and document what changed between versions.
A simple versioning approach:
v1: Initial prompt -- 72% accuracy on test set
v2: Added output format specification -- 81% accuracy
v3: Added 3 few-shot examples for ambiguous inputs -- 89% accuracy
v4: Added explicit hallucination guard -- 91% accuracy, 0 hallucinations in test set
Keep a changelog. This is especially important when working in a team, because multiple people may be iterating on the same prompts.
A/B Testing Prompts in Production
Once a prompt is deployed, monitor its performance and run A/B tests to evaluate improvements:
- Route a percentage of production traffic to the new prompt
- Measure the metrics that matter (accuracy, user satisfaction, escalation rate)
- After a statistically meaningful sample, compare results
- Promote the better prompt to 100%
Tools for Prompt Testing
Several tools help automate prompt testing:
- PromptFoo: An open-source tool for testing and comparing prompts
- LangSmith: Part of the LangChain ecosystem for tracing and evaluating LLM outputs
- OpenAI Evals: A framework for evaluating OpenAI model performance on custom tasks
- Custom scripts: For many use cases, a simple Python script that runs your prompt against a test set and calculates accuracy is sufficient
Practice Exercise
Take a prompt you have written in a previous lesson and create a test set of at least 10 inputs. Run the prompt against all 10 inputs, record the outputs, and identify any failures. For each failure, hypothesise why it failed and make one targeted change to the prompt to address it.
Key Takeaways
- Testing prompts systematically is as essential to AI engineering as unit testing is to software engineering.
- LLMs are non-deterministic, so testing your prompt once is not enough -- test across multiple inputs and runs.
- Build a test set of 20 to 50 representative inputs including edge cases before declaring a prompt production-ready.
- Common failure modes include format failures, instruction following failures, hallucinations, ambiguity handling, and context length issues.
- Version-control your prompts, document changes, and always retest everything when you make a modification.
Try it yourself
Key Takeaways
- Prompt testing is as essential to AI engineering as unit testing is to software engineering -- never deploy untested prompts.
- LLM non-determinism means you must test across many diverse inputs, not just a few examples.
- Build a test set of 20 to 50 representative inputs including edge cases before declaring a prompt production-ready.
- Common failure modes include format failures, hallucinations, instruction-following failures, and ambiguity -- each has specific fixes.
- Version-control your prompts, document changes, and always retest the full test set after every modification.
Quick Quiz
1.Why is it important to test a prompt against multiple inputs rather than just one or two examples?
2.What is a 'hallucination' in the context of LLM outputs?
3.What is the recommended approach when you make a change to fix one type of prompt failure?
4.Why should prompts be version-controlled like code?
Ready to go further?
CareerEx gives you structured 12-week training, live classes every Saturday and Sunday, real tutor feedback, and a certificate. Join the next cohort.
Join CareerEx