Prompt Evaluation Checklist: Score LLM Outputs Before Production
A practical checklist for evaluating prompt quality across AI models—instruction following, format compliance, grounding, safety, and cost—before you ship to production.
Choosing a model from a demo is easy; trusting it with your production prompts is not. A prompt evaluation checklist turns subjective “this looks fine” reviews into repeatable scores you can compare across OpenAI, Anthropic, Google, and future providers.
Use this guide with side-by-side runs in Smart AI Comparison and the downloadable AI model evaluation scorecard.
Before You Run Comparisons
Fix the prompt under test
- Prompt version is saved (filename or git tag)
- Success criteria are written in plain language
- Required output format is specified (JSON schema, markdown sections, word limit)
- Grounding rules are explicit (“only use provided context”)
Fix the evaluation set
- At least 5 representative inputs (happy path + edge cases)
- Same inputs will be sent to every model—no provider-specific tweaks during scoring
- Reviewers know the rubric before seeing outputs
Core Evaluation Dimensions
Instruction adherence
- Completes every requested sub-task
- Respects constraints (length, tone, audience, language)
- Does not add unrequested sections or disclaimers that break downstream parsing
Format compliance
- Valid JSON / markdown / CSV when required
- Field names and types match schema
- No trailing commentary outside the requested format
Factual grounding
- Claims in the answer appear in supplied context (when RAG or inline source is used)
- Unknowns are flagged instead of invented
- Citations or quotes match source text when requested
Quality and usefulness
- Correct for the business task (not just grammatically fine)
- Appropriate depth—not verbose filler or overly terse gaps
- Safe refusal when policy boundaries apply (explain why, suggest alternative)
Operational fit
- Latency acceptable for your UX (chat vs. batch)
- Token usage and estimated cost fit budget at expected volume
- Stable enough across 2–3 repeat runs on the same prompt
Scoring Workflow
- Baseline parallel run — Select 2–4 models in Smart AI Comparison; submit the same prompt for each test input.
- Blind score (optional) — Copy outputs without model labels; score 1–5 per dimension.
- Record failures — Note refusal, hallucination, format break, or latency spike with the exact model ID.
- Iterate prompt once — Fix the prompt for observed failure modes; re-run all models (avoid tuning for one vendor only).
- Freeze version — When rubric thresholds pass, tag the prompt and model list for production.
Use the evaluation scorecard to track scores per model and per test case.
Red Flags That Disqualify a Model
- Repeated JSON or schema failures after two prompt revisions
- Systematic hallucination on your domain entities
- Unacceptable refusal rate on legitimate tasks
- Cost or latency 2× above alternatives with no quality gain
Side-by-Side in Smart AI Comparison
- Connect BYOK keys for providers you want to test.
- Open Compare, pick category (text/image/video/audio as supported).
- Paste your fixed prompt; select models.
- Review outputs in the grid; export or share results with your team.
- Log scores in the scorecard; repeat after provider model updates.
Free tier includes limited daily comparisons; Pro unlocks unlimited runs for full evaluation sets.
Related guides
- Prompt testing across multiple models
- How to measure AI response consistency
- Audio and voice AI evaluation checklist