
A polished AI demonstration leaves several questions unanswered. Does the system handle the task's awkward cases? Can a reviewer distinguish a supported answer from a plausible invention? How much work remains after generation? A useful evaluation makes those questions measurable before a team relies on a tool, an LLM configuration, or a reusable prompt.
Begin with the AI screening framework, then evaluate the complete workflow below. Treat the model, prompt, supplied context, tools, and human review as parts of the same candidate. Changing any one of them can change what is being evaluated, so the record should describe the configuration that actually produced each result.
1. Define a task and an observable acceptance rule
Write a short task specification with the intended user, input, output, and permitted actions. Replace broad goals such as “produce good analysis” with observable requirements. For a document extraction task, those might include capturing specified fields, retaining the source location, using a defined format, and leaving unavailable information explicitly unresolved.
The NIST AI Risk Management Framework 1.0 organizes risk work around governance, context mapping, measurement, and management. It calls for evaluation methods, documented limitations, and monitoring appropriate to the system's use. The procedure here applies those broad ideas to a practical comparison of AI workflows; its particular test design and scoring choices should be adapted to the task.
Define unacceptable outcomes before reviewing candidates. An unsupported source reference, an unauthorized action, or disclosure outside the intended audience may deserve a separate blocking condition. Avoid letting an excellent writing score compensate for a failure that makes the output unusable.
2. Build a test set around the work and its exceptions
Collect authorized examples that represent routine work, difficult work, incomplete inputs, and cases in which the system should ask for clarification or decline to infer an answer. Describe why each example belongs in the set. Include variations in document structure, terminology, and length that matter to the actual workflow.
Create an expected-result record for each case. This may be a reference answer, required facts, acceptable alternatives, or a checklist of conditions. Some tasks admit several good outputs, so a single preferred wording can be too narrow. Document the basis for judgment and keep disputed cases visible until the rubric can handle them consistently.
Keep a development set for prompt improvements and a separate evaluation set for comparison. Once a case has repeatedly guided revisions, do not keep presenting it as unseen evidence. Preserve some difficult examples for a later check, and record any overlap with demonstration material supplied by a vendor or prompt author.
For a document assistant, include an input containing irrelevant instructions inside quoted source material. The expected behavior should remain tied to the actual task and authorized instructions. This tests a concrete boundary without requiring the evaluator to expose real confidential information.
3. Freeze the configuration and preserve a run record
Assign each candidate a version and save the exact prompt, model identifier, available settings, input preparation, retrieval material, and permitted tools. Record the date and environment. Where a provider does not expose a detail, mark it unavailable rather than claiming the configuration is fully reproducible.
For every run, retain the case identifier, candidate version, raw output, errors, elapsed time, and any measurable usage. Store sensitive examples under the controls appropriate to their source. A screenshot of a successful answer is insufficient for reconstructing the input, intermediate context, and later scoring decision.
Repeat selected cases when variation matters. Preserve all attempts, including failures and retries, rather than choosing the most favorable response. Reproducibility here means that another reviewer can understand and repeat the procedure within its stated limits; it does not promise identical text from every execution. The LLM screening criteria help keep model access and configuration assumptions visible.
4. Score outputs without candidate labels
Give outputs neutral identifiers and randomize their presentation order before human scoring. Keep the candidate mapping separate until judgments are recorded. This design helps the reviewer focus on the task evidence rather than a preferred brand, an impressive demo, or familiarity with the person who wrote the prompt.
Use a small rubric with clearly described levels. For extraction, score factual correctness, completeness, source support, and format compliance separately. For drafting, distinguish factual support from tone and usefulness. Include examples of what counts as a material error so reviewers do not apply incompatible standards.
Have reviewers explain disagreements on a shared subset of cases before scoring the full set. If the rubric changes, identify which earlier judgments need reconsideration. Report the number of evaluated cases, missing runs, and unresolved scoring disputes. An average without that context can hide weak evidence or a small number of consequential failures.
Keep the comparison paired
Evaluate candidates against the same case identifiers, source material, and acceptance rules wherever the comparison requires equal conditions. If one candidate receives additional context or manual preparation, record that difference and include the extra work in its operating cost. Decide whether you are comparing models under equivalent inputs or complete workflows with different preparation steps.
Review changes case by case before looking at a summary. Identify where a revised candidate resolves an error, introduces a new one, or leaves the result unchanged. Preserve missing and failed runs in a separate reporting category, with retries linked to the original attempt. This makes it possible to distinguish a broad improvement from a narrow gain accompanied by regressions elsewhere, and gives the next revision a specific set of cases to investigate.
5. Diagnose errors before rewriting the prompt
Classify each failed case by the observed problem and its practical consequence. Useful categories include unsupported content, omitted information, wrong source selection, malformed output, failed tool use, excessive refusal, and action beyond scope. Keep observed behavior separate from a hypothesis about its cause.
Then inspect the workflow layer most likely to explain the failure. Was the necessary source present? Was it retrieved correctly? Did the prompt describe the required output? Did the system receive conflicting instructions? Was the scoring reference itself wrong? Changing the prompt before answering these questions can conceal a data or evaluation problem.
Make one interpretable revision at a time when practical and rerun relevant development cases. After the change stabilizes, evaluate it against the reserved comparison set. The prompt screening guide provides a place to connect a reusable prompt with its version history, test evidence, and known limits.
6. Count the cost of accepted work
Measure cost across the complete task. Include generation, retrieval or tool use, retries, human checking, correction, and any required preparation. Use observed usage and the applicable pricing or internal cost assumptions, with dates attached. Keep one-time setup separate from recurring processing so comparisons remain interpretable.
A useful accounting measure is total measured task cost divided by the number of outputs that meet the acceptance rule. Define both terms before calculating. If no output passes, report that condition rather than inventing a meaningful cost per accepted output. Also record time to an accepted result, since a fast first draft can still require lengthy review.
For a hypothetical comparison, a cheaper generation step may require more checking and retries. The worksheet should show that tradeoff without assuming which candidate wins. Present quality, critical failures, cost, and completion time together; avoid collapsing them into one score unless the weighting has a clear, documented purpose.
7. Set a bounded human review process
Specify which outputs require review, who can approve them, what the reviewer must inspect, and when the workflow stops. An instruction to “check everything” is incomplete if the reviewer lacks the source material, time, or authority needed to resolve an error. Give reviewers a concrete checklist tied to the acceptance rule.
Define escalation triggers for unsupported claims, missing evidence, unexpected actions, and repeated failures. Establish what happens when the review queue exceeds its capacity. The response may be to narrow the task or pause affected processing; it should not silently become approval by default.
Keep a record of the decision and the version reviewed. If a candidate is accepted for a limited use, state that boundary and the changes that require reevaluation, such as a new model version, a different data source, or broader tool permissions.
Conclusion: make each comparison explainable
A useful AI evaluation connects a defined task to representative cases, independent scoring, traceable versions, and the actual effort needed for accepted work. Keep failures alongside successes and record the limits of every conclusion. That evidence makes the next decision easier: revise a prompt, repair a data step, narrow the task, or investigate a candidate further.
https://digitalassetscreener.com/blog/ai-llm-prompt-evaluation-workflow/


