A support assistant writes an elegant answer about a return. It includes a friendly greeting, a clear next step and one invented policy exception. A grammar check approves it. A customer satisfaction prompt might approve it too. The answer still creates an obligation the business never offered.
An AI evaluation suite tests whether variable outputs satisfy a defined task across representative and challenging cases. It combines independently specified expectations, retained inputs, suitable graders and explicit release criteria. Its job is to expose consequential failures and support a decision, not produce a reassuring average.
This chapter of The AI Software Factory follows a hypothetical return-policy drafting assistant. It may prepare replies from an approved policy but cannot send them or issue refunds. Software behavior tests cover exact permissions and state changes; this suite examines the quality of the generated draft.
Define the decision the suite must support
A suite can help choose a model, approve a prompt change, evaluate a retrieval system or monitor a live feature. Those decisions need different evidence. A small development set that helps improve a prompt does not independently establish release readiness.
For the hypothetical assistant, the first release question is whether a new prompt improves useful drafting without increasing unsupported policy claims or disclosure risk. Define the task boundary: retrieve approved policy, use the provided order facts, identify missing information and prepare a reviewable reply.
The expectation should include abstention or clarification. An answer that invents a deadline to avoid saying “I cannot confirm” is a failure. A draft that correctly identifies missing order information can be successful even though it does not complete the customer’s desired result immediately.
Write acceptance criteria before comparing candidates. If the team chooses the model first and adjusts the rubric afterward, the evaluation can become a justification exercise. Criteria can change when the task is better understood, but that change should be recorded and earlier results interpreted under their original conditions.
Separate exact invariants from judged quality
Some requirements can be checked deterministically: required output fields, allowed labels, arithmetic, whether an unauthorized tool was called or whether private data appears in a defined location. Enforce exact invariants through application logic and software tests where practical.
Other requirements involve judgment: whether the reply clearly explains a limitation, preserves the customer’s meaning or gives a useful next step. These need a rubric and examples. A single score should not blur the distinction between a malformed schema and a weak explanation.
For the return assistant, policy fidelity is central. The evaluator should inspect whether each consequential statement is supported by the approved policy and supplied facts. Warmth and clarity matter, but they cannot compensate for inventing a refund entitlement.
The suite can therefore use a layered result: exact format checks, critical policy/disclosure checks and graded communication quality. The release decision sees the layers separately. A fluent response with a critical failure remains a critical failure even if its overall quality score is high.
Build cases from the work the feature will receive
Useful cases reflect actual task variation. Include ordinary questions, ambiguous requests, unavailable facts, contradictory inputs and attempts to redirect the assistant. A set composed only of neat single-sentence examples can reward a feature that fails as soon as a customer writes naturally.
Sources can include authorized historical cases, expert-written examples and carefully labeled synthetic inputs. Preserve provenance and the reason each case belongs. Synthetic cases can probe a boundary, but they do not establish how frequently that boundary occurs in production.
OpenAI’s current evaluation guidance recommends task-specific objectives, representative and edge cases, and grading calibrated to human judgments. The methodology is useful independently of a particular vendor platform. This chapter does not require the vendor’s hosted evaluation service.
For the hypothetical assistant, case families might include a valid return within policy, a request after the deadline, missing identity, a conflicting tracking record and a message asking the assistant to ignore the policy. Each family should have cases that vary wording without changing the governing facts.
Record expected outcomes without writing one required sentence
A reference answer can help, but requiring an exact sentence may reject valid alternatives. Define the content obligations: explain the applicable policy, identify unsupported facts, avoid disclosure and propose the authorized next step.
For a missing order number, the expected outcome can require a clarification request and forbid claims about the order. It need not demand one particular greeting. For an expired return window, the expectation can require the correct boundary and any approved exception process without promising an exception the policy does not contain.
Reference material should be specific enough for the evaluator to inspect the answer. “Be helpful and accurate” gives little direction when a draft sounds helpful but invents a concession. Include examples of a successful explanation, a minor clarity defect and a critical unsupported claim.
The rubric also needs a rule for disagreement. A reviewer may find a policy ambiguous. That is a specification issue rather than a model failure with a clear correct answer. Retain the case as unresolved until the policy owner supplies the intended meaning, or exclude it from the release metric while reporting the gap.
Protect evaluation cases from development leakage
A team improves a prompt by examining failures. Once a case guides the revision, it becomes development evidence. Repeatedly calling it a fresh test can exaggerate the result.
Maintain a development set and a protected evaluation set. Limit access to the protected cases during optimization. When a protected failure is examined and used to fix the system, retain it as a regression case and refresh the independent evaluation boundary where practical.
Holdouts do not need secrecy for its own sake. Their purpose is to test performance beyond the examples used to shape the implementation. A team that cannot maintain a meaningful holdout should report that limitation rather than claim independent generalization from repeated familiar cases.
Separate case identity from its displayed wording. Slight paraphrases of the same scenario can still share the same underlying information. Splitting near-duplicates between development and holdout can make the latter easier without adding a genuinely different condition.
Version inputs, configuration and outputs together
An evaluation result needs the model or implementation version, prompt, retrieval configuration, policy version and relevant tool behavior. “The model scored well” is not reproducible when the task inputs and settings are unknown.
Retain the output for each case and the grader result. Aggregate metrics should link back to failures. If the assistant produced an invented policy exception, the team should be able to inspect the input, retrieved passages and draft that led to it.
Retrieval deserves its own evidence. A poor answer can arise because the correct policy was absent from context, because the model misread it or because the final message omitted a condition. Separating those causes lets the team choose the right repair instead of rewriting the prompt for every failure.
Configuration changes can invalidate comparisons. A candidate with a better retrieval source may outperform another even if the model itself is not better. Report the comparison at the system level unless the experiment actually isolated the model difference. Association with a model name is not proof of the cause.
Use model graders as instruments that need calibration
A model grader can scale inspection, but it has its own failure modes. OpenAI’s guidance identifies position and verbosity biases and recommends agreement checks against human labels. A judge that prefers longer answers can reward a draft that says more while introducing more unsupported claims.
Calibrate on cases reviewed by people with relevant knowledge. Compare both agreements and consequential disagreements. If the judge misses invented policy exceptions, improving its average agreement on friendly greetings is insufficient.
Blind candidate identity where practical and vary presentation order for pairwise comparisons. Give the grader the relevant evidence and a specific question. “Which answer is better?” can invite stylistic preference; “Does either answer promise a refund unsupported by these policy passages?” directs inspection toward the actual risk.
A different model name does not establish independence. The author and judge may share assumptions or overlook the same missing condition. Independent AI verification examines that correlated-failure problem. Exact checks and accountable human review remain useful complements.
Repeat runs when variability matters
One output per case can conceal a system that succeeds most of the time but occasionally invents a consequential claim. Repeated runs can expose variation, provided the budget and interpretation are explicit.
Choose repeat counts according to the decision and failure consequences. A low-risk phrasing comparison may need less repetition than an answer that can mislead a customer about money. There is no universal count that proves reliability for all AI features.
Report the distribution of outcomes rather than only the best sample. If a candidate generates a correct draft in four attempts and a critical error in one, the critical outcome should remain visible. A product cannot select the best output after the fact unless its deployed workflow has a valid way to identify it.
Repeated testing consumes resources. Record total evaluation cost and the basis for stopping. A bounded suite can provide useful evidence without pretending that unlimited repetitions would eliminate uncertainty. The results support the tested case distribution and configuration, not every future customer message.
Aggregate by severity and case family
An overall pass rate can hide a concentrated failure. Suppose a hypothetical suite contains ninety ordinary cases and ten missing-identity cases. A system passes every ordinary case and discloses order details in all ten identity cases. A ninety-percent result would conceal a severe product defect.
Report critical failures separately and break results down by meaningful case family. The denominator should be clear. A score based on individual output attempts differs from one based on cases that passed every repeat.
Use quality metrics that match the intended decision. Policy fidelity, evidence support, clarification quality and useful next steps can be more informative than a generic fluency score. The rubric should explain why each metric matters to the customer obligation.
Numerical examples here are hypothetical. A real release threshold should be agreed before testing and supported by the product’s consequences, review workflow and tolerance for residual error. A vendor’s illustrative percentage is not automatically the right threshold for a merchant app.
Connect evaluations to a release decision
The suite should end in a decision: accept the change within scope, revise it, narrow the feature or keep it behind human review. A report that lists scores without stating the implications leaves the operator to infer policy from a dashboard.
Compare the candidate with a retained baseline using the same conditions. Identify improvements, regressions and unresolved tradeoffs. A candidate may improve clarity while worsening policy fidelity. That is a reason to inspect the failures, not blend them into one attractive average.
The decision should identify the artifact and evidence it covers. A prompt edited after evaluation needs appropriate rechecking. A new tool or policy source can alter the feature’s behavior even when the model remains unchanged.
Model routing can use this evidence to compare cost and quality for a defined task. It should include retries and review effort rather than selecting a cheaper model from nominal call prices alone.
Learn from production without polluting the evidence
Production observations reveal cases the initial suite missed. A customer correction, support escalation or unsupported claim can become a retained regression case after appropriate sanitization and authorization.
Do not assume positive feedback establishes factual correctness. A customer can like an answer that invents an exception. Likewise, a negative rating can reflect an unwelcome but accurate policy. Feedback is evidence to investigate, not a complete grader.
Sample across actual traffic and important failure classes. Keep the monitoring denominator and collection method visible. A handful of voluntarily reported complaints does not estimate the overall failure rate without understanding who reports and who remains silent.
The customer outcome feedback loop connects these observations to product improvement. Evaluation preserves the distinction between learning from a case and independently testing the next change. Both are valuable when they are labeled honestly.
Keep evaluation infrastructure replaceable
Store case definitions, reference evidence and rubric versions in a portable form under the project’s existing conventions. A vendor dashboard can help run the suite, but it should not be the only home for the product’s acceptance knowledge.
Service APIs, availability and pricing can change. A suite whose cases survive those changes is easier to maintain. The application can replace a runner while preserving the obligations and historical results needed for comparison.
Avoid building a elaborate platform before the first useful suite exists. A versioned case file, repeatable runner and inspectable report can support an initial pilot. Expand the infrastructure when case volume, reviewer workflow or operational requirements justify it.
Investigate the first consequential disagreement
When a judge and a human reviewer disagree, preserve the case before changing either rubric or prompt. Ask what evidence each used and which claim produced the disagreement. A judge may have rewarded a polite phrase while the reviewer noticed an unsupported return exception. The repair may be a clearer policy-fidelity criterion rather than a stronger model.
Re-evaluate the repaired rubric on cases that were not used to write it. Otherwise a grader can appear improved because it learned the exact disagreement the team just explained. Retain the old and new judgments so the comparison remains inspectable.
Some disagreements reveal a genuine product ambiguity. If two responsible policy owners interpret a deadline differently, a model judge cannot settle the company’s obligation legitimately. Resolve the policy, update its authoritative source and then evaluate the assistant against the accepted rule.
This process can make the suite smaller and more useful. Removing an unanswerable case with a documented reason is better than forcing a numerical label that hides the missing decision. Keep the excluded case visible in the report so the release owner understands which condition still lacks a supported answer.
What Would We Do at Salars?
A proposed Forge pilot would evaluate one narrow explanation feature against safe merchant records. Deterministic margin arithmetic would remain outside the model. The explanation would need to preserve source rows, identify assumptions and avoid presenting hypothetical values as observed financial results.
The suite would include ordinary quotes, missing currency, contradictory cost assumptions and attempts to override the diagnostic boundary. An independent reviewer would define the consequential expectations. Protected cases would remain separate from prompt development, with failures retained rather than deleted to improve the score.
The report would state the tested configuration, output variability, severity-specific results, total cost and remaining review needs. It would not claim that a passing suite proves the app profitable or safe for automatic merchant actions. Those are different decisions with different evidence.
A useful evaluation makes the next product choice clearer. Its strongest result is a specific supported conclusion about behavior under known conditions, accompanied by the failures that constrain that conclusion.
Sources
- OpenAI: evaluation best practices, read October 7, 2026, for task-specific objectives, case variety, human calibration and model-grader biases. The article uses the methodology and does not depend on the hosted Evals platform.
- Salars: test an AI workflow before it acts, for a small fictional workflow case set.
- Salars AI library, for related verification guides. All suite designs, percentages and Forge pilots in this chapter are hypothetical or proposed.
Loading comments…