An AI evaluation can look mature while still being useless for a product decision. A dashboard may show that a candidate model scored 0.86 instead of 0.82, yet the team still cannot answer whether checkout support tickets resolve faster, whether a merchant accepts the suggested action, whether latency remains tolerable, or whether the change can be reversed safely.
The engineering problem is not merely how to run a grader. It is how to preserve a chain of evidence from a versioned user task to an offline result, a controlled release and an observed product outcome. That chain should survive a model swap, a prompt rewrite, a retrieval change and even the retirement of the evaluation service itself.
This matters now because OpenAI's official deprecation schedule says the Evals platform was deprecated on 3 June 2026, existing evals become read-only on 31 October 2026, and the dashboard and API are scheduled to shut down on 30 November 2026. The lesson is not “stop evaluating.” It is to make the evaluation contract portable and treat a vendor dashboard as a runner, not the system of record.
Start with the product decision
Write the decision before selecting a metric. A useful decision might be: “Promote the new support assistant to 20% of eligible conversations if refund-policy classification does not regress, unsafe tool calls remain zero in the release set, p95 response latency stays below the agreed limit, and cost per resolved conversation improves.”
That statement has a candidate, population, quality gates, operational guardrails and a promotion action. “Raise the average LLM score” has none of them. OpenAI's evaluation guidance recommends task-specific evals, representative data, continuous evaluation and human calibration instead of generic metrics or “it feels better.” Anthropic's evaluation guide similarly starts with specific, measurable and relevant success criteria and notes that most use cases need several dimensions such as task fidelity, consistency, latency, price and privacy.
Separate three questions: Can the system perform the task on a controlled dataset? Can it operate inside latency, cost and safety limits? Does it improve the user outcome after release? One score should not pretend to answer all three.
Define a portable evaluation contract
Keep the source of truth in versioned files or records your team controls. The contract should name the dataset version, task schema, candidate configuration, grader definitions, aggregation rules, thresholds, guardrails and decision rule. Candidate identity is more than a model name: include the model snapshot or alias, prompt hash, retrieval configuration, tool schema version, safety policy and relevant runtime parameters.
An illustrative contract might look like this:
{
"task": "refund-policy-resolution",
"dataset": "support-cases-2026-09-v3",
"candidate": {
"model": "provider/model-snapshot",
"prompt": "sha256:…",
"retriever": "kb-42",
"tools": "support-tools-v7"
},
"graders": ["route_exact_match-v2", "citation_rubric-v4"],
"gates": {
"critical_tool_errors": 0,
"route_accuracy_min": 0.94,
"p95_latency_ms_max": 2500
},
"decision": "shadow_then_canary"
}This is a structural example, not a universal threshold. Store the raw per-case result and the aggregate, because an average can hide one severe failure class. Give every run an immutable ID that links the contract, code revision, dataset, outputs, grader results and environment.
Build the dataset from real task distribution
A portable dataset is a set of task records with stable IDs, inputs, expected properties and slice labels. Include common traffic, high-value journeys, rare failures, ambiguous inputs and adversarial cases. OpenAI recommends mixing domain, historical, production, synthetic and human-curated data, then expanding continuously from new failures. Anthropic emphasizes mirroring the real task distribution and including edge cases.
Do not copy production logs into an eval store without a data contract. Remove or tokenize personal data, define retention, record consent and access boundaries, and keep sensitive payloads out of metric labels. Store enough context to reproduce the task, but avoid preserving an entire conversation when a minimal excerpt and structured facts are sufficient.
Use fixed release sets for comparability and a rotating discovery set for new failure modes. If every failed production case immediately enters the release set, teams can overfit the prompt to recent incidents. Keep a held-out set that candidate authors cannot inspect, and periodically retire cases that no longer represent the product.
Grade with a ladder, not one judge
Use the cheapest reliable grader for each property. Start with code: schema validity, exact labels, required citations, tool arguments, arithmetic, executable code or database assertions. Move to pairwise or rubric-based model grading when the property is semantic. Use domain experts for calibration, severe-risk cases and disagreements.
Both OpenAI and Anthropic recommend clear rubrics and favor constrained judgments such as pass/fail, classification or pairwise comparison over vague open-ended scoring. Anthropic's guide orders code-based grading as the fastest and most reliable, human grading as flexible but expensive, and LLM grading as scalable only after its reliability is tested. OpenAI warns about judge biases and recommends maintaining agreement with human feedback.
Version the grader like product code. A candidate score cannot be compared with a previous run if the rubric changed silently. Record grader model, prompt and normalization logic. Run a calibration set whenever a judge changes, and report agreement by important slices rather than only as one global percentage.
Connect eval results to the product funnel
Offline correctness is only the first layer. For an AI suggestion feature, instrument a funnel such as eligible → suggestion generated → suggestion shown → user reviewed → accepted → action executed → outcome retained. Rejection, editing, undo and escalation are product signals, not mere model failures.
Measure cost per successful task, not only cost per request. OpenAI's deployment checklist explicitly recommends comparing task success, latency, token use and cost per successful task when choosing a model. A cheaper model that triggers more retries or human correction may cost more per resolved outcome.
Define outcome windows. Acceptance in five seconds does not prove the suggestion was useful if the user reverses it an hour later. A support answer may be “positive” in a thumbs-up event yet fail to resolve the case. Join immediate interaction events with the downstream state change that represents value, while respecting privacy and data minimization.
Create one trace across evaluation and production
Use a stable evaluation run ID, case ID and candidate version across logs, traces and result records. In production, create a decision ID that links model execution to the product action without putting raw user text into attributes. Record latency, token categories, retrieval result IDs, tool-call outcome, grader or policy decisions and final product status as controlled fields.
OpenTelemetry Semantic Conventions provide standardized names across traces, metrics, logs, profiles and resources. Use the relevant standard attributes where they fit, then add low-cardinality product fields such as feature version, experiment cohort and outcome class. Do not use customer IDs, prompts or free-form model output as metric labels.
Keep observability evidence separate from eval truth. A trace explains what happened during one execution. The eval store explains whether that execution satisfied a versioned criterion. Joining them is useful; collapsing them into the same unstructured log is not.
Release through shadow and canary stages
Run the candidate offline first. If it passes, use shadow traffic where the candidate receives representative inputs but cannot affect the user or call mutating tools. Compare its decisions with the active version, inspect disagreements and verify real latency and cost.
Then expose a small eligible cohort with a reversible routing rule. Keep severe-risk constraints outside randomization: a candidate that can issue refunds, publish content or change customer data needs explicit tool permissions and approval boundaries. Promotion should read the contract's gates automatically, but a passing gate is evidence for a decision, not an obligation to ship.
Rollback must restore the complete previous candidate configuration: model, prompt, retrieval index, tools and policies. Rolling back only the model while leaving a new tool schema or prompt active creates a configuration that was never evaluated. Keep the previous bundle deployable until the outcome window closes.
Migrate the system of record before the runner disappears
For teams using OpenAI Evals, the calendar is concrete: inventory definitions and results now; stop creating new dependencies on dashboard-only state; export datasets, rubrics and run evidence before the 31 October read-only date; and validate the replacement runner before the 30 November shutdown. OpenAI links a Promptfoo migration path, but the durable decision is broader than choosing that specific tool.
Build an adapter around the runner. It receives your contract and cases, invokes the chosen model or workflow, maps outputs into your result schema and returns evidence. A local harness, CI job or another hosted service can implement the same boundary. Keep thresholds and release decisions in your repository or product-control service rather than reproducing them manually in each dashboard.
Do not migrate only aggregates. Preserve case-level outputs, grader versions, annotations and run metadata so regressions remain explainable. Verify parity by running the old and new runners on the same frozen candidate and dataset, then investigate disagreements before switching the release gate.
Hypothetical example: merchant action suggestions
Imagine an assistant that suggests a response and an optional action for a merchant's support ticket. This example is hypothetical. The offline set contains routine questions, Arabic and English variants, missing order identifiers, conflicting policy text, refund requests and malicious attempts to trigger an unauthorized action.
Code graders validate the response schema, cited source IDs and tool arguments. A rubric grader checks whether the response uses the supplied policy and asks for missing information instead of inventing it. Critical cases require a human-reviewed expected action. The release gate prohibits unauthorized mutating calls and requires minimum routing accuracy, while latency and cost remain guardrails.
In shadow mode, the candidate cannot execute anything. In canary mode, low-risk response suggestions appear to a small merchant cohort, but refunds still require explicit approval. Product telemetry distinguishes shown, accepted unchanged, edited, rejected, executed, failed and undone. The team promotes only when the candidate clears offline gates and improves a downstream outcome such as time to a correctly resolved ticket without increasing reversals or support escalations.
Ship the evidence, not the score
A production-ready AI change should carry an evidence bundle: decision statement, contract version, candidate bundle, dataset and slice report, grader calibration, latency and cost distribution, severe-case review, shadow comparison, canary outcomes and rollback reference.
The practical Product Engineering principle is this: an eval is not a dashboard number. It is a repeatable argument for changing user exposure. Own the contract, keep runners replaceable, connect technical quality to product outcomes, and preserve enough evidence to explain both promotion and rollback.
Official references
These references document the tools discussed. Examples and design decisions are illustrative and should be adapted to the project and its versions.
- OpenAI — API deprecations: Evals platform timeline, verified 30 September 2026
- OpenAI — Evaluation best practices, verified 30 September 2026
- OpenAI — API deployment checklist, verified 30 September 2026
- Anthropic — Define success criteria and build evaluations, verified 30 September 2026
- OpenTelemetry — Semantic conventions, verified 30 September 2026
Prepared by: Noor Yasser
Working through a similar engineering challenge?
I help teams turn architecture decisions into a clear scope and dependable, reviewable implementation.




