AI agents · Tool evaluation

Evaluate MCP tool-using agents: score the whole trajectory

A production evaluation contract for AI agents that discovers, selects, approves and calls MCP tools without hiding unsafe or duplicated effects behind a good answer.

TOPIC HUBAI, RAG & Vector Search
Conceptual illustration of an AI agent trajectory passing through tool-contract, approval, execution and evaluation gates before entering a feedback dataset.
An editorial interpretation of the topic, followed by a practical execution diagram.

An AI agent with MCP tools should be evaluated as a trajectory, not only by its final sentence. A production test must verify which tool the agent selected, whether its arguments matched the schema and tenant scope, whether a sensitive action requested approval, what the tool actually returned, whether retries duplicated an effect, and whether the final answer faithfully reflected that result. This layered approach is useful for assistants that read business data or take actions; it is unnecessary for a deterministic workflow that ordinary application code can execute more safely.

This is a practical engineering guide based on the Model Context Protocol specification and official OpenAI and Anthropic documentation, verified on 5 October 2026. It is not a report of a new model release. The scoring contract, dataset structure and release gates below are application-design recommendations built on those documented protocol and evaluation behaviors.

Treat one run as a typed trajectory

The MCP tools specification defines a discoverable tool by name, description and `inputSchema`, with an optional `outputSchema`. The model may decide which tool to invoke, but the host remains responsible for access control, validation, timeouts, logging and confirmation of sensitive operations. That means “the answer looked right” observes only the last node of a much larger system.

Represent a run as ordered events: user request, policy scope, available tool set, model decision, tool name, arguments, approval decision, server result, retry or handoff, and final response. Attach stable IDs for run, tenant, user, tool version and operation. Do not store secrets or unrestricted personal data in the trace; retain redacted inputs or hashes where the evaluator needs proof without the original value.

OpenAI's agent-evaluation guidance describes traces as end-to-end records of model calls, tool calls, guardrails and handoffs. Trace grading is useful while debugging workflow-level behavior, then repeatable datasets and eval runs make changes comparable. The important product decision is not which dashboard holds the trace. It is that every release can reconstruct why a tool was exposed, selected, approved and accepted.

Build scenarios around decisions, not prompts

A prompt list is not yet an agent eval. Each case needs initial state, authenticated scope, permitted tools, expected decision points, acceptable alternative paths, forbidden effects and a final-state assertion. For an order assistant, “cancel order 481” means different things when the order belongs to another tenant, has already shipped, requires approval or has already been cancelled by a previous attempt.

Create five scenario classes. Happy paths confirm the shortest safe route. Boundary cases cover missing fields, ambiguous identities and pagination. Authorization cases ask for another tenant's resource or a tool outside the caller's role. Adversarial cases put instructions inside retrieved documents or tool results. Recovery cases inject timeouts, partial failures, duplicated callbacks and resumed sessions.

Use real production distributions after removing sensitive data, then add curated rare cases that production logs underrepresent. Version the fixture state and expected outcomes. A case whose backing database changes between runs cannot distinguish model regression from test drift. OpenAI's evaluation best practices recommend task-specific evals, early and repeated testing, production-derived cases, continuous evaluation and human calibration of automated scores.

Evaluate tool discovery and selection separately

First ask whether the correct capability was visible. A perfect model cannot select a tool omitted from discovery, and a secure policy should intentionally hide tools the caller cannot use. MCP supports `tools/list`; clients should account for pagination and for tool-list change notifications when the server advertises that capability. Cache discovery only with a clear invalidation rule.

Then score selection. For each case define required tools, allowed alternatives and forbidden tools. “Read order, then summarize” may permit `orders.get` and no mutation. “Cancel order after approval” may require a read before `orders.cancel`. Do not grade only exact sequence when two routes are operationally equivalent; grade invariants such as “no mutation before ownership check.”

Expose the smallest relevant tool surface. OpenAI's MCP guidance notes that `allowed_tools` can reduce cost and latency when a server exposes many tools. It is also a useful policy boundary, although it does not replace authorization inside the server. Measure unavailable-tool attempts, unnecessary calls and tool-selection accuracy by intent class.

Evaluate each layer independently: discovery, selection, arguments, policy, approval, execution, authoritative state and grounded response; feed failures back into the dataset.
Evaluate each layer independently: discovery, selection, arguments, policy, approval, execution, authoritative state and grounded response; feed failures back into the dataset. Open for a larger view

Test descriptions and schemas as product interfaces

Tool names and descriptions are instructions consumed by a probabilistic planner. Anthropic's tool-use documentation emphasizes detailed descriptions: what the tool does, when to use it, when not to use it, parameter meaning and important limitations. Treat a description change like a routing change and run the selection suite before publishing it.

Validate arguments twice. The client or server should reject values that do not match `inputSchema`; business code must then enforce facts JSON Schema cannot express, such as tenant ownership, resource state, spending limit or allowed transition. Never let the model supply the authoritative tenant ID when the authenticated session already determines it.

Use an `outputSchema` for stable structured results. The MCP specification says servers must conform when one is declared, and clients should validate structured results. Test missing fields, unexpected enum values, oversized payloads and a server returning text where the client expects structured content. Tool annotations are hints, not trusted policy, unless the server itself is trusted.

type ToolEvalCase = {
  identity: { tenantId: string; roles: string[] };
  request: string;
  allowedTools: string[];
  expected: {
    requiredCalls?: string[];
    forbiddenCalls?: string[];
    approvalBefore?: string[];
    finalState: Record<string, unknown>;
  };
  faults?: Array<'timeout' | 'duplicate-result' | 'stale-read'>;
};

const result = await runAgentWithRecordedTools(testCase);
assertSchemaValid(result.toolCalls);
assertNoForbiddenCalls(result.trace, testCase.expected);
assertApprovalOrder(result.trace, testCase.expected);
assertAuthoritativeState(result.state, testCase.expected.finalState);

This code illustrates the shape of a harness; it is not a vendor SDK. The decisive assertion is authoritative final state, not a substring in the model's answer.

Grade arguments, not only tool names

An agent can choose the correct tool and still cause the wrong result. Score required arguments, normalized values, omitted defaults and policy-derived fields. For search, verify filters, limit and tenant scope. For a payment, verify currency, amount, recipient and idempotency key. For a destructive action, verify the exact resource and expected version.

Prefer deterministic checks for structured fields. A JSON Schema validator, ownership assertion or database query is more reliable than asking another model whether the value “looks sensible.” Use an LLM grader when correctness depends on semantic equivalence, such as whether a justification accurately reflects a tool result. Calibrate that grader against human labels and record disagreement.

Separate valid alternatives from errors. A date may be normalized to UTC or preserved with an offset; both can be correct if the contract allows them. Encode that in the evaluator instead of forcing one serialized string. Conversely, do not accept a different resource merely because the final prose sounds plausible.

Make approval a testable state transition

OpenAI's MCP integration requests approval by default before data is shared with a remote server, and exposes approval request and response items. The MCP specification also recommends human confirmation for sensitive operations and clear visibility into invoked tools. In your application, define which operations are read-only, externally visible, financially meaningful or destructive, then map them to approval policy.

Test the order of events: the user sees the tool name, meaningful arguments and consequence before the call; a denial causes no side effect; an approved call uses the reviewed arguments; a later retry does not silently bypass approval with changed parameters. Bind the approval to a digest of tool, arguments, tenant and expiry. If any material field changes, request a new decision.

Do not use approval prompts for every harmless read. Repeated low-value confirmations train users to accept without reading. The eval should score both missing approval and unnecessary approval, because either can damage the product.

Verify the effect outside the model trace

A tool result saying “success” is a claim, not proof. After a mutation, query the authoritative system or inspect its durable event. Confirm that exactly one intended effect occurred and that no forbidden field changed. This catches a server that returned success before committing, a timeout after commit and a retry that created a duplicate.

Give every retryable mutation an application-generated operation ID. The server should store that ID with the effect and return the prior result when it is replayed. In the harness, inject a timeout after commit, resume the agent and confirm the state remains singular. Also test a tool execution error, a protocol error, a rate limit and a malformed result; MCP distinguishes protocol failures from execution failures, and the agent should not handle them as one generic message.

Use record-and-replay carefully. Recorded tools make model comparisons fast and reproducible, but they do not prove the current server contract or authorization policy. Keep a deterministic replay suite for planning and a smaller live sandbox suite for end-to-end effects. Never point automated destructive evals at production accounts.

Score the final answer against evidence

After the trajectory is safe, evaluate communication. The final answer should state what actually happened, distinguish an attempted action from a completed one, surface approval or recovery requirements, and avoid inventing fields absent from tool output. For a failed call, “the order was cancelled” is wrong even if the chosen tool and arguments were correct.

Attach claims to trace evidence: call result, authoritative read or policy decision. Grade factual support, completeness, uncertainty and whether sensitive values were unnecessarily repeated. An agent that completes the task but leaks another customer's identifier fails the run.

Keep product outcome separate from workflow correctness. The workflow may be correct while the user goal remains unresolved because stock is unavailable or policy forbids the action. Report both: safe execution rate and task-resolution rate. Combining them into one score makes a strict authorization check look like a quality regression.

Create a layered scorecard and release gate

Use hard gates for safety and deterministic contract violations: cross-tenant access, forbidden tool invocation, missing required approval, duplicate financial effect, invalid output schema and unsupported claim of success. One occurrence can block a release. Use rates for softer behavior: correct tool selection, unnecessary calls, argument accuracy, recovery success, final-answer support, latency and cost.

Slice results by intent, tool, tenant tier, language, model, prompt version and failure mode. A high aggregate can hide that Arabic requests select the wrong shipping tool or that one expensive MCP server dominates latency. Record confidence intervals where the sample is small, and keep a fixed holdout set so repeated tuning does not turn the benchmark into training data.

Run deterministic contract tests on every tool or policy change, then a bounded agent suite on prompt, model or tool-surface changes. Add sampled production traces only after redaction and consent controls are satisfied. OpenAI recommends continuous evaluation and growing the dataset as nondeterministic cases appear; the useful implementation is a release gate with owners and rollback thresholds, not a dashboard no one reviews.

Know when not to use an agent or MCP

Use an MCP-connected agent when the task is genuinely variable: the user expresses goals in natural language, several tools may be relevant, and the system benefits from model planning while application controls can contain risk. Use ordinary code when the flow is known: validate input, call API A, then API B, and return a defined result. A deterministic state machine is easier to test, operate and secure.

MCP standardizes discovery and invocation; it does not grant trust to a server, make an unsafe API safe, or guarantee the model will choose correctly. Do not expose broad administrative tools merely because the protocol can describe them. Place authorization and invariants in the tool implementation, keep the available surface narrow, and require approval where consequences justify it.

The production contract is layered: the right tool is visible, the agent selects it for the right reason, arguments satisfy schema and policy, approval precedes sensitive effects, the server produces a valid result, retries are idempotent, authoritative state matches the goal, and the final answer is grounded in the trace. Evaluate each layer independently, then decide whether the whole trajectory is safe to release. For related patterns, see the portable AI evaluation contract, durable AI agent execution and AI integration services.

Official references

These references document the tools discussed. Examples and design decisions are illustrative and should be adapted to the project and its versions.

Prepared by: Noor Yasser

FROM DECISION TO DELIVERY

Working through a similar engineering challenge?

I help teams turn architecture decisions into a clear scope and dependable, reviewable implementation.

Book a 30-minute callRelated serviceAI integrations & retrieval systemsRelevant projectAI Action Studio