AI · RAG

How to evaluate a RAG pipeline before changing the model

Separate retrieval failures from answer failures with a small evaluation set, source checks and realistic commerce examples.

TOPIC HUBAI, RAG & Vector Search
Selected knowledge cards gathered into one answer, illustrating retrieval.
An editorial interpretation of the topic, followed by a practical execution diagram.

When a store assistant gives a wrong answer, switching the model may fix the wrong layer. The relevant product might be absent from the index, the policy might be stale, or the answer might require a live API call. Separate ingestion, retrieval and answer generation so that each failure has an observable cause.

Build a representative question set

Collect anonymized questions covering exact product names, vague descriptions, spelling errors, mixed languages, comparisons and unanswerable requests. Record the expected source and acceptable behavior. For a request such as “waterproof shoes in size 42,” product properties and current stock are separate requirements. A polished demo containing only easy questions does not measure this distinction.

Chunk around meaningful boundaries

Keep a product description or policy clause coherent. Store its identifier, title, language, tenant and update time as metadata. Fixed character cuts can separate an exception from its rule, while very large chunks dilute relevance. Compare alternative boundaries on the same evaluation set instead of treating one chunk size as universal.

Inspect retrieval before judging the answer

Record returned documents and rank. Check whether the required source appears within the first k results and whether other results introduce irrelevant or cross-tenant content. Recall@k measures the fraction of required reference sources recovered in those results. With one reference source per question, the share of questions retrieving that source is a useful starting signal. Human inspection still matters when reference labels are incomplete.

Evaluate retrieved evidence and the generated answer separately.
Evaluate retrieved evidence and the generated answer separately. Open for a larger view

Separate documents from live facts

A maintained shipping policy is suitable retrieval material. Inventory and order status usually need an authorized live lookup. Treat retrieved text as untrusted data: it must not redefine tool permissions. Verify identity and order ownership in application code before an order tool runs.

Question -> tenant-scoped retrieval -> source review
         -> authorized live lookup, when needed
         -> grounded answer -> evaluation record

Measure quality, latency and cost separately

Review claim accuracy, citation support and appropriate refusal. A linked source is not proof that it supports the sentence. Track retrieval and generation latency and per-request cost. Change one variable at a time and keep a held-out set that was not used while tuning. This makes regressions visible instead of hiding them behind an attractive average score.

Change the model when the evidence points there

If correct context is consistently present but the model still cannot follow instructions or compose an accurate response, compare models. If the source is missing, improve ingestion and retrieval first. RAG cannot guarantee zero hallucinations; design measurable limits and a human escalation path for insufficient evidence.

Scenario: recommending an unavailable product

Retrieval can correctly find a shoe matching the question while stock has changed since indexing. Classify this as a freshness or tool-use issue rather than automatically blaming semantic search. Record the retrieved document, the authorized API result and the exact context passed to generation. That separation reveals whether the assistant ignored correct evidence or never received it.

Testing knowledge changes

After a shipping-policy update, rerun the same question before and after reindexing and verify that obsolete evidence is retired. Test product deletion and knowledge-source removal too. Removing information should change answer behavior rather than leaving forgotten copies in a cache or index. Keep a consistent comparison sheet and record why a change was accepted or rejected.

Official references

These references document the tools discussed. Examples and design decisions are illustrative and should be adapted to the project and its versions.

Prepared by: Noor Yasser

FROM DECISION TO DELIVERY

Working through a similar engineering challenge?

I help teams turn architecture decisions into a clear scope and dependable, reviewable implementation.

Book a 30-minute callRelated serviceAI integrations & retrieval systemsRelevant projectMurshid