When a store assistant gives a wrong answer, switching the model may fix the wrong layer. The relevant product might be absent from the index, the policy might be stale, or the answer might require a live API call. Separate ingestion, retrieval and answer generation so that each failure has an observable cause.
Build a representative question set
Collect anonymized questions covering exact product names, vague descriptions, spelling errors, mixed languages, comparisons and unanswerable requests. Record the expected source and acceptable behavior. For a request such as “waterproof shoes in size 42,” product properties and current stock are separate requirements. A polished demo containing only easy questions does not measure this distinction.
Chunk around meaningful boundaries
Keep a product description or policy clause coherent. Store its identifier, title, language, tenant and update time as metadata. Fixed character cuts can separate an exception from its rule, while very large chunks dilute relevance. Compare alternative boundaries on the same evaluation set instead of treating one chunk size as universal.
Inspect retrieval before judging the answer
Record returned documents and rank. Check whether the required source appears within the first k results and whether other results introduce irrelevant or cross-tenant content. Recall@k measures the fraction of required reference sources recovered in those results. With one reference source per question, the share of questions retrieving that source is a useful starting signal. Human inspection still matters when reference labels are incomplete.
Separate documents from live facts
A maintained shipping policy is suitable retrieval material. Inventory and order status usually need an authorized live lookup. Treat retrieved text as untrusted data: it must not redefine tool permissions. Verify identity and order ownership in application code before an order tool runs.
Question -> tenant-scoped retrieval -> source review
-> authorized live lookup, when needed
-> grounded answer -> evaluation recordMeasure quality, latency and cost separately
Review claim accuracy, citation support and appropriate refusal. A linked source is not proof that it supports the sentence. Track retrieval and generation latency and per-request cost. Change one variable at a time and keep a held-out set that was not used while tuning. This makes regressions visible instead of hiding them behind an attractive average score.
Change the model when the evidence points there
If correct context is consistently present but the model still cannot follow instructions or compose an accurate response, compare models. If the source is missing, improve ingestion and retrieval first. RAG cannot guarantee zero hallucinations; design measurable limits and a human escalation path for insufficient evidence.
Scenario: recommending an unavailable product
Retrieval can correctly find a shoe matching the question while stock has changed since indexing. Classify this as a freshness or tool-use issue rather than automatically blaming semantic search. Record the retrieved document, the authorized API result and the exact context passed to generation. That separation reveals whether the assistant ignored correct evidence or never received it.
Testing knowledge changes
After a shipping-policy update, rerun the same question before and after reindexing and verify that obsolete evidence is retired. Test product deletion and knowledge-source removal too. Removing information should change answer behavior rather than leaving forgotten copies in a cache or index. Keep a consistent comparison sheet and record why a change was accepted or rejected.
Official references
These references document the tools discussed. Examples and design decisions are illustrative and should be adapted to the project and its versions.
Prepared by: Noor Yasser
Working through a similar engineering challenge?
I help teams turn architecture decisions into a clear scope and dependable, reviewable implementation.




