AI Engineering4 min read

Production RAG: retrieval, permissions and cost control

Build a retrieval system with document permissions, grounded answers, useful evaluations and a cost model you can measure.

Documents passing through permission checks and retrieval to a cited answer

The short answer

A reliable RAG system needs permission-aware retrieval, source citations, an explicit no-answer path and evaluation on representative questions. Hybrid search and caching are options to test, not guarantees of accuracy.

Define what a correct answer means

Retrieval-augmented generation (RAG) supplies a model with relevant source material before it answers. It can help with changing or private knowledge, but retrieval does not guarantee the retrieved material is correct, complete or authorized for the user.

Start with a narrow use case such as answering questions about a versioned product manual. Write examples of acceptable answers, the citations they need, and when the system should decline to answer or escalate. Keep answer generation separate from tools that change business records.

Prepare documents for retrieval

Preserve headings, table context, version numbers and source URLs. Give every chunk a stable document identifier and retain its access rules and revision date. Split on useful semantic boundaries rather than assuming one token count works for every document. Test questions whose answers span two sections or depend on a footnote.

Document deletion and permission changes must propagate to search indexes and caches. Re-embedding a document should not leave old, accessible chunks behind. Keep a record of ingestion failures so missing evidence is visible to operators.

Compare retrieval approaches

Exact terms such as error codes, SKUs and version numbers often benefit from keyword search. Embeddings can help with conceptual similarity. A hybrid system can combine both result sets and use rank fusion or reranking, but it adds cost and latency.

Build a labelled question set before adding stages. Measure whether relevant authorized passages appear in the retrieved set, then whether the final answer is supported by those passages. Include ambiguous, unanswerable and outdated questions. Precision at a chosen cutoff is a retrieval metric, not a measure of overall factual accuracy.

Enforce access before exposing context

  1. Authenticate the user and establish tenant and document permissions in trusted application code.
  2. Apply authorization constraints to retrieval and validate the final selected passages before they reach the model.
  3. Treat retrieved text as untrusted data. A document instruction to reveal secrets or call a tool must not override application policy.
  4. Restrict tools to necessary operations; validate arguments, permissions and business rules independently of model output.
  5. Require an appropriate confirmation or review step for consequential actions and retain a minimal audit trail.

Control costs without sharing the wrong answer

Track retrieval, embedding, reranking and generation costs separately, along with p50 and p95 latency. Limit context to passages that help answer the question, cap output length, and choose model capability against your evaluation set.

Start caching with exact matches and a clear invalidation strategy. Scope keys to tenant, access-policy version, document version, model and prompt version. Similar wording is not enough to establish that two users are entitled to the same answer. A semantic cache needs additional testing for false matches and stale answers; it still incurs retrieval or embedding work.

Release with a measurable acceptance gate

Set thresholds appropriate to the risk of the use case. Test citations, refusals, permission boundaries, malicious documents, timeouts and tool failures. Structured JSON can make parsing reliable, but a valid schema does not make a claim true or an action authorized.

Review failed answers and retrieval misses after release. Add them to a held-out regression set rather than tuning only to a demonstration. Show the user the supporting source and make uncertainty visible when the available material cannot resolve the question.

Frequently asked questions

Does RAG stop an AI system from inventing answers?

No. Retrieval can supply useful evidence, but the passages may be incomplete, outdated or irrelevant, and the model may misinterpret them. Evaluate whether the final answer is supported by the selected sources and define when the system should decline or escalate.

How do I keep one customer from seeing another customer’s documents?

Enforce tenant and document permissions in trusted application code before evidence reaches the model. Recheck the selected passages and scope caches to the relevant access policy and content versions. A prompt asking the model to respect permissions is not an authorization boundary.

Should I use keyword search, vector search or both?

Test representative questions first. Exact identifiers and error codes often need keyword matching, while embeddings can help with conceptual similarity. Compare retrieval quality, answer support, latency and cost before adding hybrid retrieval or reranking.

When is caching an AI answer safe?

Only when the cache key and invalidation rules preserve the answer’s permissions and freshness requirements. Include the tenant, policy, document, model and prompt versions where relevant. Similar wording alone does not prove that two users can receive the same answer.

Sources and further reading

Keep exploring

Explore AI application development or the queue and idempotency guide for reliable background workflows.

Paul Edward

Written by Paul Edward

Senior full-stack web developer working with PHP, Laravel, WordPress and AI-assisted web systems.

More about Paul

Leave a Reply

Your email address will not be published. Required fields are marked *

Loading a quick check… (this needs JavaScript)

Project brief Step 1 of 2 · The work

What do you want built?

A paragraph is genuinely enough to start. If it isn't work I'm right for, I'll say so and point you somewhere better.

The work

Pick everything that applies.

Platform

No idea is a perfectly good answer.

What are you trying to build, and what does it have to do for the people who use it? Write it the way you'd say it out loud.

0 / 1200

Two steps. Under a minute.