RAG development combines search infrastructure with language-model generation so an application can answer questions using controlled, up-to-date source material. A production system must do more than retrieve similar text: it must ingest and version documents, enforce permissions, select useful context, produce traceable citations, handle failure, and expose enough telemetry to improve quality over time.
The basic flow is straightforward: a user submits a question, the system retrieves relevant content, an LLM uses that content to compose an answer, and the application returns the response with appropriate citations. The engineering difficulty is in the boundaries between those steps. Poor chunking can hide the answer. Weak retrieval can supply plausible but irrelevant context. Excessive context can increase latency and cost. Missing authorization checks can expose information a user should never see.
This makes RAG development a software architecture problem, not merely a prompt-writing exercise. The application layer, retrieval services, model calls, data pipelines, evaluation process and human-review paths must work together.
What a production RAG architecture contains
A useful reference architecture separates the system into several responsibilities:
- Source systems: document stores, knowledge bases, ticketing systems, databases, file shares or approved external sources.
- Ingestion and transformation: extraction, cleaning, metadata enrichment, chunking, deduplication and version tracking.
- Indexing: keyword indexes, vector embeddings or a hybrid combination of both.
- Query processing: authentication, authorization filtering, query normalization and optional query expansion.
- Retrieval and reranking: candidate generation followed by relevance scoring and context selection.
- Generation: model routing, prompt construction, citation formatting and structured output validation.
- Application workflow: user interface, feedback, escalation, audit logging and business-system actions.
- Evaluation and observability: quality datasets, traces, cost data, latency measurements and failure analysis.
These components do not need to be separate products or services from the beginning. They do need clear ownership and interfaces. A Python service may handle ingestion, embeddings and retrieval while Laravel or another application layer owns users, permissions, workflows and administration. The right split depends on existing systems, team skills and operational requirements; a distributed architecture is not automatically better.
For broader implementation considerations, see AI development and the related guide to custom software development.
Ingestion determines what retrieval can find
Retrieval quality is constrained by the material placed in the index. Ingestion should therefore be treated as a repeatable data pipeline rather than a one-time upload script.
Extract structure before creating chunks
Documents often contain headings, tables, lists, page numbers, footnotes and access-control metadata. Flattening everything into plain text can remove the relationships that make an answer understandable. Preserve meaningful structure where possible, and record attributes such as source identifier, document version, section title, effective date, language and access scope.
Choose chunks for meaning and use
Fixed-size chunks are easy to implement, but they may split a definition from its conditions or separate a table heading from its values. Structure-aware chunking can keep related sections together. Overlapping chunks may help recall, but excessive overlap duplicates context and increases index size and generation cost.
Chunking should be tested against representative questions. A chunk is useful when it gives the retrieval and generation stages enough context to answer accurately without requiring the model to reconstruct a document from disconnected fragments.
Track versions and deletion
Production ingestion needs idempotency. If a source document changes, the system should be able to update or replace its indexed representations without leaving stale copies. Deletions matter just as much as additions, particularly when documents contain confidential or regulated information. A source-of-truth identifier and version strategy make reprocessing and audit work more predictable.
Combine retrieval methods instead of assuming embeddings solve search
Vector search is valuable when a question and an answer use different wording. It can identify semantically related passages even when there is limited exact term overlap. However, vector similarity is not a universal relevance signal.
Keyword or lexical search can be stronger for product codes, policy identifiers, names, error messages and exact legal language. A hybrid approach can combine lexical and semantic candidates before reranking. Metadata filters are also essential: the system may need to restrict results by tenant, department, document status, region, product, date or user permission.
Storage choices should follow operational requirements. A team may use an existing relational database with vector support, a dedicated search platform, or a managed vector service. The decision involves query patterns, scale, filtering, backups, data residency, team expertise, migration effort and total operating responsibility. The comparison in pgvector vs. Pinecone examines those trade-offs without treating one option as universally correct.
Use reranking to improve the final context
Initial retrieval is usually optimized for recall: return enough candidates that the relevant passage is unlikely to be missed. Reranking then applies a more precise relevance model or scoring process to order those candidates before context is sent to the LLM.
This two-stage design helps address a common failure mode: the correct document appears somewhere in the candidate set but is pushed below more broadly similar passages. Reranking can also reduce context size by selecting the most useful results rather than passing every retrieved chunk to the model.
Reranking introduces additional computation, so it should be applied where its quality benefit justifies added latency and cost. Some systems can use a lightweight path for simple queries and a deeper path for ambiguous or high-risk requests. The choice should be measured with a representative evaluation set rather than inferred from a few successful demonstrations.
Design citations as part of the response contract
Citations are not decorative footnotes. They provide a way for users to verify claims and for product teams to investigate incorrect answers. A citation should identify the source and, where practical, the relevant section, page, record or passage.
The generation layer should receive citation metadata alongside retrieved content. The application can then render citations from structured fields rather than asking the model to invent URLs or document references. This reduces formatting errors and makes it possible to validate that every cited source was actually retrieved.
Define what happens when the answer cannot be supported. Depending on the use case, the system may respond that the available sources do not establish an answer, ask a clarifying question, show relevant documents without summarizing them, or route the request for human review. A confident answer without adequate evidence is often worse than a transparent refusal.
Permissions and privacy must apply before generation
Authorization cannot be added only to the user interface. Retrieval must be scoped to the authenticated user, tenant, role or other access boundary before content reaches the model context. Filtering results after generation is too late: sensitive information may already have influenced the response.
Review the full data path, including source connectors, temporary files, embedding requests, model providers, logs, traces, caches and evaluation datasets. Minimize retained data where appropriate, redact sensitive fields when feasible, and define which teams can inspect prompts and retrieved passages. The exact controls depend on the data and applicable obligations, but the architectural principle is consistent: know where information travels and enforce access at every relevant boundary.
Evaluate retrieval and generation separately
A response can be wrong because retrieval failed, because the model misunderstood good context, or because the application formatted the result incorrectly. A single end-to-end score will not explain which layer needs work.
Build a test set containing realistic questions, expected source material and acceptable answer characteristics. Evaluate at least:
- Retrieval relevance: whether the correct source or passage is present among the candidates.
- Context quality: whether selected passages are sufficient, non-duplicative and appropriately scoped.
- Grounding: whether claims are supported by the supplied sources.
- Citation accuracy: whether references point to sources that actually support the answer.
- Abstention behavior: whether the system avoids unsupported answers when evidence is missing.
- Workflow success: whether users can complete the intended task with the response.
Automated evaluation can accelerate iteration, but human review remains important for ambiguous questions, nuanced policies and high-impact workflows. Store failed examples with enough trace context to reproduce them, while applying appropriate privacy controls.
Plan for latency, cost and model failure
A RAG request may involve authentication, query processing, multiple retrieval operations, reranking, one or more model calls and citation formatting. Each stage contributes latency and operating cost. Instrument them independently so teams can identify whether a slow response comes from the search layer, a remote model, oversized context or application code.
Model routing can match tasks to different models or processing paths. A simple factual lookup may not need the same generation model as a multi-document analysis. Caching can help with stable retrieval or repeated requests, but cached results must respect permissions and document freshness. Timeouts, retries and circuit breakers should prevent a slow provider from blocking the entire product experience.
Fallbacks should be explicit. The application might show search results when generation fails, use a smaller model for a constrained response, queue a longer analysis, or escalate to a person. These behaviors protect the underlying workflow instead of making the entire feature depend on one successful model call.
Connect RAG to product workflows carefully
RAG becomes commercially useful when it supports a real workflow: answering an internal policy question, summarizing a support history, preparing a case brief or helping an employee find approved procedures. The surrounding application should control identity, permissions, state changes and audit records.
For example, an AI support feature may retrieve knowledge-base articles and ticket history, draft a response with citations, and require approval before sending it to a customer. The model can assist with interpretation, while the application remains responsible for ticket ownership, status transitions and outbound communication. This separation makes human review and rollback practical.
When an existing web product owns users and business processes, Python can provide focused AI services without forcing the entire application to be rewritten. The architecture should define contracts for authentication, retrieval requests, structured responses, errors and observability. The article on Laravel and Python AI architecture explores this division in more detail.
Production readiness checklist for RAG development
- Identify authoritative sources and define how freshness is measured.
- Preserve document metadata, access scope and version information during ingestion.
- Test chunking against real questions rather than relying on a default size.
- Use lexical, semantic and metadata filtering where the domain requires it.
- Measure reranking quality and its effect on latency and cost.
- Generate citations from validated source metadata.
- Apply authorization before retrieved content enters model context.
- Separate retrieval, grounding and workflow evaluation.
- Log traces, latency, token usage, model selection and user feedback with privacy controls.
- Define refusals, fallbacks and human-review paths before launch.
- Version prompts, retrieval settings, models and evaluation datasets.
- Reprocess updates and deletions reliably.
Effective RAG development is disciplined systems engineering. Embeddings and an LLM may form the visible center of the feature, but dependable behavior comes from the surrounding pipeline: controlled data, relevant retrieval, defensible citations, permission-aware context, measurable quality and resilient product workflows. Designing those components together is what turns a retrieval demo into maintainable custom software.
For related model-application patterns, see OpenAI API integration and the comparison of RAG versus fine-tuning.