Skip to content
AI & Automation

Retrieval-Augmented Generation (RAG): A Complete Guide for Businesses

What retrieval-augmented generation is and how to build it: ingestion, chunking, embeddings, hybrid retrieval, reranking, grounded generation with citations, evaluation, costs and common failure modes.

Quick answer

Retrieval-augmented generation (RAG) answers questions using your own information. At indexing time, documents are collected, parsed, split into chunks, embedded and indexed with metadata and permissions. At question time, the system retrieves the most relevant chunks (ideally with hybrid keyword and vector search plus reranking), gives them to a language model with instructions to answer only from them and cite sources, and refuses when evidence is missing. Evaluate retrieval and answers separately, because most failures start with retrieving the wrong content.

Where This Fits

This is the hub for ZSpace Labs' RAG guides. Deeper topics: chunking, embeddings, vector databases, hybrid search, reranking, enterprise RAG architecture, GraphRAG and RAG vs fine-tuning. For a product view, see AI knowledge base.

For building complete generative AI products around RAG, see generative AI application development.

How RAG Works

RAG has two phases. Indexing prepares your content: connectors pull documents, parsers extract text and structure, chunking splits it into retrievable passages, an embedding model converts passages into vectors, and an index stores vectors, text, metadata and permissions. Query time finds and uses that content: the question may be rewritten, retrieval finds candidate passages, a reranker orders them, and the model generates an answer from the top passages with citations.

Query time is where most quality is won or lost.

Ingestion and Parsing

Garbage in, garbage retrieved. Parse documents in a way that keeps structure: headings, lists, tables and page numbers. PDFs with columns, scanned pages and complex tables need specialised parsing or OCR. Store metadata (source, title, section, date, owner, access groups) with every chunk; it powers filtering, citations and freshness. Re-index on change rather than on a slow schedule where content changes often.

Chunking and Embeddings

Chunks should be small enough to be specific and large enough to make sense on their own. Structure-aware chunking (by section) usually beats fixed-size splitting for business documents; see RAG chunking strategies. Embeddings turn chunks into vectors for semantic search; the model you choose affects quality, cost and language support; see vector embeddings explained.

Retrieval: Hybrid Search, Filters and Reranking

Vector search finds passages with similar meaning; keyword search (BM25) finds exact terms such as product codes, names and error messages. Combining them in hybrid search usually beats either alone. Apply metadata filters (permissions, product, date) during retrieval. Then rerank a larger candidate set with a cross-encoder or similar model so the best passages reach the prompt.

TechniqueWhat it fixes
Query rewritingVague or conversational questions
Hybrid searchMissed exact terms and codes
Metadata filtersWrong product, region, date or permission
RerankingRelevant passages ranked too low
Parent-document retrievalChunks too small to answer alone

Generation: Grounded Answers With Citations

Instruct the model to answer only from the provided sources, cite them, say when the sources do not contain the answer and avoid speculation. Keep the context focused: more passages are not always better, and irrelevant text can confuse the model and raises cost. Validate citations where accuracy matters (does the cited passage support the claim?). For structured outputs from documents, see AI document extraction.

Building an AI assistant on your company's documents?

ZSpace Labs builds RAG systems with permission-aware retrieval, citations and evaluation, connected to the sources your teams already use.

Start a Project

Evaluating RAG

Build a question set from real queries with expected answers and the sources that contain them. Measure retrieval (are the right sources in the top results?) separately from generation (is the answer faithful to the sources, correct and complete?). Use deterministic checks where possible and calibrated LLM judges for faithfulness. Re-run on every change to parsing, chunking, embeddings, retrieval settings or model. See AI evaluation.

  • Retrieval recall at k: is a correct source in the top k?
  • Ranking quality: how high does the correct source appear?
  • Faithfulness: are all claims supported by retrieved text?
  • Answer correctness and completeness
  • Correct refusals when the answer is not in the sources
  • Latency and cost per question

Security and Permissions

A RAG system must not show people documents they cannot open in the source system. Copy access controls with content and filter at retrieval time by the user's identity and groups. Treat retrieved text as untrusted: a document could contain instructions aimed at the model, so retrieval-only assistants should not have powerful tools. The OWASP Top 10 for LLM Applications lists vector and embedding weaknesses among its risks. See enterprise RAG architecture.

Costs

Indexing costs come from parsing and embedding (mostly up front, plus updates). Query costs come from retrieval infrastructure, reranking and model tokens, which scale with the number of passages included. Keep context lean, cache frequent answers where safe and choose the smallest model that meets quality targets; see LLM cost optimization.

Advantages and Limitations

AdvantagesLimitations
Answers from current, private informationQuality depends on content quality and coverage
Citations make answers checkableRetrieval can miss or misrank the right passage
Update knowledge by re-indexing, not retrainingComplex questions across many documents are hard
Permissions can mirror source systemsNeeds ongoing evaluation and content ownership

How to Build a RAG System Step by Step

  • 1. Define the questions users need answered and collect real examples
  • 2. Inventory sources and their owners, formats and permissions
  • 3. Build ingestion with structure-preserving parsing and metadata
  • 4. Choose chunking and embeddings and test on your questions
  • 5. Implement hybrid retrieval with filters and reranking
  • 6. Write generation instructions for grounded, cited answers and refusals
  • 7. Evaluate retrieval and answers separately and fix the weakest stage
  • 8. Launch with feedback and a process for content owners to fix gaps

RAG Architecture Patterns

PatternHow it worksWhen to use
Basic RAGRetrieve top chunks, generate oncePrototypes, simple FAQs
Advanced RAGQuery rewriting, hybrid search, reranking, citationsMost production systems
Parent-document RAGMatch small chunks, pass larger sectionsLong structured documents
Agentic RAGAn agent decides when and what to retrieve, possibly several timesComplex, multi-part questions
GraphRAGKnowledge graph and community summariesRelationship and corpus-wide questions

Tools and Technology Choices

A RAG stack typically includes connectors and parsers, an embedding model, an index (Postgres with pgvector, a dedicated vector database or a search engine with vector support), a reranker, a language model and an evaluation harness. Frameworks such as LlamaIndex and LangChain speed up assembly; cloud platforms and enterprise search products offer managed options. Choose components based on your sources, scale, permission model and data residency, and keep them swappable behind your own interfaces. Storage choices are compared in vector databases for AI.

RAG Use Cases by Function

FunctionQuestions RAG answersTypical sources
HR and peopleLeave, benefits, policies, onboardingHandbooks, policy sites
Customer supportProduct how-to, troubleshooting, policiesHelp centre, runbooks, resolved tickets
SalesProduct capabilities, pricing rules, security answersProduct docs, approved security questionnaires
Legal and complianceClause positions, policy interpretationPlaybooks, contract libraries
Engineering and ITArchitecture decisions, runbooks, incidentsWikis, repositories, postmortems
OperationsProcedures, specifications, standardsSOPs, manuals, specifications

Operating RAG in Production

Launching is the start. Production RAG needs ingestion monitoring (failed syncs, parsing errors, document counts), freshness tracking, evaluation runs after every pipeline change, sampled answer reviews, user feedback triage and cost and latency dashboards. Assign owners: an engineering owner for the pipeline and content owners for each source area. Schedule re-evaluation when you change embedding models, chunking, retrieval settings or the generation model, and keep the previous index available until the new one is proven.

  • Ingestion health and freshness alerts
  • Evaluation in CI for pipeline and prompt changes
  • Weekly review of low-rated answers and unanswered questions
  • Content owner reports per source area
  • Cost per question and latency percentiles
  • Versioned indexes with rollback

Worked Example

An illustrative scenario, not a client case: an engineering firm's RAG assistant gives vague answers about project standards. Evaluation shows the right document is retrieved but split mid-table, and part numbers are missed by vector search. Switching to section-aware chunking that keeps tables intact and adding keyword search with rank fusion fixes most failures; a reranker improves the rest.

Common Mistakes

  • Tuning prompts when retrieval is the problem
  • Vector-only search for content full of codes and names
  • Losing tables and headings during parsing
  • No permission filtering
  • No evaluation set
  • Stale indexes nobody owns

Want a RAG system your team can trust?

Talk to ZSpace Labs about RAG development and data integration and deployment.

Start a Project

Conclusion

RAG is a retrieval problem first and a generation problem second. Invest in parsing, chunking, hybrid retrieval, reranking, permissions and evaluation, and keep content owners involved. Next: enterprise RAG, hybrid search and AI knowledge base.

FAQ

Common questions

A technique where an AI system first retrieves relevant information from your own sources, such as documents or databases, and then gives it to a language model to generate an answer grounded in that information, usually with citations.

Get in touch

Have a project in mind?

Whether you're building a new digital product, improving an existing website, or looking to automate part of your business — let's talk.

Keep exploring
AI & Automation
6 min read

Enterprise RAG Architecture: How to Build AI Systems With Company Data

How to architect RAG for an organization: source connectors, ingestion pipelines, access control sync, permission-aware retrieval, indexing, freshness, monitoring, governance and deployment.

Read article
AI & Automation
6 min read

RAG vs Fine-Tuning: Which Approach Should You Choose for AI Applications?

How RAG and fine-tuning differ: knowledge versus behaviour, freshness, data needs, cost, citations and maintenance, with a decision process and when combining them makes sense.

Read article
AI & Automation
6 min read

AI Knowledge Base: How to Build an AI Assistant That Uses Company Documents

How to build an AI knowledge base assistant: choosing sources, ingestion, permissions, retrieval, cited answers, refusals, feedback loops, content ownership, rollout and measurement.

Read article