Grounding AI in What's Actually True.
A language model without retrieval is a brilliant amnesiac — it knows a lot, but only up to a point in time, and it can't access your data, your documents, or your domain knowledge. Retrieval-Augmented Generation changes that. RAG gives AI systems a memory that can be updated, queried, and trusted. It's the difference between an AI that hallucinates confidently and one that answers from evidence. I think it's one of the most practically important techniques in applied AI, and I build RAG pipelines into almost everything I deploy.
What Is RAG?
"RAG is the difference between asking someone what they remember and handing them the relevant documents before they answer."
Retrieval-Augmented Generation (RAG) is a technique that combines a language model's reasoning capabilities with a retrieval system's ability to find relevant information at query time. Instead of relying solely on what the model learned during training, RAG retrieves relevant documents or chunks from an external knowledge base and injects them into the model's context before generating a response.
The pipeline has three core stages: at index time, documents are chunked, embedded into vectors, and stored in a vector database. At query time, the user's question is embedded, the most semantically similar chunks are retrieved, and those chunks are passed to the language model as context. The model then generates a response grounded in the retrieved evidence.
The result is an AI system that can answer questions about your specific data, stay current without retraining, cite its sources, and fail gracefully when it doesn't know something — rather than inventing a plausible-sounding answer.
Why RAG Matters for Agent Systems
For agents that need to act in the real world, hallucination isn't just an annoyance — it's a failure mode. RAG is how you give agents grounded, verifiable knowledge.
Eliminates Hallucination on Domain Knowledge
Language models hallucinate when asked about things outside their training data. RAG replaces guessing with retrieval — the model answers from documents you control, not from statistical patterns in its weights.
Keeps Knowledge Current
Training cutoffs make models stale. RAG knowledge bases can be updated continuously — new documents, new data, new policies — without retraining or fine-tuning the model. The knowledge layer and the reasoning layer evolve independently.
Enables Source Attribution
When an agent answers from retrieved chunks, it can cite exactly which document, page, or section it drew from. This makes AI outputs auditable and verifiable — critical for any production use case where trust matters.
Scales to Arbitrary Knowledge
A model's context window is finite. A vector database is not. RAG lets agents reason over millions of documents by retrieving only the relevant subset at query time — the model sees what it needs, not everything.
Composable with MCP
RAG pipelines exposed through MCP servers become reusable tools any agent can call. An agent doesn't need to know how retrieval works — it just calls the knowledge tool and gets grounded context back.
Graceful Degradation
A well-built RAG system knows when it doesn't know. When retrieval returns low-confidence results, the system can say so — rather than generating a confident but wrong answer. Calibrated uncertainty is a feature, not a bug.
Systems I'm Building
Active RAG projects in the lab — each one solving a real knowledge retrieval problem.
Lab Knowledge Base
A RAG system over my own lab documentation — architecture decisions, runbooks, experiment notes, and research papers. Exposed as an MCP server so any agent in the lab can query it. The system ingests markdown, PDFs, and structured notes, and returns answers with source citations.
Use case
Internal agent knowledge layer
Stack
pgvector · PostgreSQL · OpenAI text-embedding-3-large · TypeScript
Technical Documentation RAG
A retrieval system over technical documentation for the tools and frameworks I use — API references, SDK docs, changelog entries. Agents can query it when they need accurate, up-to-date technical details rather than relying on potentially stale training knowledge.
Use case
Agent tool-use grounding
Stack
pgvector · PostgreSQL · OpenAI Embeddings · Cheerio scraper
Multi-Collection Federated RAG
An experimental system that routes queries across multiple specialized knowledge collections — routing the query to the most relevant collection before retrieval, then merging and reranking results. Designed for cases where a single flat index isn't the right structure.
Use case
Cross-domain agent knowledge
Stack
pgvector · LangChain · Cohere Rerank · TypeScript
How I Structure RAG Pipelines
Every RAG system I build follows the same two-phase architecture — index time and query time — with deliberate decisions at each stage.
Document Ingestion
Source documents are loaded from their origin — file system, S3, database, web scraper, or API. Each document is normalized to a consistent internal format with metadata preserved.
Chunking
Documents are split into chunks sized for the embedding model and retrieval context. Chunk boundaries respect semantic units — paragraphs, sections, code blocks — not arbitrary character counts.
Embedding
Each chunk is embedded into a high-dimensional vector using an embedding model. The vector captures the semantic meaning of the chunk, enabling similarity search rather than keyword matching.
Storage
Vectors and their source chunks are stored in pgvector. Metadata — source document, page, section, timestamp — is stored alongside for filtering and attribution.
Query Embedding
The user's query is embedded using the same model used at index time. This ensures the query vector lives in the same semantic space as the document vectors.
Retrieval
The top-k most similar chunks are retrieved via approximate nearest-neighbor search. Metadata filters can narrow the search to specific collections, date ranges, or document types.
Reranking
Retrieved chunks are reranked using a cross-encoder model that scores each chunk against the full query. This catches cases where vector similarity missed the most relevant result.
Generation
The top reranked chunks are injected into the language model's context as grounding evidence. The model generates a response that cites its sources and stays within the retrieved evidence.
Chunking & Embedding Strategy
Chunking is where most RAG systems fail. The right strategy depends on the document type, the query patterns, and the embedding model.
Semantic Chunking
Split on semantic boundaries — paragraphs, sections, headings — rather than fixed character counts. A 512-token chunk that cuts mid-sentence retrieves worse than a 600-token chunk that ends at a natural boundary.
Hierarchical Chunking
Store chunks at multiple granularities — document summary, section summary, paragraph. Retrieve at the paragraph level but include section context in the generation prompt. Gives the model both precision and context.
Sliding Window with Overlap
Chunks overlap by 10–20% to avoid splitting relevant content across chunk boundaries. The overlap ensures that a sentence at the end of one chunk also appears at the start of the next.
Structured Extraction
For structured data — tables, JSON, code — extract and embed the structured representation directly rather than treating it as prose. A table row embeds better as 'field: value' pairs than as raw HTML.
I use OpenAI text-embedding-3-large as the default embedding model — 3072 dimensions, strong multilingual performance, and good separation between semantically distinct chunks. For latency-sensitive applications I drop to text-embedding-3-small with minimal quality loss.
Retrieval & Reranking
Vector similarity gets you close. Reranking gets you accurate. I treat them as two distinct stages with different jobs.
Stage 1: Vector Retrieval
Approximate nearest-neighbor search over pgvector using cosine similarity. Fast, scalable, and good enough to get the right documents into the candidate set. I retrieve top-20 to top-50 candidates — more than I'll use, but enough to give the reranker good material to work with.
Stage 2: Cross-Encoder Reranking
A cross-encoder model scores each candidate chunk against the full query — not just the query vector, but the actual query text. Cross-encoders are slower than vector search but dramatically more accurate at identifying the most relevant chunks. I use Cohere Rerank for production and a local cross-encoder for offline use.
Stage 3: Context Assembly
The top reranked chunks are assembled into a context block with source citations. I include the chunk text, the source document name, and the section heading. The generation prompt instructs the model to answer only from the provided context and to cite sources explicitly.
Evaluation & Quality Metrics
A RAG system you can't measure is a RAG system you can't improve. I track these metrics across every system I build.
Retrieval Recall@k
What fraction of the time does the correct document appear in the top-k retrieved results? This measures the retrieval stage in isolation, before reranking. A low recall@k means the problem is in the index or the embedding model, not the generator.
MRR (Mean Reciprocal Rank)
How high does the most relevant chunk rank in the retrieved results? MRR penalizes systems that find the right document but bury it at position 8. A good reranker should push the most relevant chunk to position 1 or 2.
Answer Faithfulness
Does the generated answer stay within the retrieved context, or does the model add information from its training weights? I use an LLM-as-judge approach to score faithfulness — asking a separate model to verify each claim in the answer against the retrieved chunks.
Answer Relevance
Does the generated answer actually address the question asked? A faithful answer that doesn't answer the question is still a failure. Relevance is scored separately from faithfulness.
Context Precision
Of the chunks retrieved, what fraction were actually useful for answering the question? High context precision means the retrieval is tight. Low precision means the model is wading through noise to find the signal.
Latency P95
End-to-end latency at the 95th percentile — from query received to response returned. RAG adds latency at every stage. I track P95 rather than average because tail latency is what users actually experience.
Challenges & Lessons Learned
Building RAG systems that work in production is harder than building ones that work in demos. These are the lessons that cost me the most time.
Chunking Strategy Is the Highest-Leverage Decision
I spent weeks tuning embedding models and retrieval parameters before realizing the real problem was chunking. Bad chunk boundaries — splitting mid-sentence, mid-table, mid-code-block — degraded retrieval quality more than any other factor. Fix chunking first.
Vector Search Alone Is Not Enough
Pure vector similarity retrieval has a ceiling. It finds semantically similar chunks, but 'semantically similar' isn't always the same as 'most relevant to this specific question.' Adding a reranking stage consistently improved answer quality by 15–30% on my evaluation sets.
Metadata Filtering Is Underrated
Filtering by document type, date range, or collection before vector search dramatically improves precision — especially in multi-collection systems. Don't make the retrieval system search everything when you know the answer is in a specific subset.
Evaluation Is Not Optional
Without a test set and metrics, you're flying blind. Every change I made to chunking, embedding, or retrieval that I thought would help sometimes hurt. You can't know without measuring. Building an evaluation harness early is the best investment in a RAG project.
The Generation Prompt Matters as Much as Retrieval
A good retrieval stage can be undone by a bad generation prompt. The prompt needs to explicitly instruct the model to stay within the retrieved context, cite sources, and express uncertainty when the context doesn't contain the answer. Without this, models drift back to their training weights.
Tech Stack
The tools and frameworks I use to build and run RAG systems.
Storage & Indexing
Embedding & Reranking
Pipeline & Orchestration
Explore More of the Lab
RAG systems are the knowledge layer that grounds agent reasoning. See how they connect to the rest of the stack.