Understanding, Designing, and Deploying Retrieval-Augmented Generation Systems
Chapter 1: Why RAG? The Case for Retrieval-Augmented Generation
- The Knowledge Problem with Language Models
- A Simple Example: What Goes Wrong Without Retrieval
- The Core Idea of RAG in One Sentence
- What This Book Will Teach You
Chapter 2: Language Model Limitations and the Need for External Knowledge
- How Pretrained Language Models Store Knowledge
- Hallucinations: Causes and Manifestations
- The Training Cutoff Problem
- Parameter-Efficient Knowledge Updating
- Why Finetuning Alone Is Not Enough
- The Scalability Argument for Retrieval
Chapter 3: Information Retrieval Foundations
- What Is Information Retrieval
- The Vector Space Model and TF-IDF
- The BM25 Algorithm
- Inverted Indexes and Boolean Retrieval
- Relevance Metrics: Precision, Recall, NDCG
- The History That Led Us to RAG
Chapter 4: Embeddings and Vector Representations
- From Words to Vectors: Word Embeddings to Sentence Embeddings
- What an Embedding Actually Represents
- Architecture of Modern Embedding Models
- Cosine Similarity and Vector Distance Metrics
- Evaluating Embedding Quality
- Choosing an Embedding Model
Chapter 5: Similarity Search and Vector Databases
- The Nearest Neighbor Problem at Scale
- Brute-Force Search and Its Limits
- Approximate Nearest Neighbor Algorithms: HNSW, IVF, PQ
- Vector Database Architecture and Design
- Choosing a Vector Database
- Vector Index Maintenance
Chapter 6: Document Ingestion and Preprocessing
- The Ingestion Pipeline End-to-End
- Document Format Parsing: PDFs, HTML, Office Documents
- Text Cleaning and Normalization
- Handling Special Content: Tables, Code, Images, Math
- Metadata Extraction and Structuring
- Building a Robust Ingestion System
Chapter 7: Chunking Strategies
- Why Chunking Matters
- Fixed-Size and Sliding-Window Chunking
- Semantic and Recursive Chunking
- Chunking for Code, Legal Documents, and Technical Manuals
- Parent-Child and Hierarchical Chunking
- Measuring and Tuning Chunk Quality
Chapter 8: Index Construction
- Embedding and Storing Chunks
- Building Vector Indexes at Scale
- Combining Dense and Sparse Indexes
- Metadata Indexing and Filtering
- Incremental Indexing and Data Freshness
- Index Versioning and Rollbacks
Chapter 9: Basic Retrieval
- The Retrieval Operation
- Dense Retrieval Mechanics
- Sparse Retrieval Mechanics
- Combining Dense and Sparse: Hybrid Search
- Metadata Filtering and Faceted Search
- Retrieval Performance Tuning
Chapter 10: Query Transformation and Expansion
- The Query Transformation Problem
- Query Rewriting for Clarity and Completeness
- Multi-Query Retrieval
- Hypothetical Document Embeddings (HyDE)
- Query Decomposition for Complex Questions
- When Query Transformation Hurts
Chapter 11: Reranking
- Why Reranking Is Necessary
- Cross-Encoder Architecture
- Reranker Training and Fine-Tuning
- Late-Interaction and ColBERT Models
- Practical Reranking: Latency vs. Quality Trade-Offs
- Choosing and Configuring a Reranker
Chapter 12: Context Construction and Prompt Design
- From Chunks to Context
- Formatting Retrieved Context for the LLM
- Context Compression and Summarization
- Token Budget Management
- Prompt Design for Grounded Generation
- Structuring Multi-Chunk Contexts
Chapter 13: The Generation Step
- Guiding the LLM with Retrieved Context
- Temperature, Top-P, and Generation Control
- Enforcing Faithfulness to Context
- Handling Conflicting or Insufficient Evidence
- Streaming Responses and Partial Answers
- Generation Failure Modes
Chapter 14: Building a Naïve RAG Pipeline
- Designing the Naïve RAG Architecture
- Project Structure and Dependencies
- Ingestion and Indexing Code
- Retrieval Implementation
- Generation Integration
- End-to-End Execution and Testing
Chapter 15: Advanced RAG Patterns
- Hybrid Search with Reranking: The Standard Advanced Pattern
- Contextual Compression Retrieval
- Self-Correction and Self-RAG
- Adaptive Retrieval Strategies
- Implementing Advanced Patterns
Chapter 16: Modular RAG Architectures
- The Case for Modularity
- Toolformer-Inspired Approaches
- Modular Design Patterns
- Routing and Orchestrating Modules
- Implementing a Modular RAG System
Chapter 17: Hierarchical and Parent-Child Retrieval
- The Granularity Problem
- Parent-Child Retrieval
- Multi-Level Hierarchical Indexes
- Small-to-Big Retrieval
- Implementing Hierarchical Retrieval
Chapter 18: Multi-Hop and Agentic RAG
- Single-Hop vs. Multi-Hop Reasoning
- Iterative Retrieval Strategies
- Agentic RAG: Reasoning, Planning, Acting
- Search and Retriever Tools for Agents
- Controlling Agentic RAG Complexity
- Implementing Multi-Hop Retrieval
Chapter 19: GraphRAG and Knowledge Graphs
- Why Graphs for RAG
- Knowledge Graph Construction from Documents
- Graph-Based Retrieval
- GraphRAG Architecture Patterns
- Implementing GraphRAG
Chapter 20: Conversational RAG and Multimodal RAG
- Conversational RAG: Maintaining Context Across Turns
- Conversational Query Understanding
- Conversation History in the Prompt
- Multimodal Embeddings and Retrieval
- Image, Audio, and Video in RAG
- Implementing Conversational RAG
Chapter 21: Long-Context Approaches and Structured Data
- Long-Context LLMs: RAG Competitors or Complements
- Attention at Scale
- When Long Context Is Preferable to RAG
- Hybrid: RAG with Long-Context Models
- Retrieval from Structured and Tabular Data
- SQL Retrieval and Text-to-SQL
- Retrieval from Structured and Tabular Data
Chapter 22: RAG Evaluation: Measuring What Matters
- What to Evaluate: A Complete Taxonomy
- Retrieval Quality Metrics: Recall, Precision, MRR, NDCG
- Answer Quality and Faithfulness Metrics
- Hallucination Detection
- Automated Evaluation: RAGAS, TruLens, and Custom Metrics
- Human Evaluation and Benchmark Design
Chapter 23: Debugging RAG Systems
- Classifying RAG Failures
- Retrieval Failures: Diagnosing and Fixing
- Generation Failures: Diagnosing and Fixing
- End-to-End Debugging Workflow
- Logging, Tracing, and Observability
- Continuous Improvement Loops
Chapter 24: Security and Access Control for RAG
- Prompt Injection: Direct Attacks
- Prompt Injection: Indirect Attacks Through Retrieved Documents
- Authorization-Aware Retrieval
- Data Leakage
- Tenant Isolation in Multi-Tenant Systems
- Malicious Documents and Content Poisoning
- Red-Teaming and Security Testing
Chapter 25: Production Architecture and Deployment
- Small-Scale Architecture: Single-Service Design
- Enterprise Architecture: Modular and Scalable Design
- Multi-Tenant Architecture
- Deployment Options: Cloud, On-Premise, and Hybrid
- Scalability Patterns
- High Availability and Fault Tolerance
Chapter 26: Production Operations and Optimization
- Monitoring and Observability
- Latency Optimization
- Cost Optimization
- Data Freshness and Incremental Indexing
- Versioning, Rollback, and Continuous Deployment
Chapter 27: Real-World Case Studies and Patterns
- Case Study 1: Customer Support Bot for SaaS Company
- Case Study 2: Enterprise Knowledge Management System
- Case Study 3: Financial Research and Analysis
- Case Study 4: Healthcare Information System
- Case Study 5: Developer Documentation Assistant
- Case Study 6: E-Commerce Product Information System
- Lessons Across Case Studies
Chapter 28: Conclusion: The RAG Landscape and Future Directions
- The Core Insight: Why RAG Endures
- Key Principles for Effective RAG
- The RAG Technology Landscape
- Future Directions: Where RAG Is Heading
- Getting Started: Practical Guidance
- The Promise of RAG