Front Matter
- Preface
- How to Read This Book
- Read This First — RAG Vocabulary at a Glance
- About the Author
Part I — Foundations
- The AI Landscape
- What an LLM is — and why parametric memory hallucinates
- The 2026 vocabulary: LLM, RAG, AI Agent, Agentic AI
- The honest comparison: RAG vs Fine-Tuning vs Long-Context Prompting
- When NOT to use RAG
- The evolution arc: RAG → Agentic RAG → AI Memory → KAG
- MCP, RAG, and Skills — three complementary architectures
- The two phases: indexing and query
- Your First RAG in 10 Minutes — Hello World
- The .NET Toolkit for RAG Development
- The two abstractions the code is built on
- The May 2025 renames you cannot afford to miss
- Ollama: use OllamaSharp, not Microsoft.Extensions.AI.Ollama
- AddSmartDocsCore unpacked
- Configuration: appsettings.json and user-secrets
- Migrating from Semantic Kernel to MAF
- Migrating from MAF preview to 1.0
- TokenCounter, the utility you'll use everywhere
- The companion solution layout
- One version, set once: Central Package Management
- The package map: what to add, and when
- HttpClient resilience: Microsoft.Extensions.Http.Resilience
- Local development: native services, not cloud-by-default
Part II — The RAG Pipeline
- Embeddings: Turning Text into Vectors
- What an embedding actually is
- Token limits and silent truncation
- Cosine, dot product, Euclidean: which one to use
- Asymmetric embeddings: query vs document
- The 2026 embedding leaderboard
- Domain adaptation: when default embeddings aren't enough
- The EmbeddingService domain port
- Batching and rate-limit handling
- Multilingual: BGE-M3 vs query translation
- Matryoshka representations and dimension trimming
- Benchmarking embeddings on your own corpus
- Chunking and Contextual Retrieval
- Why chunking is the highest-leverage tuning knob
- The four classic strategies
- The chunk-size question
- Characters versus tokens: sizing chunks the model's way
- Where every classic strategy breaks: the context problem
- Anthropic's Contextual Retrieval
- The numbers Anthropic reported, and ours
- The cost question, and prompt caching
- When the document is code: Roslyn-aware chunking
- Wiring chunking into the SmartDocs pipeline
- When contextual retrieval isn't worth it
- A cheaper cousin: Late Chunking
- The Ch04_ChunkingPlayground sample
- Where this thread continues
- Multimodal RAG
- What multimodal embedding models do
- Three architectural patterns
- One space, one dimension: the constraint everyone forgets
- PDF processing: the silent killer
- Tables: the in-between case
- Wiring multimodal into SmartDocs
- Grounding an answer in the image itself
- When pure text wins
- Vector Databases
- What HNSW actually does
- Tuning HNSW: the three knobs that move recall
- How much RAM will this need?
- The 2026 .NET landscape
- IVectorStore: the abstraction SmartDocs actually uses
- The MEVD record shape, for reference
- ID design and why upsert is idempotent
- Payload co-location: keep the text next to the vector
- Creating a collection: the pitfalls
- Filtering: a preview, not this chapter's job
- Quantization: the memory lever
- Measuring it: recall@k and latency percentiles
- The comparison sample
- Hybrid stores (forward reference)
- Choosing the right backend
- Indexing Strategies
- Strategy 1: full-chunk indexing (the default)
- Strategy 2: summary indexing
- Strategy 3: sub-chunk (small-to-big) indexing
- Strategy 4: hypothetical-question (query) indexing
- Where embedding and storage actually happen
- What the strategies actually score
- Choosing per use case
- The cost surface
- Wiring into SmartDocs
- Keeping the index fresh: updates, deletes, and re-embedding
- The Retriever
- Dense retrieval (vector search)
- Sparse retrieval (BM25 and friends)
- Hybrid retrieval: dense + sparse together
- Reciprocal Rank Fusion: the math
- Why hybrid wins
- Tuning hybrid: weights and oversampling
- ColBERT and late-interaction (briefly)
- Choosing K
- Diversity: trimming redundant results
- Score thresholding and abstention
- The retriever contract
- Latency budget
- Wiring into SmartDocs
- Reranking
- What a cross-encoder does differently
- The 2026 reranker landscape
- The IReranker interface
- The headline numbers
- Late-interaction (ColBERT) reranking
- LLM-as-reranker
- A self-hosted cross-encoder (ONNX BGE)
- Latency budget for rerank
- When NOT to rerank
- Thresholding after rerank
- Wiring into the SmartDocs pipeline
- The Complete RAG Pipeline
- The pipeline shape
- The augmentation step
- Generation: streaming with IChatClient
- Grounding: extracting and validating citations
- The API endpoint: SSE
- The Blazor consumer
- Error handling and partial responses
- Cost of a query: the prompt budget
- Putting it together: the trace of one query
Part III — Query Intelligence
- Metadata Filtering and Query Construction
- The two filtering positions
- Designing the metadata schema
- Filter expressions: MetadataFilter
- Query construction: from natural language to filter
- The filter-then-search pipeline
- Self-Query: a special case worth knowing
- When the filter is too tight: relaxation
- Operators beyond equality
- Validating extracted values against the vocabulary
- Security filters are auth-derived, never extracted
- A few more filtering concerns
- When NOT to construct queries
- Filter caching
- Wiring into SmartDocs
- Query Routing and Conversational Queries
- Why route at all
- The IQueryRouter interface
- Three routing strategies, plus a cascade
- The cheap→expensive cascade
- Confidence, fallback, and the refuse path
- Structured-output robustness
- Conversational queries: the rewrite problem
- Query rewriting with ConversationalQueryRewriter
- Managing conversation history beyond a magic number
- Wiring it all together
- Fan-out and fusion: when a query hits more than one silo
- Security: the router and rewriter are an injection surface
- Cost economics of routing + rewriting
- Routing telemetry: the signals that matter
Part IV — Graph and Hybrid Storage
- Graph Databases for RAG
- What a graph database stores that a vector store doesn't
- Neo4j with Neo4j.Driver 6.2.1
- Building the entity graph at ingest time
- Cypher in 60 seconds
- The graph retriever interface
- Route to one, or run both and fuse
- Vector + graph: GraphRAG (forward reference)
- When NOT to reach for a graph
- Wiring into SmartDocs
- Testing graph code without a live Neo4j
- Hybrid Databases
- The 2026 single-store options
- Postgres + pgvector + tsvector: the practical default
- Qdrant native hybrid
- Azure AI Search: hosted hybrid + semantic reranker
- Choosing among the three
- Filter interplay
- When two stores still win
- Wiring into SmartDocs
Part V — RAG Design Patterns
- Classic RAG Enhancements
- HyDE: Hypothetical Document Embeddings
- RAPTOR: Recursive Abstractive Processing for Tree-Organized Retrieval
- RAG-Fusion: query expansion with RRF
- Step-back prompting: retrieve the principles first
- Self-RAG: when to retrieve at all
- CRAG: Corrective RAG
- Lost in the middle: ordering chunks in the prompt
- Contextual compression: trimming chunks before generation
- Chunk pre-injection: loading context before the first turn
- Composing patterns
- What we're not covering
- Wiring into SmartDocs
- Vectorless RAG
- The kind of corpus this works on
- Building the structural index
- Retrieval over the structural index
- The query construction problem
- When structural beats vector cleanly
- When to fold structural into your stack
- Lexical retrieval: the other embedding-free pillar
- Choosing a vectorless retriever
- Normalizing the identifier before you look it up
- Versioning and point-in-time retrieval
- From retrieved nodes to a cited answer
- Five things that bite at scale
- The Ch16 sample
- GraphRAG and LazyGraphRAG
- The global-question problem
- GraphRAG: the eager indexing pipeline
- GraphRAG by the numbers
- The three query methods
- LazyGraphRAG: the cost-shifting trick
- The cost curve
- When to choose eager
- Implementation in SmartDocs
- Switching between eager and lazy
- What GraphRAG can't do
- Model Context Protocol
- What MCP actually is
- The two standard transports
- The .NET MCP SDK
- Tenant scoping at the boundary: security without a filter type
- When the server needs something from the caller: multi-round-trip requests
- The client side: calling MCP tools from a MAF agent
- Notifications and subscriptions
- The MCP-vs-A2A-vs-OpenAPI decision tree
- Security: the part the spec under-emphasizes
- Tool granularity: the design decision that decides UX
- Structured output and streaming progress
- Composing MCP servers
- Observability and deployment
- Testing MCP servers
- The two SmartDocs servers
- Agentic, Multi-Agent, and Memory
- From single-agent RAG to multi-agent workflows
- The Workflow graph: conditional and parallel edges
- Agentic memory: semantic, episodic, procedural
- CodeAct (a.k.a. Code Mode): one generated program instead of N tool calls
- From demo to production: the five concerns the prototype skips
- Tool design for agents (vs APIs for humans)
- Putting it together: the SmartDocs production agent
- The Ch19_MultiAgentOrchestration sample
Part VI — Production Concerns
- Evaluation and Metrics
- Why eval is the unsexy multiplier
- The five metrics that matter
- The .NET-native eval library: Microsoft.Extensions.AI.Evaluation
- RAGAS and DeepEval: when to leave .NET
- LLM-as-judge: how it actually works (and where it fails)
- Building the seed dataset
- The CI gate
- The promotion ratchet
- Latency and cost as eval metrics
- Ground-truth evaluation, the framework-native way (verified on MAF 1.20.0)
- The eval-driven development loop
- Evaluating for safety, not just quality
- Online versus offline evaluation
- What the eval can't tell you
- Latency, Cost, and Performance
- The 2026 cost economics
- The four-layer cache
- The other half of cost: send fewer tokens
- The latency budget per pipeline stage
- .NET-specific performance levers
- Resilience: Polly + OpenTelemetry
- Provisioned throughput vs pay-as-you-go
- Cost attribution: who pays for what
- Avoid the premature optimizations
- Freshness, Drift, and Migration
- Freshness as a first-class metric
- Three ingest strategies
- Embedding drift: the silent killer
- The Drift-Adapter pattern
- Cache invalidation under document churn
- Orphan chunks and re-chunking on update
- GDPR Article 17 and the right to erasure
- Schema migrations
- Backfill scheduling
- The Ch22_DriftAdapterMigration sample
- Security
- The threat model
- Layer 1: input layer
- Layer 2: embedding layer
- Layer 3: retrieval layer
- Layer 4: output layer
- Audit logs
- Secret scanning at ingest
- Supply-chain security
- MCP-server trust hierarchy
- Keyless authentication and infrastructure
- Agentic-action security: least privilege and human approval
- The 25 red-team cases
- Incident response shape
- Trust by Design
- Impact assessments before deployment: DPIA and FRIA
- Provider vs deployer: who owes what
- The GPAI layer underneath your RAG system
- The shape of the EU AI Act obligations
- Conformity assessment, CE marking, and EU-database registration
- Citations as grounding evidence
- Audit logs as action evidence
- Consent and transparency
- Lawful basis is not consent
- Data residency, sovereignty, and cross-border transfers
- Faithfulness as a contractual quality bar
- Bias, fairness, and non-discrimination
- Model cards and technical documentation
- Annex IV technical documentation is not the model card
- Human oversight
- Post-market monitoring
- Right to explanation
- Copyright, corpus licensing, and output IP
- The compliance dashboard
- The 'right to be forgotten' meets the embedding
- The Ch24 sample: trust by design, end to end
Part VII — Capstone Project
- Production Capstone
- The reference topology
- Bicep: the resource graph as code
- The GitHub Actions deployment workflow
- NuGet packaging for the public surface
- Secret management
- Smoke tests after deploy
- Production hardening
- Day-2 operations
- The pre-flight checklist
- Capacity planning
- The day after launch
- An alternative shape: vertical slices
- Where this book ends and your work begins
- Closing
Appendices
- Appendix A — Practice Exercise Solutions
- Appendix B — Design Pattern Quick Reference
- Appendix C — Vector Database Comparison
- Appendix D — Embedding Benchmark
- Appendix E — Python → .NET Rosetta
- Appendix F — Companion Repository Guide
- Appendix G — Debugging Checklist
- Appendix H — Prompt Engineering Patterns
- Appendix I — The Math Behind RAG
- Other Books You’ll Enjoy
- Index
