A Practical Engineering Guide to Benchmarking, Optimization, and Production Deployment
Introduction
- Why evaluation matters more for local models
- What you will be able to do after reading this book
- Who this book is for
- How this book is organized
- A note on pace and specificity
Chapter 1: Defining the Landscape
- What is a Small Language Model
- What is a Local Language Model
- Why Local Matters
- The Hardware Landscape
- Cost and Privacy Implications
- How This Book is Structured
Chapter 2: Model Architectures and Scaling
- Transformer Fundamentals
- Parameter Count and Scaling Laws
- Attention Mechanisms and Variants
- Grouped-Query and Multi-Query Attention
- Mixture of Experts Architectures
- Context Window Mechanics
- Architectural Choices that Affect Inference
Chapter 3: Model Quality and Capability Dimensions
- Instruction Following and Alignment
- Reasoning and Chain-of-Thought
- Coding and Technical Tasks
- Structured Output and Format Control
- Multilingual Performance
- Long-Context Understanding
- Tool Use and Function Calling
- Hallucination and Factuality
- Refusal Behavior and Safety
Chapter 4: Model Weights, Formats, and Storage
- Safetensors and PyTorch Weights
- GGUF Format and llama.cpp
- ONNX and Cross-Platform Export
- Model Card Information
- Quantization-Aware Format Choices
- Storage Requirements and Disk I/O
- Weight Loading Performance
Chapter 5: Quantization Fundamentals
- Why Quantize
- Full-Precision vs. Quantized Representations
- INT8 and INT4 Quantization
- Per-Tensor vs. Per-Channel vs. Per-Token
- Quantization-Aware Training and Post-Training
- GGUF Quantization Schemes
- Measuring Quantization Impact
- When Quantization Fails
Chapter 6: The Local Inference Stack
- Hardware Abstraction Layers
- GPU Runtimes and Drivers
- Memory Management and Allocation
- Context Switching and Scheduling
- Process Isolation and Resource Limits
- Container Considerations
- The Inference Stack Architecture
Chapter 7: llama.cpp and GGUF Ecosystem
- llama.cpp Architecture and Design
- GGUF Model Loading
- Quantization Selection in llama.cpp
- Threading and CPU Optimization
- Metal and GPU Offloading
- API and Embedding Integration
- Performance Tuning and Flags
- Limitations and Gotchas
Chapter 8: Ollama, MLX, and Alternative Runtimes
- Ollama Architecture and Modelfile
- Running and Customizing Ollama
- Apple MLX Framework
- vLLM for Local Deployment
- Text Generation Inference
- Choosing Your Runtime
Chapter 9: Tokenization and Prompt Processing
- How Tokenization Works
- Vocabulary Size and Efficiency
- Tokenizer Choice Matters
- Measuring Tokenization Performance
- Special Tokens and Formatting
- Prompt Templates and System Messages
- Token Counting and Estimation
- Multilingual Tokenization Issues
Chapter 10: Latency and Throughput Fundamentals
- Time-to-First-Token
- Inter-Token Latency
- Tokens Per Second
- Prompt Processing Speed
- Batch Processing and Parallelism
- What Drives Each Metric
- Hardware Bottleneck Identification
- Measuring Without Contamination
Chapter 11: Memory Behavior and KV Cache
- What is the KV Cache
- KV Cache Growth and Context Windows
- Memory Planning and Budgeting
- KV Cache Quantization
- Sliding Windows and Attention Limits
- Prompt Caching Strategies
- Eviction and Reuse
- Practical Memory Monitoring
Chapter 12: Designing Sound Evaluation Methodology
- Evaluation vs. Benchmarking
- Task Selection and Workload Definition
- Representative Test Construction
- Prompt Design and Variance
- Deterministic vs. Stochastic Evaluation
- Temperature and Sampling Parameters
- Sample Size and Saturation
- Evaluation Framework Design
Chapter 13: Statistical Rigor and Experimental Design
- Reproducibility and Seeding
- Confidence Intervals for Model Comparison
- Statistical Significance Testing
- Variance Sources and Control
- Measurement Error and Noise
- Warm-up Effects and Caching
- Normalizing Across Hardware
- Reporting Standards
Chapter 14: Public Benchmarks and Their Limitations
- Multiple-Choice Knowledge Benchmarks
- Math and Reasoning Benchmarks
- Code Generation Benchmarks
- Instruction and Chat Benchmarks
- Benchmark Contamination
- Overfitting to Benchmarks
- Aggregate Scores and Their Fallacies
- When to Trust Public Benchmarks
Chapter 15: Building Custom Evaluation Harnesses
- Evaluation Harness Architecture
- Implementing a Basic Harness
- Streaming and Telemetry Collection
- Automated Experiment Orchestration
- Result Storage and Retrieval
- Structured Output Validation
- Comparative Analysis and Reporting
- Visualization and Dashboards
Chapter 16: Measuring Model Quality in Depth
- Automated Quality Metrics
- Factuality Verification
- Hallucination Measurement
- Instruction Following Assessment
- Reasoning Quality Evaluation
- Code Execution Verification
- Long-Context Fidelity Testing
- Model-as-Judge Evaluation
Chapter 17: Workload-Specific Evaluation
- Conversational Assistants
- Coding and Development Tools
- Document Analysis and Summarization
- Retrieval-Augmented Generation Systems
- Agentic Workflows
- Edge and Offline Deployments
- Enterprise and Private Assistants
- Embedding and Classification Tasks
Chapter 18: Safety, Robustness, and Reliability
- Safety Evaluation Methodology
- Adversarial Prompt Testing
- Jailbreak Resistance
- Refusal Calibration
- Consistency and Reliability Under Load
- Degradation Modes
- Monitoring for Production Safety
- Compliance and Policy Checks
Chapter 19: Advanced Optimization Techniques
- Speculative Decoding
- Continuous Batching
- Prompt and KV Cache Compression
- Compiler Optimizations
- Distillation and Adapters
- Model Merging Strategies
- Hardware-Specific Optimizations
- Distributed and Multi-Node Inference
- Techniques for Improving Performance Without Compromising Quality
Chapter 20: Decision Frameworks and Production Deployment
- Multi-Criteria Decision Frameworks
- Model Selection by Workload
- Hardware Procurement Decisions
- Cost Modeling and TCO
- Operational Considerations
- Case Study: Local Coding Assistant
- Case Study: Private RAG System
- Case Study: Edge Deployment