A Systems Engineer’s Guide to Edge Inference on Resource-Constrained Devices
Introduction
- What This Book Covers
- How to Read This Book
- Technology Context
- The Central Thesis
Chapter 1: The Physics of Tiny
- Why Memory Rules Everything
- SRAM, DRAM, and Flash: Capacity, Speed, and Cost
- Memory Hierarchy and Cache Behavior
- Compute vs. Bandwidth: Arithmetic Intensity
- Power Budgets and Thermal Walls
- The Numbers: How Much Memory and Bandwidth Does an LLM Actually Need
Chapter 2: Inside an LLM Inference Step
- Tokenization: Strings to Integers
- The Embedding Layer and Positional Encoding
- Transformer Blocks: Self-Attention and Feed-Forward
- Multi-Head and Grouped-Query Attention
- Layer Normalization and Residual Connections
- Autoregressive Decoding and the KV Cache
- The Complete Inference Dataflow
Chapter 3: Measuring Feasibility
- Parameter Storage: Bits, Bytes, and Formats
- Activation Memory: What You Cannot See Coming
- KV Cache: The Hidden RAM Killer
- Arithmetic Intensity and Compute Requirements
- Bandwidth Calculations: The True Bottleneck
- The Feasibility Checklist
Chapter 4: Model Architectures for Constrained Inference
- Dense Decoder-Only Transformers: The Baseline
- Mixture of Experts: Selective Activation
- Recurrent and Hybrid Architectures
- Compact Language Models: Design for Small
- Architecture Choices That Matter for Embedded: Context Window, Heads, FFN Width
- Selecting a Model Family for Your Target
Chapter 5: Quantization Fundamentals
- Why Quantization Works: Redundancy in Floating-Point Weights
- Symmetric vs. Asymmetric Quantization: Scales and Zero-Points
- Per-Tensor vs. Per-Channel vs. Group-wise
- Calibration Strategies
- The Quantization-Error-Accuracy Tradeoff
- Quantization Formats: FP16, BF16, INT8, and Lower
Chapter 6: Weight-Only Quantization
- How Weight-Only Quantization Works
- Dequantization on the Fly: The Compute Cost
- Packing Schemes: 4-bit, 5-bit, 6-bit Weights
- Weight-Only Quantization and Activation Outliers
- Implementing Weight-Only Matrix Multiplication
- When Weight-Only Is Enough (and When It Is Not)
Chapter 7: Activation Quantization and Mixed Precision
- Why Activations Are Harder to Quantize
- Activation Outliers and Their Impact
- Mixed Precision Strategies
- Quantization-Aware Training
- Post-Training Quantization with Activation Calibration
- INT4/INT3 Activation Quantization: The Cutting Edge
Chapter 8: Sparsity, Pruning and Structured Compression
- Unstructured vs. Structured Sparsity
- Pruning Methods and Magnitude Thresholding
- Compressed Sparse Row and Other Formats
- The Problem of Sparsity on Small CPUs
- Block-wise and Channel-wise Pruning
- Low-Rank Factorization
- When to Use Pruning and Sparsity
Chapter 9: Knowledge Distillation and Architecture Modification
- Distillation: Making Small Models Smarter
- Architecture Surgery: Head Merging, Layer Thinning
- Vocabulary Reduction and Tokenizer Optimization
- Context Window Management and Compression
- Speculative Decoding: Drafting and Verifying
- When to Modify the Model vs. When to Optimize the Runtime
Chapter 10: KV Cache Engineering
- KV Cache Layout: Memory Mapping the Context
- KV Cache Quantization and Compression
- Eviction Policies and Sliding Windows
- Activation Checkpointing and Recomputation
- Streaming and Weight-Reuse Strategies
- KV Cache for Long Context on Tiny Memory
Chapter 11: CPU, SIMD and Vector Optimization
- Vector Units: NEON, SSE, AVX, and RISC-V V-Extensions
- Writing Vectorized Matmul Kernels
- Packing and Unpacking for SIMD
- Instruction-Level Parallelism and Loop Unrolling
- Memory Access Patterns and Cache Locality
- Tuning for Specific CPU Microarchitectures
Chapter 12: GPU, NPU and DSP Acceleration
- Mobile GPUs: Mali, Adreno, PowerVR
- Neural Processing Units: NPU Architectures and Constraints
- DSP Acceleration for Embedded Systems
- Tensor Cores and Matrix Multiply Accelerators
- Accelerator APIs and Vendor SDKs
- When Accelerators Help (and When They Get in the Way)
Chapter 13: Inference Runtimes and Software Architecture
- The Runtime Stack: From Model Files to Generated Tokens
- Graph Representations and IR Formats
- Lightweight Transformer Runtimes
- Vendor-Specific Runtimes vs. Portable Solutions
- Kernel Selection and Code Generation
- Designing Your Own Minimal Inference Runtime
Chapter 14: Model Conversion and Deployment Pipelines
- Exporting Models: ONNX, GGUF, safetensors, and Others
- Tokenizer Packaging and Conversion
- Quantization Pipelines: Tools and Techniques
- Cross-Compilation and Build Systems
- Binary Size, Flash Constraints, and OTA Updates
- Model Versioning and Update Strategies
Chapter 15: Performance Engineering and Benchmarking
- Metrics That Matter: TTFT, ITL, Throughput, and Energy
- Benchmarking Methodology: Warm-Up, Sampling, Reproducibility
- Profiling: Finding the Real Bottleneck
- Compute-Bound vs. Memory-Bound vs. Synchronization-Bound
- Compiler Optimization and Its Surprises
- Performance Regression and Change Management
Chapter 16: Power, Energy and Thermal Management
- Measuring Power and Energy in Real Deployments
- Energy per Token: A Practical Metric
- Dynamic Voltage and Frequency Scaling
- Thermal Throttling and Sustained Performance
- Power-Optimized Scheduling and Batching
Chapter 17: Debugging and Failure Analysis
- Incorrect Outputs: Numerical Error vs. Model Behavior
- Memory Exhaustion and Fragmentation
- Stack Overflow and Heap Corruption
- Quantization Bugs: Zero-Point and Scale Errors
- Hardware-Specific Issues: Alignment, Endianness, FPU Modes
- Silent Accuracy Degradation and Regression Testing
Chapter 18: Security and Privacy
- The Security Model of Local Inference
- Model Extraction and Protection
- Supply Chain Risks: Malicious Models and Tooling
- Input and Output Security: Prompt and Data Handling
- Secure OTA Updates and Integrity Verification
- Privacy Advantages and Their Limits
Chapter 19: Case Study : An LLM in 256 MB of RAM
- The Hardware and Its Constraints
- Choosing and Preparing the Model
- Memory Budget: Weights, KV Cache, Runtime Overhead
- Implementation Decisions and Trade-offs
- Benchmarking Results
- Lessons Learned
Chapter 20: Case Study : Offline Edge Assistant on an ARM SBC
- System Requirements and Architecture
- Model and Tokenizer Selection
- Building the Inference Engine
- Application Integration: Input, Output, Streaming
- Robustness: Error Handling, Watchdogs, Recovery
- Production Readiness Checklist
Chapter 21: Case Study : Pushing the Limits on MCU-Class Hardware
- MCU Constraints: Kilobytes of RAM, Megahertz Clocks
- Micro-Models: Can Transformers Run on an MCU?
- Extreme Quantization and Architecture Reduction
- Alternative Approaches: Recurrent and State-Space Models
- Measured Performance and Practical Viability
- The Real Limit: Where LLMs Cannot Go
Chapter 22: Production Deployment Patterns
- Firmware Integration and Boot-Time Constraints
- Multi-Model and Fallback Strategies
- Handling Concurrent Requests on Constrained Systems
- Observability: Logging, Metrics, Telemetry
- A/B Testing and Model Rollout at the Edge
- Maintenance and Lifecycle Management
Chapter 23: Future Directions
- Architectural Innovations: Mamba, RWKV, and State-Space Models
- Hardware Roadmaps: What Next-Gen Edge Silicon Will Bring
- Algorithmic Advances: Better Quantization, Sparsity, Distillation
- Compiler and Runtime Innovations
- The Economics of Edge Inference
- The Next Frontier: What Comes After Tiny LLMs
Conclusion: The Engineer’s Playbook
- The Decision Tree: Choosing Your Optimization Strategy
- A Checklist for Edge LLM Deployment
- Accepting Trade-offs: No Solution Is Free
- Final Thoughts: The Promise and Limits of Edge Intelligence