A Complete Guide to Virtual Memory, Allocators, Paging, and Performance in the Modern Linux Kernel*
Introduction
Chapter 1: Foundations of Memory Abstraction
- The Memory Problem: Why Programs Cannot Share Physical Memory Directly
- Virtual Memory as an Abstraction Layer
- Processes, Address Spaces, and Isolation
- The Role of the Operating System in Memory Management
- Hardware Support: MMUs and Protection
- The Linux Memory Management Philosophy
Chapter 2: Hardware Foundations and Address Translation
- Memory-Mapped I/O and Physical Addressing
- The Memory Management Unit Architecture
- Page Tables: Structure, Levels, and Walks
- Translation Lookaside Buffers and TLB Shootdowns
- x86-64 Paging: Four-Level and Five-Level Page Tables
- ARM64 and RISC-V Virtual Memory Systems
- Memory Types, Caching Attributes, and Access Permissions
- Walking a Page Table by Hand: A Practical Example
Chapter 3: Linux Address Spaces and the Virtual Memory Layout
- The Process Address Space in Linux
- User-Space Layout: Code, Data, Heap, Stack, and Mapped Regions
- Kernel Virtual Address Space and Direct Mapping
- Architecture-Specific Layouts: x86-64, ARM64, and RISC-V
- Address Space Layout Randomization and Security
- The mm_struct and vma_struct: Core Data Structures
- Memory Region Management with mmap, brk, and mremap
- Practical Experiment: Inspecting Your Process’s Memory Map
Chapter 4: Physical Memory Management and the Buddy Allocator
- Tracking Physical Pages: The page Struct and Memmap
- Memory Zones and Node Architecture
- The Buddy Allocator Algorithm and Implementation
- Free Page Lists and Order Management
- Allocation Paths: GFP Flags and Context Sensitivity
- Fragmentation Analysis and Mitigation Strategies
- NUMA-Aware Allocation and Policy
- Source Walkthrough: mm/page_alloc.c
Chapter 5: Kernel Memory Allocators: Slab, Slub, and SLOB
- Why Kernel Needs Specialized Allocators
- The Slab Allocator Design and History
- Slub: Simplification and Performance Improvements
- SLOB: Minimalist Allocation for Embedded Systems
- Per-CPU Caches and Lockless Fast Paths
- kmalloc, kcalloc, and the Kernel Allocation API
- Cache Colocation, False Sharing, and NUMA Effects
- Building a Custom Kernel Allocator Module
Chapter 6: User-Space Memory Allocation and vmalloc
- The malloc Contract and POSIX Requirements
- glibc ptmalloc: Arenas, Bins, and Thread Safety
- Alternative Allocators: jemalloc and tcmalloc
- mmap vs brk: Strategies for Large and Small Allocations
- vmalloc and the Kernel’s Virtual Memory Allocator
- vmap, ioremap, and Special Mapping Regions
- Performance Comparison: Benchmarking User-Space Allocators
- Practical Experiment: Tracing malloc Behavior with eBPF
Chapter 7: Demand Paging, Page Faults, and Copy-on-Write
- The Lazy Loading Philosophy: Demand Paging
- Page Fault Types and Handling Flow
- Major and Minor Faults: Performance Implications
- File Backed Pages and the Page Cache Introduction
- Copy-on-Write Semantics and Implementation
- Fork, Exec, and Memory Efficiency
- Shared Memory Regions and mmap SHARED vs PRIVATE
- Source Walkthrough: The Page Fault Handler in mm/memory.c
Chapter 8: The Page Cache and Buffer Head Architecture
- Why Caching File Data in RAM Transforms Performance
- Page Cache Data Structures: Radix Trees and XArrays
- Page State Machine and Lifecycle Management
- Readahead Algorithms and Prefetching Strategies
- Dirty Pages, Writeback, and Flush Daemons
- Buffer Heads, Extents, and Historical Evolution
- Interaction with Filesystems and Block Devices
- Practical Experiment: Measuring Page Cache Hit Rates
Chapter 9: Swapping, Reclaim, and Page Replacement Policies
- When RAM Runs Out: The Reclaim Problem
- Active and Inactive Page Lists
- Page Replacement Policies and LRU Approximation
- Swap Spaces, Swap Devices, and File-Based Swapping
- The Swap-In and Swap-Out Pathways
- zswap, z3fold, and Compressed Caching
- Thrashing Detection and Prevention
- Practical Experiment: Tuning vm.swappiness and Observing Behavior
Chapter 10: Memory Compaction, Defragmentation, and Huge Pages
- Fragmentation in Long-Running Systems
- The Compaction Algorithm: Migration and Scanning
- Direct Reclaim vs Kswapd vs Compact Daemons
- Huge Pages and TLB Efficiency
- Transparent Huge Pages: Automatic Promotion and Demotion
- hugetlbfs and Preallocated Huge Page Pools
- Performance Impact: Benchmarks and Trade-offs
- Practical Experiment: Enabling and Measuring THP Effects
Chapter 11: NUMA, Memory Policy, and Scalability
- The NUMA Problem: Distance Matters
- Node, Zone, and Memory Hierarchy in Linux
- Memory Policies: Default, Interleave, Bind, and Prefer
- Task Migration and Memory Placement Coordination
- Per-NUMA Node Allocators and Caches
- Scalability Challenges on Large Systems
- Performance Measurement and NUMA-Aware Programming
Chapter 12: Memory Cgroups, Controllers, and Resource Limits
- Containerization and the Need for Memory Isolation
- Cgroup Hierarchy and Memory Controller Architecture
- Memory Limits, Soft Limits, and Thresholds
- Accounting: Tracking Usage Per-Cgroup
- Reclaim Under Constraints: Hierarchical Pressure
- Swap Accounting and Memory+Swap Limits
- OOM Behavior Within Cgroups
- Practical Experiment: Setting Up Memory-Limited Containers
Chapter 13: Out-of-Memory Handling, Panic, and Recovery
- When Reclaim Fails: The Last Resort
- The OOM Killer Algorithm and Victim Selection
- Scoring Factors and Heuristics
- OOM Score Adjustment and Task Protection
- System-wide vs Cgroup-local OOM Events
- Kernel Panic Conditions and Memory Exhaustion Scenarios
- Recovery Strategies and Production Best Practices
- Real-World Case Studies: OOM Incidents and Post-Mortems
Chapter 14: Synchronization, Concurrency, and Locking in Memory Management
- Concurrency Challenges in a Shared Allocator
- Locking Hierarchy and Deadlock Prevention
- Page Table Locks and RCU Protection
- Zone Locks, PGDAT Locks, and Fine-Grained Contention Control
- Per-CPU Data Structures and Lockless Fast Paths
- Reference Counting and RCU in Page Lifecycle Management
- Memory Barriers, Ordering, and Cache Coherency
- Scalability Analysis: Contention on Large Systems
Chapter 15: Performance Optimization, Profiling, and Debugging
- Memory Performance Metrics That Matter
- Using /proc and /sys for Runtime Inspection
- Perf Tools: Flame Graphs and Memory Profiling
- eBPF and BCC: Custom Tracing Instruments
- ftrace and Kernel Function Tracing
- Kernel Module Development for Memory Experiments
- Crash Dump Analysis with crash and kdump
- Practical Experiment: Building a Complete Debugging Workflow
Chapter 16: Security Implications, Exploits, and Mitigations
- Memory Corruption and Kernel Exploitation
- Use-After-Free and Slab-Based Attacks
- Information Leaks Through Page Reuse
- ASLR, KASLR, and Address Randomization
- Stack Protection and Canary Values
- SMAP, SMEP, and Supervisor Mode Protections
- MDS, Spectre, Meltdown, and TLB Side Channels
- Kernel Hardening: CONFIG Options and Best Practices
Chapter 17: Architecture-Specific Implementations and Source Walkthroughs
- Architecture Abstraction Layers in the Kernel
- x86-64 Implementation Details and Optimizations
- ARM64 Memory Management and Attribute Indirection
- RISC-V Sv39, Sv48, and Sv57 Paging Modes
- TLB Shootdown Differences Across Architectures
- IOMMU and DMA Mapping Considerations
- Navigating the Kernel Source Tree: A Practical Guide
- Building and Booting a Custom Kernel for Experimentation
Conclusion: Synthesis and Future Directions
- Key Design Principles
- Emerging Trends
- Final Perspective