From Architecture to Optimization - A Practical Guide to Making Software Faster
Introduction: The Machine Beneath the Abstraction
- Why Performance Still Matters
- Why This Book Is Different
- How to Read This Book
- About the Repository Analysis
- What This Book Will Not Do
- Getting Started
- Chapter Summary
Chapter 1: Why Performance Matters - and Why It Is Hard
- The End of Free Lunch - Moore’s Law and the Performance Wall
- Latency vs. Throughput - Two Different Problems
- The Cost of Performance Mistakes - Real-World Economics
- Why Intuition Fails - Counterintuitive Behavior of Modern CPUs
- A Scientific Approach - The Performance Engineering Methodology
- What This Book Will and Will Not Cover
- Chapter Summary
Chapter 2: Inside the Modern CPU - A Mental Model
- The Von Neumann Architecture and Its Limitations
- The Fetch-Decode-Execute Cycle - And Why It Is Not That Simple
- Cores, Threads and SMT - Understanding Compute Resources
- The Processor as a Factory - Pipeline, Buffer and Queue Analogies
- Why Clock Speed Is Only Part of the Story
- Instruction Set Architectures - x86-64 and ARM64 Compared
- Chapter Summary
Chapter 3: Instructions and the Instruction Set
- What Is an Instruction - Anatomy and Encoding
- RISC vs. CISC - History, Philosophy and Modern Reality
- Micro-ops - How x86 Instructions Become Executable Work
- Instruction Latency vs. Instruction Throughput - Two Different Numbers
- Addressing Modes and Their Performance Cost
- Reading Agner Fog and Processor Reference Manuals
- Chapter Summary
Chapter 4: Pipelines and Superscalar Execution
- The Pipeline - Stages, Hazards and Throughput
- Superscalar Design - Issuing Multiple Instructions Per Cycle
- Execution Ports and Resource Constraints
- Pipeline Depth - Benefits and the Cost of Mispredictions
- Frontend vs. Backend Bottlenecks
- Architectural Implications - How to Write Code for Superscalar CPUs
- Chapter Summary
Chapter 5: Out-of-Order Execution and Instruction-Level Parallelism
- The Problem - Dependencies and Stall Cycles
- Out-of-Order Execution - Scoreboarding and Tomasulo’s Algorithm
- The Reorder Buffer and Commit Stage
- True Dependencies, Anti-Dependencies and Output Dependencies
- Instruction-Level Parallelism - Measuring and Exploiting It
- What You Can Control - Software Techniques to Increase ILP
- Chapter Summary
Chapter 6: Branch Prediction and Speculative Execution
- The Branch Problem - Why Control Flow Kills Throughput
- Branch Prediction Mechanisms - Static and Dynamic Predictors
- Speculative Execution - Running Before You Know
- Branch Misprediction Penalty - Quantifying the Cost
- Writing Predictable Code - Data-Dependent Branches and Conditional Moves
- Security Implications - Spectre, Meltdown and Their Microcode Fixes
- Chapter Summary
Chapter 7: Instruction Scheduling and Code Layout
- Instruction Fetch - Width, Alignment and Cache Lines
- Code Alignment and Its Impact on Pipeline Performance
- Software Pipelining - Manual Instruction Scheduling
- Compiler Scheduling - What Compilers Do and What They Cannot See
- Loop Layout and Branch Placement
- Reading and Optimizing Assembly Output
- Chapter Summary
Chapter 8: Caches - The Memory Hierarchy
- The Memory Wall - Why Main Memory Is Too Slow
- Cache Levels - L1, L2, L3 and Their Roles
- Cache Internals - Lines, Sets, Ways and Associativity
- Cache Access Types - Hits, Misses and Conflict Misses
- False Sharing - A Subtle Performance Killer
- Cache-Friendly Programming - Locality, Stride and Data Structures
- Chapter Summary
Chapter 9: Cache Coherency and Multi-Core Communication
- The Cache Coherency Problem - Multiple Copies, One Truth
- MESI and MOESI Protocols - How CPUs Keep Caches in Sync
- The Snooping Bus and Directory-Based Coherency
- Coherency Traffic - Measuring Its Impact
- Writing Cache-Coherency-Aware Code
- NUMA and Cache Coherency in Large Systems
- Chapter Summary
Chapter 10: Virtual Memory and Translation Lookaside Buffers
- Virtual Memory - Isolation, Abstraction and Performance Cost
- Page Tables and Page Faults - Walking the Tree
- TLBs - Purpose, Hierarchy and Miss Penalties
- TLB Shootdowns - The Multi-Core Cost
- Huge Pages - Reducing Translation Overhead
- Memory Mapping Strategies for Performance
- Chapter Summary
Chapter 11: Memory Access Patterns and Bandwidth
- Sequential vs. Random Access - The Performance Gap
- Spatial and Temporal Locality - Exploiting Hardware Prefetching
- Memory Bandwidth - Measuring and Maximizing It
- Prefetching - Hardware and Software Techniques
- Data Layout Transformations - AOS vs. SOA and Beyond
- Benchmarking Memory Performance - Real Numbers on Real Hardware
- Chapter Summary
Chapter 12: SIMD and Vectorization
- SIMD Fundamentals - Doing More Work Per Instruction
- x86 SIMD Families - SSE, AVX, AVX2, AVX-512
- ARM NEON and SVE - Vectorization on ARM
- Auto-Vectorization - When Compilers Succeed and Fail
- Manual Vectorization - Intrinsics and Assembly
- Common Pitfalls - Alignment, Reductions and Control Flow
- Chapter Summary
Chapter 13: Multithreading and Hardware Concurrency
- Hardware Multithreading - Simultaneous Multithreading (SMT)
- Software Threads - Mapping to Hardware
- Scaling Performance - Amdahl’s Law and Gustafson’s Law
- Thread Affinity and NUMA Awareness
- Load Balancing - Dynamic and Static Scheduling
- When Multithreading Hurts - Resource Contention and False Sharing
- Chapter Summary
Chapter 14: Synchronization and Atomic Operations
- Atomic Operations - The Hardware Foundation
- Memory Ordering and Barriers - Sequential Consistency vs. Weaker Models
- Locks - Mutexes, Spinlocks and Their Overhead
- Lock-Free and Wait-Free Algorithms - Possibilities and Limits
- Read-Write Locks and Fine-Grained Locking
- Lock Coarsening, Elimination and False Sharing
- Chapter Summary
Chapter 15: Performance Counters and Hardware Monitoring
- What Are Performance Counters - Hardware Events and Counters
- Intel PMU, AMD IBS and ARM PMU - Vendor-Specific Features
- Key Metrics - IPC, Cache Miss Rates, Branch Mispredictions
- Using perf - The Linux Performance Toolkit
- Statistical Sampling vs. Precise Events
- Designing Counter-Based Experiments
- Chapter Summary
Chapter 16: Profiling - Understanding What Your Code Actually Does
- Sampling vs. Instrumentation - Trade-Offs
- Flame Graphs - Visualizing Execution Profiles
- CPU Profiling - Finding Hot Paths
- Cache Profiling - Identifying Memory Bottlenecks
- Branch Profiling - Understanding Control Flow Performance
- Profiling in Production - Sampling, Tracing and Overhead
- Chapter Summary
Chapter 17: Benchmarking - Measuring What Matters
- Microbenchmarks vs. Macrobenchmarks - When to Use Each
- Experimental Design - Controls, Variables and Repetition
- Statistical Variation - Warmup, Cooling and Outliers
- Measuring Latency - Precision Timers and Statistical Analysis
- Measuring Throughput - Sustained Performance and Sustaining It
- Common Benchmarking Mistakes - And How to Avoid Them
- Chapter Summary
Chapter 18: Compiler Optimization - Working With the Compiler
- Compiler Optimization Pipeline - What Happens Under the Hood
- Optimization Levels - What -O2 and -O3 Actually Do
- Compiler Flags - Tuning for Your Target
- Reading Optimization Reports - -Rpass and -Rpass-missed
- LTO and PGO - Link-Time and Profile-Guided Optimization
- When the Compiler Fails - Inline Assembly, Intrinsics and Hand-Tuning
- Chapter Summary
Chapter 19: CPU Frequency, Power Management and Thermal Limits
- Turbo Boost and Dynamic Frequency Scaling - How It Works
- Power Limits - TDP, PL1, PL2 and Their Real-World Impact
- Thermal Throttling - When Heat Becomes the Bottleneck
- Measuring Real Frequency - Is Your CPU Running at Full Speed?
- Disabling Power Management for Benchmarks
- Performance per Watt - Optimizing for Efficiency, Not Just Speed
- Chapter Summary
Chapter 20: Operating System Interactions and System Configuration
- Scheduler Effects - Context Switches and Priority Inversion
- Interrupt Handling - Softirqs, Hardirqs and Latency
- CPU Isolation - Isolating Cores for Latency-Sensitive Workloads
- Kernel Configuration - Tuning for Performance
- cgroups and Resource Controls - Containment and QoS
- Virtualization Overhead - When Your VM Hurts Performance
- Chapter Summary
Chapter 21: Performance Analysis Methodologies - A Scientific Approach
- The Performance Engineering Workflow - Measure, Hypothesize, Optimize, Validate
- Establishing a Baseline - Reproducibility and Regression Prevention
- Formulating Hypotheses - From Symptoms to Root Causes
- The Bottleneck Taxonomy - CPU-Bound, Memory-Bound, I/O-Bound
- Prioritizing Optimizations - Impact vs. Effort
- Validating Results - Avoiding Premature Celebration
- Chapter Summary
Chapter 22: Case Studies in Performance Engineering
- Case Study 1 - Optimizing a Hash Table: Cache Locality and False Sharing
- Case Study 2 - Vectorizing Image Processing: From Naive to AVX-512
- Case Study 3 - Reducing Lock Contention in a High-Throughput Server
- Case Study 4 - NUMA-Aware Memory Allocation in a Database
- Lessons Learned - Common Patterns Across Cases
- Chapter Summary
Chapter 23: Critical Analysis of the CPU-Performance-Engineering Repository
- Repository Overview - Structure and Scope
- Strengths - What It Gets Right
- Weaknesses and Omissions - What Is Missing or Under-Explained
- Accuracy Assessment - Comparing to Primary Sources and Reference Manuals
- Outdated or Misleading Material - Specific Examples and Corrections
- How to Use This Repository - Recommendations and Caveats
- Chapter Summary
Conclusion: The Performance Engineering Mindset
- The Key Principles - A Condensed Checklist
- What Changes and What Stays the Same
- Future Trends - Chiplets, AI Accelerators and New Memory Technologies
- Continuing Your Journey - Resources and Communities
- Final Thoughts
