CUDA Programming from Scratch
- From First Principles to Production-Grade GPU Applications
Introduction: Why GPU Computing Matters Today
Chapter 1: The GPU Revolution — Architecture and History
- From Graphics to General-Purpose Computing
- GPU vs CPU: Divergent Design Philosophies
- The CUDA Platform Ecosystem
- GPU Architecture Roadmap: Fermi through Blackwell
Chapter 2: The CUDA Programming Model — Threads, Blocks, and Warps
- SIMT Execution: Single Instruction, Multiple Threads
- Thread Hierarchy: Threads, Warps, Blocks, Grids, and Clusters
- Kernel Launch Syntax and Configuration
- The Grid-Stride Loop Pattern
- Occupancy: Theory, Calculation, and the Occupancy API
Chapter 3: Memory Models — The GPU Memory Hierarchy
- Registers and Local Memory
- Global Memory and Coalesced Access
- Shared Memory: Scope, Latency, and Bank Conflicts
- Constant and Texture Memory
- L1/L2 Cache Architecture
- Memory Alignment and Vectorized Access
Chapter 4: Writing and Optimizing Kernels
- Your First CUDA Kernels: Vector Addition, Matrix Multiply
- Tiling and Shared Memory Optimization
- Avoiding Bank Conflicts: Padding and Swizzling
- Warp-Specialized Kernels and Producer-Consumer Patterns
- Common Pitfalls: Divergence, Race Conditions, Out-of-Bounds
Chapter 5: Synchronization — From Warps to Grids
- Warp-Level Synchronization (Implicit and Explicit)
- Block-Level Barriers (syncthreads)
- Cooperative Groups: Thread Block Tiles, Cluster Groups, Grid Groups
- Scoped Atomics and Thread Scopes
- Asynchronous Barriers and cuda::barrier (Hopper+)
Chapter 6: Warp-Level and Intrinsics Programming
- Warp Shuffle Primitives: __shfl_sync, __shfl_down_sync, __shfl_up_sync, __shfl_xor_sync
- Vote and Mask Operations: __ballot_sync, __any_sync, __all_sync, __activemask
- Warp-Level Reductions and Scans
- Inline PTX Assembly for Performance-Critical Code
- When to Use Warp Primitives vs. Cooperative Groups
Chapter 7: Asynchronous Execution — Streams, Events, and Overlap
- CUDA Streams: Default and User-Created
- Events for Synchronization and Timing
- Overlapping Data Transfers with Computation
- Multi-Stream Pipelining Patterns
- CUDA Graphs: Capture, Replay, and Constant-Time Launch
Chapter 8: Unified Memory and Advanced Memory Management
- The Problem with Explicit Host-Device Transfers
- cudaMallocManaged and Page Migration Engine
- Pinned (Page-Locked) Memory
- Unified Memory Performance: When It Works, When It Doesn’t
Chapter 9: Tensor Cores and Mixed-Precision Computing
- Evolution of Tensor Cores: Volta through Blackwell
- Matrix Multiply-Accumulate (MMA) Operations
- Data Precision Formats
- Writing Tensor Core Kernels with WGMMA
- The Transformer Engine and Dynamic Scaling
Chapter 10: Hopper Innovations — TMA, Barriers, and Pipelines
- Tensor Memory Accelerator (TMA): Architecture and Programming Model
- cuda::memcpy_async and Asynchronous Data Copies
- CUDA Pipelines: Producer-Consumer Patterns with Multi-Buffering
- Warp-Specialized Kernels for Maximum Utilization
- Cluster-Sized Thread Blocks and GPC-Level Scheduling
Chapter 11: CUDA Libraries — Building on NVIDIA’s Foundation
- Linear Algebra: cuBLAS, cuSOLVER, cuSPARSE
- Signal Processing: cuFFT
- Parallel Primitives: CUB and Thrust
- Random Numbers: cuRAND
- Image and Video: NPP, nvJPEG, nvCodec
- When to Use Libraries vs. Custom Kernels
Chapter 12: Dynamic Parallelism and Multi-GPU Programming
- Dynamic Parallelism: Child Kernels, Nested Launches
- Multi-GPU Architecture: PCIe vs NVLink
- Peer-to-Peer Memory Access (GPUDirect P2P)
- NCCL: Collective Communication Primitives
- Multi-GPU Design Patterns and Scaling Considerations
Chapter 13: Profiling, Debugging, and Performance Engineering
- CUDA-GDB and Nsight Debugger
- Compute Sanitizer: memcheck, racecheck, synccheck, initcheck
- Nsight Systems: System-Wide Profiling
- Nsight Compute: Kernel-Level Metrics and Analysis
- The APOD Framework (Assess, Parallelize, Optimize, Deploy)
- Performance Engineering Case Studies
Chapter 14: Real-World Applications — AI, HPC, and Scientific Computing
- Convolutional Kernels for Image Processing
- Matrix Multiplication at Scale: From Naive to Tensor Core Optimized
- Sparse Linear Algebra for Scientific Computing
- AI Training and Inference Pipelines
- Mini-Project: GPU-Accelerated Particle Simulation
