Building Real DRM Drivers for the Modern Linux Kernel
Chapter 1: Before We Begin
- Understanding What a GPU Driver Is
- Prerequisites and Assumptions
- Choosing a Kernel Version
- Setting Up the Build Environment
- Testing in a Virtual Environment
- The Tools You Will Use
- How to Read This Book
Chapter 2: GPU Architecture Fundamentals
- From CPU to GPU: Why Graphics Need Special Hardware
- The GPU Command Processor
- Ring Buffers and Submission Queues
- GPU Memory Hierarchy
- Execution Units
- Display Engines
- The Role of the Driver in the Control Chain
- Hardware Register Interfaces
- What Comes Next
Chapter 3: The Linux Graphics Stack
- A Brief History: From Framebuffer to DRM
- The Modern Stack: Layer by Layer
- The Complete Path: A Drawing Call End to End
- Userspace vs Kernel Responsibilities
- Render Nodes vs Card Nodes
- Mesa and the Kernel: A Partnership
- What Comes Next
Chapter 4: Direct Rendering Manager
- What DRM Is and Why It Exists
- The drm_device Structure
- The drm_driver Structure
- Device Registration Flow
- Device Management with devm
- DRM File Operations
- DRM Minor Devices
- DRM Capabilities
- Debugging Support
- What Comes Next
Chapter 5: Kernel Mode Setting
- From UMS to KMS
- Display Pipeline Objects
- How Objects Are Connected
- Display Modes
- EDID and DDC
- Atomic Mode Setting
- The ww_mutex Locking Scheme
- Vblank and Timing
- Standard Properties
- What Comes Next
Chapter 6: Graphics Execution Manager
- What GEM Is and What It Solves
- The drm_gem_object Structure
- GEM Object Operations
- GEM Initialization
- Creating GEM Objects
- GEM Handles
- GEM mmap
- GEM Prime and Buffer Sharing
- GEM CMA Helper
- What Comes Next
Chapter 7: Translation Table Maps
- What TTM Is and Why It Exists
- TTM Memory Types
- The ttm_bo Object
- TTM Driver Interface
- TTM Initialization
- TTM and GEM Integration
- When to Use TTM vs GEM
- What Comes Next
Chapter 8: PCI Integration and Device Discovery
- PCI Enumeration: How Devices Are Found
- The PCI Device Driver Model
- PCI Configuration Space
- PCI BARs: Memory-Mapped Registers
- DMA and DMA Masks
- PCI Interrupts: MSI and MSI-X
- Device Tree for Embedded GPUs
- What Comes Next
Chapter 9: Building the Minimal DRM Driver
- Directory Structure and Kernel Module Layout
- Building the Driver
- Loading and Verifying
- Unloading the Module
- Troubleshooting
- Summary
Chapter 10: Device Registration and Driver Lifecycle
- Module Init and Exit
- Probe: Discovering and Initializing Hardware
- Probe Error Handling
- Remove: Cleaning Up
- Runtime Power Management
- System Sleep: Suspend and Resume
- Debugging Registration Failures
- Summary
Chapter 11: GPU Memory Management
- Allocating GEM Objects from Kernel Space
- DMA Allocation: Coherent vs Streaming
- Userspace GEM Handles and Ioctls
- Mapping GEM Objects
- GEM Object Lifecycle and Reference Counting
- Mapping GEM Objects into GPU Address Space
- Implementing a Minimal GEM Memory Manager
- Testing the Memory Manager
- Summary
Chapter 12: Command Submission
- Command Streams: Structure and Semantics
- Ring Buffers and Circular Queues
- The DRM Command Submission Ioctl
- Parsing and Validating Command Buffers
- Hardware Submission: Writing to GPU Registers
- Error Paths: Invalid Commands and Buffer Overruns
- Summary
Chapter 13: GPU Virtual Memory
- Why GPUs Use Virtual Memory
- GPU Address Spaces and Contexts
- Page Table Structures
- The DRM GPUVA Manager
- Mapping and Unmapping Buffers
- Handling GPU Page Faults
- Summary
Chapter 14: Synchronization and Fences
- Why Synchronization Is Harder with GPUs
- DRM Fences: The Fundamental Synchronization Object
- Software Fences vs Hardware Fences
- Fence Contexts and Timelines
- DMA Fences and Cross-Device Synchronization
- Implementing Fence Signaling from Interrupts
- Userspace Fence Waiting
- Summary
Chapter 15: Interrupts and GPU Completion
- GPU Interrupts: Types and Sources
- MSI vs Legacy INTx Interrupts
- Requesting and Sharing IRQs
- The Interrupt Handler Structure
- Signaling Fences and Wakeups from Interrupt Context
- Debugging Interrupt Issues
- Summary
Chapter 16: Display Pipeline and Framebuffers
- Framebuffer Registration and Pixel Formats
- Implementing CRTC Enable/Disable and Mode Set
- Plane Composition and Layering
- Connector Detection and EDID Reading
- Atomic Commit for the Display Pipeline
- Scanout and Vblank Timing
- Summary
Chapter 17: The Userspace Interface
- DRM Ioctls: Standard vs Driver-Private
- Designing a Clean Ioctl Interface
- Versioning and Backward Compatibility
- The DRM File Private Data Structure
- Userspace Workflow: Open, Get Capabilities, Allocate, Submit, Wait
- libdrm Integration and uAPI Stability
- Mesa Driver Hooks: Gallium, VK, and Driver Layers
- Summary
Chapter 18: Power Management
- Runtime Power Management: Suspend and Resume
- System Sleep: Freeze, Thaw, Poweroff, Restore
- GPU Power States and DVFS
- Clock Gating and Power Gating
- Idle Detection and Autosuspend
- Resuming GPU State After Sleep
- Power Debugging and Measurement
- Summary
Chapter 19: Debugging GPU Drivers
- Kernel Log Analysis: dmesg, printk Levels
- Dynamic Debug: Enabling Per-File Logging
- ftrace and Function Graph Tracing
- DRM-Specific Debugfs Entries
- DRM Tracepoints for GPU Operations
- Using QEMU and Virtual Hardware for Testing
- GDB with Kernel Modules
- Crash Dumps and Stack Traces
- Summary
Chapter 20: Concurrency and Safety
- Locking Primitives: Mutex, Spinlock, Rwlock
- DRM’s Reservation Locking for Shared Objects
- Lock Ordering and Deadlocks
- Reference Counting: Kref, Drm_Ref, Custom
- Memory Ordering: Barriers and Ordering Guarantees
- Race Conditions Specific to GPU Drivers
- RCU for Read-Mostly Data Structures
- Summary
Chapter 21: GPU Hangs and Error Recovery
- What a GPU Hang Is and What Causes It
- Detecting Hangs: Watchdogs and Timeouts
- The DRM Reset Infrastructure
- Reset Sequence: Stopping Submission, Resetting, Resuming
- Recovering GPU State: Contexts, Page Tables, Buffers
- Handling Stale Fences and Waiting Processes
- GPU Reset and Userspace Notification
- Fault Injection for Testing
- Summary
Chapter 22: Performance Optimization
- Profiling GPU Drivers: Where Time Is Spent
- Reducing CPU Overhead in the Fast Path
- Batching and Coalescing Command Submissions
- Prefetching and Cache-Friendly Data Structures
- Avoiding Page Faults in GPU Memory
- Interrupt Coalescing and Rate Limiting
- Measuring and Comparing Performance
- Summary
Chapter 23: Security Considerations
- The GPU Driver as an Attack Surface
- IOMMU and DMA Protection
- Command Stream Validation
- GPU Sandboxing and VM Isolation
- Privilege Escalation via the Driver
- Information Disclosure through GPU Memory
- Secure Boot and Signed Drivers
- Best Practices for Security-Hardened Drivers
- Summary
Chapter 24: Reading Upstream DRM Drivers
- Choosing Drivers to Study: Simple vs Complex Examples
- Anatomy of a Modern DRM Driver: File Layout
- Understanding DRM Core Abstractions in Practice
- How to Read and Understand Complex Initialization Sequences
- Identifying Patterns: Common Idioms Across Drivers
- Using Coccinelle and Other Tools to Explore Code
- Summary
Chapter 25: Upstream Development
- The Linux Graphics Mailing Lists and Maintainers
- Finding a Maintainer for Your Driver or Subsystem
- Patch Series Structure and Cover Letters
- Writing Commit Messages for Graphics Drivers
- Code Review Expectations and Common Feedback
- The Review Cycle: What to Expect
- Handling Objections and Iteration
- Tools: B4, Git, Patchwork
- Summary
Chapter 26: Conclusion
- The Complete Driver: From Nothing to Functional
- What We Learned and Why It Matters
- Emerging Trends: Vulkan Drivers, GPU Compute, Async Scheduling
- Continuing to Learn: Where to Go Next
- Final Encouragement