Leanpub Header

Skip to main content

Writing GPU Kernels with CUDA Rust

A Practical Guide to High-Performance GPU Computing from Rust

Writing GPU Kernels with CUDA Rust
This book is 100% completeLast updated on 2026-09-17

Unlock NVIDIA GPU performance from Rust. This practical guide takes you from GPU architecture and memory hierarchies to writing and optimizing real CUDA Rust kernels. Explore both official Rust CUDA approaches with runnable examples and learn how to build fast, production-ready GPU code without giving up Rust’s safety and clarity.

Minimum price

$25.00

$35.00

You pay

Author earns

$

Also available for 1 book credit with a Reader Membership

PDF
EPUB
WEB
APP
210
Pages
About

About

About the Book

GPU parallel computing has fundamentally changed how software achieves performance, but writing kernels has historically demanded deep expertise in low-level hardware details and C++ era programming patterns. This book shows Rust developers how to harness NVIDIA GPUs using CUDA Rust, the emerging ecosystem that brings memory safety, expressive type systems, and modern tooling to GPU programming.

You will learn both of NVIDIA's official CUDA Rust tracks: the SIMT-style cuda-oxide compiler backend for fine-grained thread-level control, and the tile-based cutile-rs framework for safe, high-level kernel authoring. Along the way, you will build a solid understanding of GPU architecture, memory hierarchies, synchronization, and performance optimization. Every chapter includes complete, runnable code examples and explains the why behind each concept, not just the how. This is a practical guide for Rust programmers ready to write production-quality GPU kernels without sacrificing the safety and clarity that make Rust compelling.

Author

About the Author

Steve Publications

Steve is a technology professional with more than 20 years of experience in software development, server infrastructure, cybersecurity, vulnerability research and reverse engineering. Throughout his career, he has designed, secured, analyzed and tested complex software and infrastructure, with a particular focus on understanding how systems fail and how they can be made more secure.

Outside of work, Steve enjoys sharing knowledge with the technology community. He collaborates with researchers, industry experts and technology professionals to write practical books covering software development, cybersecurity, cloud computing, networking, DevOps, artificial intelligence and enterprise technologies. His books focus on practical learning through clear explanations, real-world examples and hands-on exercises. With more than two decades of industry experience, his goal is to help IT professionals, students and technology enthusiasts build useful skills and stay current in a rapidly changing industry.

We believe readers deserve to know how our books are created. Most of our authors are not native English speakers, so we use AI to help translate, proofread manuscripts, fix grammar, improve sentence structure and make technical explanations easier to read. AI is used as an editing tool only. It does not replace the research, technical knowledge or hands-on experience behind our books. Some of our authors also prefer to remain anonymous for privacy or professional reasons. In those cases, we publish their work under a different name. The author's name may be different, but the quality of the content and our review process remain the same.

Every book is written, reviewed and maintained by experienced technology professionals, with contributions from our private technical community of more than 420 engineers and researchers. We spend far more time validating technical accuracy and keeping our content up to date than generating text. We are always interested in working with experienced professionals who have deep expertise in a particular technology or domain. If you would like to publish a book with us or help review an existing manuscript, we'd love to hear from you. Send us a message describing your area of expertise. We are especially interested in niche technologies, specialized skills and emerging topics that are underrepresented in existing technical literature.

If you look through the contents of our books, you'll see practical examples, detailed explanations and material that is regularly updated. Our goal is to publish books that professionals can actually rely on, not low-effort AI-generated content. If you ever feel that one of our books does not meet that standard, Leanpub offers a 60-day money-back guarantee. Feel free to request a refund if you are not satisfied with your purchase.

Contents

Table of Contents

A Practical Guide to High-Performance GPU Computing from Rust

Introduction

Chapter 1: Why GPU Computing, Why Rust

  1. The Parallel Computing Revolution
  2. Where GPUs Excel and Where They Do Not
  3. The State of GPU Programming in 2025
  4. Why Rust for GPU Development
  5. How This Book Is Organized

Chapter 2: GPU Architecture and Execution Model

  1. From CPU to GPU: A Hardware Perspective
  2. Streaming Multiprocessors and Execution Units
  3. The Thread Block and Grid Execution Model
  4. Warp Execution and SIMT
  5. Memory Hierarchies on the GPU
  6. How a Kernel Launch Traverses the Hardware

Chapter 3: CUDA Fundamentals: The C++ Reference Model

  1. The CUDA C++ Kernel Syntax
  2. Host Code Versus Device Code
  3. Launch Configuration and Execution Dimensions
  4. Device Memory Allocation and Data Transfer
  5. A Complete CUDA C++ Example Walked Through
  6. What CUDA Rust Will Build Upon

Chapter 4: Introducing CUDA Rust: History and Vision

  1. The Road to GPU Programming in Rust
  2. NVIDIA’s CUDA Rust Initiative
  3. The Two Tracks: Native and C Bindings
  4. Design Goals and Philosophy
  5. Current Status, Stability, and Versioning
  6. What CUDA Rust Is Not

Chapter 5: Toolchain Setup and Project Configuration

  1. Prerequisites: Hardware, CUDA Toolkit, and Drivers
  2. Installing the CUDA Rust Compiler and Dependencies
  3. Project Layout with cargo-cuda
  4. Build Configuration and Environment Variables
  5. Your First CUDA Rust Project
  6. Troubleshooting Common Setup Issues

Chapter 6: The CUDA-Rust API Surface

  1. The core_cuda and cuda-rs Crates
  2. Device Pointer Types and Abstractions
  3. The Kernel Trait and Launch Mechanics
  4. Memory Allocation and Transfer APIs
  5. Utility Types: Dimensions, Streams, Events
  6. How the API Maps to CUDA C++ Concepts

Chapter 7: Device Code Fundamentals

  1. Marking Code as Device Code
  2. Computing Thread and Block Indices
  3. Your First Kernel: Vector Addition
  4. Understanding the Launch Dimensions
  5. Compilation: How Device Code Becomes PTX and SASS
  6. Safety Considerations in Device Code

Chapter 8: Memory Management: Allocation and Transfer

  1. Device Memory Allocation Patterns
  2. Host-to-Device and Device-to-Host Transfers
  3. Pinned Memory and Asynchronous Transfers
  4. The Cost Model of Data Movement
  5. Memory Safety at Host and Device Boundaries
  6. Complete Example: Matrix Initialization and Transfer

Chapter 9: Kernel Launch Patterns

  1. Launch Configuration and Grid Size Calculations
  2. Occupancy and Heuristics for Block Size
  3. Streams and Concurrent Execution
  4. Synchronous Versus Asynchronous Launches
  5. Kernel Chaining and Dependencies
  6. Example: Multi-Kernel Pipeline

Chapter 10: Shared Memory and Block-Level Optimization

  1. Shared Memory: Purpose and Performance
  2. Declaring and Using Shared Memory in Kernels
  3. The Matrix Transpose Optimization
  4. Warp Coalescing and Memory Access Patterns
  5. A Reduction Kernel with Shared Memory
  6. Barriers and Synchronization Within a Block

Chapter 11: Global Memory and Data Access Patterns

  1. Global Memory Architecture and Caching
  2. Coalesced Versus Uncoalesced Access
  3. Bank Conflicts and Shared Memory Layout
  4. Texture and Constant Memory
  5. Designing Kernels for Memory-Bound Workloads
  6. Example: Image Filter Kernel

Chapter 12: Synchronization and Coordination

  1. Thread-Level Synchronization: __syncthreads
  2. Warp-Level Primitives
  3. Cross-Block Synchronization Limitations
  4. Stream-Based Synchronization
  5. Host-Device Synchronization and Events
  6. Deadlock and Starvation Pitfalls

Chapter 13: Ownership, Lifetimes, and GPU Safety

  1. Ownership Semantics for Device Resources
  2. Lifetimes at the Host-Device Boundary
  3. RAII for CUDA Handles
  4. Aliasing and the Device Pointer Model
  5. Avoiding Use-After-Free on the GPU
  6. Safe Abstractions Over Unsafe CUDA Operations

Chapter 14: Error Handling and Robust GPU Code

  1. CUDA Error Types and Categories
  2. Propagating Errors in Host Code
  3. Detecting Kernel Launch Errors
  4. Runtime Versus Compile-Time Errors
  5. Defensive Patterns for Production Kernels
  6. Logging and Diagnostic Infrastructure

Chapter 15: Numerical Computing on the GPU

  1. Float32, Float64, and Mixed Precision
  2. GPU Floating Point Semantics
  3. Denormals, NaN, and Infinity
  4. Atomic Operations and Race Conditions
  5. Numerical Stability in Parallel Algorithms

Chapter 16: Interoperability: CUDA Rust and the Wider World

  1. Calling CUDA C++ Libraries from Rust
  2. FFI Patterns and Safety
  3. Interfacing with cuBLAS and Linear Algebra
  4. cuFFT and Transform Libraries
  5. CUDA Graphs and Advanced APIs
  6. Building Hybrid CUDA C++ and CUDA Rust Projects

Chapter 17: Integration with Rust Application Ecosystems

  1. Multi-Threaded Hosts and CUDA Contexts
  2. CUDA with Rust Async Runtimes
  3. CPU-GPU Hybrid Algorithms
  4. Designing a GPU-Accelerated Rust Library
  5. Example: Async GPU Processing Pipeline

Chapter 18: Debugging GPU Kernels

  1. Common GPU Kernel Bugs
  2. Using cuda-memcheck and Sanitizers
  3. Debug Builds and GDB for CUDA
  4. Print-Based and Trace-Based Debugging
  5. Debugging Deadlocks and Timing Issues
  6. Systematic Approaches to Bug Isolation

Chapter 19: Profiling and Performance Analysis

  1. NVIDIA Nsight Systems and Compute Profilers
  2. Occupancy Metrics and Utilization
  3. Memory Bandwidth Analysis
  4. Identifying Kernel Bottlenecks
  5. Roofline Model and Performance Reasoning
  6. Iterative Optimization Case Study

Chapter 20: Performance Optimization Patterns

  1. Register Pressure and Spilling
  2. Instruction-Level Parallelism
  3. Loop Unrolling and Predicate Elimination
  4. Warp Divergence and Control Flow
  5. Using Specialized Instructions
  6. Example: Optimizing a Convolution Kernel

Chapter 21: Production Deployment and Best Practices

  1. Testing GPU Kernels
  2. Continuous Integration for CUDA Rust
  3. API Design for Reusable Kernels
  4. Documentation and Team Conventions
  5. Portability and Hardware Targeting
  6. When Not to Use CUDA Rust

Conclusion

  1. The CUDA Rust Journey So Far
  2. Where the Ecosystem Is Heading
  3. Complementary Technologies
  4. A Final Word on GPU Computing with Rust

References

Get the free sample chapters

Click the buttons to get the free sample in PDF or EPUB, or read the sample online here

The Leanpub 60 Day 100% Happiness Guarantee

Within 60 days of purchase you can get a 100% refund on any Leanpub purchase, in two clicks.

See full terms...

Earn $8 on a $10 Purchase, and $16 on a $20 Purchase

We pay 80% royalties on purchases of $7.99 or more, and 80% royalties minus a 50 cent flat fee on purchases between $0.99 and $7.98. You earn $8 on a $10 sale, and $16 on a $20 sale. So, if we sell 5000 non-refunded copies of your book for $20, you'll earn $80,000.

(Yes, some authors have already earned much more than that on Leanpub.)

In fact, authors have earned over $15 million writing, publishing and selling on Leanpub.

Learn more about writing on Leanpub

Free Updates. DRM Free.

If you buy a Leanpub book, you get free updates for as long as the author updates the book! Many authors use Leanpub to publish their books in-progress, while they are writing them. All readers get free updates, regardless of when they bought the book or how much they paid (including free).

Most Leanpub books are available in PDF (for computers) and EPUB (for phones, tablets and Kindle). The formats that a book includes are shown at the top right corner of this page.

Finally, Leanpub books don't have any DRM copy-protection nonsense, so you can easily read them on any supported device.

Learn more about Leanpub's ebook formats and where to read them

Write and Publish on Leanpub

You can use Leanpub to easily write, publish and sell in-progress and completed ebooks and online courses!

Leanpub is a powerful platform for serious authors, combining a simple, elegant writing and publishing workflow with a store focused on selling in-progress ebooks.

Leanpub is a magical typewriter for authors: just write in plain text, and to publish your ebook, just click a button. (Or, if you are producing your ebook your own way, you can even upload your own PDF and/or EPUB files and then publish with one click!) It really is that easy.

Learn more about writing on Leanpub