Building fast LLM inference is about far more than turning the right knobs. This book takes you inside vLLM and the systems behind it, showing how memory, scheduling, GPUs and distributed serving shape real-world performance. Learn how to measure what matters, find bottlenecks and build inference infrastructure that scales without wasting money.
Go beyond basic Linux tuning and learn how performance really works under the hood. Explore kernel internals, hardware behavior, benchmarking and production optimization with practical tools, real trade-offs and copy-paste-ready examples. Build the skills to diagnose bottlenecks and solve performance problems with confidence.
What happens when AI starts making the decisions behind your database queries? This practical guide explores how modern optimizers work, then shows how machine learning and reinforcement learning can improve them. Learn to build, test and deploy AI-driven optimization systems that deliver real performance gains.
Go beyond writing code that works. Learn what really happens inside .NET, from C# and the JIT to memory, CPU execution and the operating system. This practical guide shows experienced developers how to measure performance, find real bottlenecks and make targeted improvements that hold up in production.
What really makes a compression system fast, efficient and effective? Using bzip3 as a practical case study, this book breaks down the building blocks behind modern compression, from data transforms and match finding to entropy coding and parallel processing. Learn how to design, implement and tune compression systems for real-world performance.
PostgreSQL Under Load shows what really happens when production gets busy. Learn how connections, transactions, MVCC, locks, WAL and autovacuum interact under pressure, then use that knowledge to diagnose incidents, tune PostgreSQL and build systems that stay reliable when concurrency spikes.
Build a real LLM inference engine from the ground up in portable C. Learn how transformers work, load SafeTensors weights, implement KV caching and optimize inference on the CPU. By the end, you will run TinyLlama locally and generate text with a complete working runtime.
What does it take to run an LLM on hardware that barely has enough memory to breathe? This practical guide takes you from transformer math and model compression to inference runtimes, benchmarking and real-world deployment, with hands-on case studies for building fully offline LLM systems on tiny devices.
What if you could build an AI that writes just like you? In Building Your LLM Twin, you'll follow a hands-on journey from raw writing data to a fully deployed AI that captures your unique voice. Learn to build, fine-tune and scale real-world LLM systems with practical code, proven techniques and the tools to take your project from concept to production.
Kill N+1 queries, read EXPLAIN-friendly EF patterns, and ship faster ASP.NET APIs with a dense EF Core performance cookbook.
Master the ideas behind Java concurrency, from threads, race conditions, and synchronization to CompletableFuture, virtual threads, and structured concurrency. With practical patterns, runnable examples, and a companion GitHub repository, this book helps you move from understanding concurrency to applying it in real programs.
DuckDB is changing how developers think about analytics. This practical guide takes you from its architecture and SQL capabilities to performance tuning, cloud storage and production deployments. Learn how DuckDB works, where it shines and how to build fast, flexible analytical systems around it.
Python is entering a new performance era. Explore free-threaded execution, Rust extensions and modern tools for building faster, more scalable Python systems. Learn how CPython is changing, when Rust makes sense and how to turn these ideas into production-ready software.
Why is your system slow, and how do you know what to fix? This practical guide takes you from CPU and memory behavior to databases, networks and distributed systems. Learn how to profile, benchmark and diagnose real performance problems with working examples across modern languages and tools.