The Leanpub 60 Day 100% Happiness Guarantee
Within 60 days of purchase you can get a 100% refund on any Leanpub purchase, in two clicks.
See full terms...
Lab Guide — GPU Infrastructure, Model Serving and Operations
Somewhere there is a server with two GPUs for which no operational process exists. This book turns that into an operable platform — hands-on, in 23 labs: vLLM, KServe, LiteLLM, the NVIDIA GPU Operator, Keycloak, OpenBao, ArgoCD, pgvector. Not a tutorial: a reference work that shows the derivations behind every setting.
Minimum price
$25.00
$39.00
About the Book
Somewhere there is a server with two GPUs for which no operational process exists.
Developers have long been using ChatGPT and Claude, often through private accounts. Business units have built proofs of concept that nobody operates. The data protection officer has questions. Procurement has seen an invoice that nobody can attribute to a team.
Whoever is to turn that into an operable platform needs different knowledge from the data scientist and different knowledge from the application developer: GPUs in the cluster, the mechanics of inference servers, access control on models, cost attribution — and the question of which data may reach which model.
HANDS-ON: 23 LABS
The book builds the platform piece by piece on a reference lab you can rebuild: one NVIDIA GPU with 24 GiB, passed through to a virtual machine, RKE2 across several nodes. Every lab names its prerequisite and its expected output — so you can tell whether it worked, and how you notice when it did not.
One lab needs more: the tensor parallelism measurement in Chapter 9 rents a two-GPU instance by the hour, one to three euros for the whole series. Where the hardware runs out, the book says so instead of pretending.
WHAT YOU OPERATE IN THE LABS
vLLM — batching, prefix caching, chunked prefill, preemption and the KV cache made visible through the scheduler's own metrics; engine arguments set from a calculation instead of by guessing.
KServe — ServingRuntime and InferenceService, declarative model serving, canary rollouts, autoscaling on the queue, scale-to-zero.
LiteLLM as the AI gateway — virtual keys, budgets, rate limits, model allowlists bound to the data class, fallbacks, spend logs per team.
NVIDIA GPU Operator, driver, Container Toolkit, DCGM Exporter — from PCI passthrough to a GPU the scheduler can allocate; MIG and time-slicing, including the memory isolation test that shows why one of them is unsuitable for tenants.
Keycloak and OpenBao — OIDC claims down to team and data class, and provider keys that no developer ever sees.
Prometheus and Grafana — the metrics that carry meaning for inference, and the ones that mislead.
ArgoCD and Harbor — moving a 20 GB model artefact through a GitOps chain, with a rollback that takes minutes.
pgvector — a retrieval path with permission filtering inside the query, and erasure you can demonstrate.
Also covered: Ray, Qdrant, Milvus, PostgreSQL, Redis, External Secrets, Proxmox VE, RKE2.
NOT A TUTORIAL
The labs are there to make things measurable, not to be followed once and used up. The aim is not to explain tools — tools age. The aim is to make the derivations available: why the KV cache and not the weights is the capacity bottleneck, why 100 percent GPU utilization can mean nothing at all, why self-hosting pays off above a certain utilization and not above a certain number of users.
Every chapter carries a version stamp naming what it was written and checked against. Chapters 2 to 16 each end with a troubleshooting table; every chapter ends with interview questions at senior level, with what separates a weak answer from a convincing one.
From the contents: I LLM Inference Fundamentals · II GPU Infrastructure · III Model Serving · IV The Enterprise Layer · V Operations · VI Synthesis
Assumes Kubernetes, GitOps, identity providers and monitoring in production. This is not a Kubernetes book and not a machine learning book — it starts where those leave off.
386 pages · 17 chapters · 4 appendices · 23 labs · PDF and EPUB
About the Author
Thomas Zachmann has spent more than twenty years building systems, and the last several of them building Kubernetes platforms for organisations where security and auditability are the requirement rather than the aspiration: on-premises, on bare metal, in air-gapped networks, and in European sovereign clouds — under European regulatory and audit regimes.
His work is the hardening of container platforms and everything that has to hold up when an auditor asks: policy enforcement with Kyverno, identity and access with Keycloak, supply-chain evidence with Harbor, Trivy and Dependency Track — and secrets management with HashiCorp Vault, which he has been deploying and operating since 2017. He works in the public clouds as well, which is what makes the on-premises argument in these books a comparison rather than a preference.
He is a Certified Kubernetes Security Specialist (CKS), Certified Kubernetes Administrator (CKA) and Certified Kubernetes Application Developer (CKAD). Those twenty years began in software engineering — performance and systems programming at SAP in C and C++, database performance analysis at IBM, and Go since. That background is why the automation in his platforms is written rather than bought, and why his books ask you to type rather than to read.
He is based in Hamburg and works across the DACH region.
Click the buttons to get the free sample in PDF or EPUB, or read the sample online here
Within 60 days of purchase you can get a 100% refund on any Leanpub purchase, in two clicks.
See full terms...
We pay 80% royalties on purchases of $7.99 or more, and 80% royalties minus a 50 cent flat fee on purchases between $0.99 and $7.98. You earn $8 on a $10 sale, and $16 on a $20 sale. So, if we sell 5000 non-refunded copies of your book for $20, you'll earn $80,000.
(Yes, some authors have already earned much more than that on Leanpub.)
In fact, authors have earned over $15 million writing, publishing and selling on Leanpub.
Learn more about writing on Leanpub
If you buy a Leanpub book, you get free updates for as long as the author updates the book! Many authors use Leanpub to publish their books in-progress, while they are writing them. All readers get free updates, regardless of when they bought the book or how much they paid (including free).
Most Leanpub books are available in PDF (for computers) and EPUB (for phones, tablets and Kindle). The formats that a book includes are shown at the top right corner of this page.
Finally, Leanpub books don't have any DRM copy-protection nonsense, so you can easily read them on any supported device.
Learn more about Leanpub's ebook formats and where to read them
You can use Leanpub to easily write, publish and sell in-progress and completed ebooks and online courses!
Leanpub is a powerful platform for serious authors, combining a simple, elegant writing and publishing workflow with a store focused on selling in-progress ebooks.
Leanpub is a magical typewriter for authors: just write in plain text, and to publish your ebook, just click a button. (Or, if you are producing your ebook your own way, you can even upload your own PDF and/or EPUB files and then publish with one click!) It really is that easy.