Leanpub Header

Skip to main content

Apache CloudStack — AI Factory Edition

Building and running a flexible, small-scale yet powerful local AI factory on commodity enterprise hardware

This book is 60% completeLast updated on 2026-08-07

A model you rent can be repriced, throttled or revoked. A model you own cannot. Build a sovereign local AI factory on commodity enterprise hardware: Apache CloudStack, frontier-class open models in 768GB of RAM, GPU passthrough, RAG, fine-tuning, and a multi-agent support desk — with nothing leaving the building. You don't need to be a hyperscaler. Grow your own, cook your own.

Minimum price

$19.99

$34.99

You pay

Author earns

$
PDF
EPUB
WEB
APP
Discussion Forum
About

About

About the Book

There is a particular kind of letter that changes how an engineer thinks. For some it is a licensing renewal. Lately, for a growing number of us, it has been the notice that a model you had built into your product is being withdrawn, repriced, or placed behind an export control you cannot satisfy. A model you rent can be revoked by someone else's decision. A model you own cannot.

This book builds, on hardware you can buy today and rack this week, a local AI factory: a small, owned, sovereign platform that runs frontier-class open-weight models on your own iron, serves them to multiple users and tenants, and puts them to work earning their keep. No per-token meter. No revocation risk. Nothing leaves the building.

The platform is Apache CloudStack bent to a new purpose, across four tiers of mostly refurbished commodity iron:

  • The Flex — the chapter that sells the book: two Dell R7525 servers redeployed from saved templates into a four-node inference cluster, a two-node aggregated cluster, or four isolated single-tenant boxes. Three machines from one pair of servers; the templates double as the backups. A sovereignty story you can hand to an auditor
  • The quality brain — a frontier-class open model (roughly 744 billion parameters, mixture-of-experts) held in 768GB of system RAM via llama.cpp, with its honest tokens-per-second envelope on real DDR4
  • An ARM model farm on an Ampere workstation, and a 100Gb RDMA inference fabric reality-checked rather than oversold — cross-host pairs, not a mythical clean mesh
  • The brains put to work — retrieval over your own documents, fine-tuning small models with LoRA/QLoRA on a single 16GB consumer GPU, an n8n multi-agent support desk that triages, answers, guardrails and escalates, and an AI-first web front-end built by your own local coding model
  • The regulated-firm case made concrete — multi-tenant isolation, data governance, and operating the whole estate by the Mars test: if this node were on Mars, could you still recover it?

The reader I have most in mind is the engineer inside an organisation whose data legally cannot leave its walls — finance, legal, healthcare, government. But it is equally for anyone who has watched rented infrastructure become a leash and decided to take it back. You do not need to be an AI researcher; you need a Linux command line, a lab to break things in, and the discipline to measure rather than assume.

The book's contract is honesty: every figure is measured, dated, labelled a vendor's claim, or flagged VERIFY for your own build. You do not need to be a hyperscaler to own a capable AI factory — and you should not pretend to be one.

Author

About the Author

Michael Hinsley

Mike Hinsley has spent more than three decades with his hands on real infrastructure — from electronics and engineering through operations, security, cloud and SaaS — and has never lost the habit of wanting to understand the machine all the way down. He is the founder of a UK infrastructure consultancy that helps organisations escape per‑core licensing lock‑in by migrating from VMware/Broadcom to sovereign, self‑owned platforms built on Apache CloudStack, with documented client savings of up to 94%. His guiding principle is one he calls GYOCYO — Grow Your Own, Cook Your Own: own the means of a capability rather than rent it as a finished product. He lives it literally. From a smallholding in rural Cheshire he runs large‑scale aquaponics, IoT‑monitored beehives, and a 28.8 kW solar array backed by 90 kWh of battery storage — the same instrument‑everything, owe‑nothing‑to‑anyone thinking he brings to enterprise clients, proven first on his own land. His HIVE‑DC concept — a data centre in a beehive — grew directly out of that overlap. Mike previously presented "Aquaponics, Apiculture, and Advanced Networking with Apache CloudStack" at the CloudStack European User Group in London. He publishes under the Leaf Spine Books imprint, and is preparing to establish a sustainable‑technology education centre uniting aquaponics, beekeeping and cloud infrastructure under one roof.

Contents

Table of Contents

Introduction — Read This First

  1. What this book is
  2. Who it is for
  3. The thesis: you do not need to be a hyperscaler
  4. How to use this book
  5. The shape of the journey
  6. A method, not a product

Apache CloudStack — The Sovereign AI Factory

  1. On owning what you run

The Book in Full — Annotated Contents

  1. Part I — Build
  2. Part II — The Flex
  3. Part III — The Brains
  4. Part IV — Make It Pay
  5. Front and back matter
  6. The expanded edition

Chapter 1 — Why Own an AI Factory

  1. The Story: the second letter
  2. The thesis: you do not have to be a hyperscaler
  3. What this factory is — and what it is not
  4. The strongest reason: data that cannot leave the building
  5. The live hook: the week a model could be pulled
  6. GYOCYO: Grow Your Own, Cook Your Own
  7. Return on Research
  8. The Mars heuristic
  9. The map of the journey ahead
  10. The honesty contract

Chapter 2 — The Build of Materials

  1. The Story: counting what the building already holds
  2. How to read this bill of materials
  3. Tier 1 — The responsive tier: two Dell R7525s
  4. Tier 2 — The quality brain: two Dell R740XDs, 768GB each
  5. Tier 3 — The ARM model farm: one Ampere Ultra
  6. The graphics-card pool, counted honestly
  7. Tier 4 — The management plane: an R660, or a pair of Raspberry Pis
  8. Where these four tiers live: three clusters in one zone
  9. Proving the iron: a hardware-readiness pass
  10. What is not on the headline list — but you still need
  11. The shape of the cost — without a figure

Chapter 3 — Racking, Power and the Off-Grid Interlock

  1. The Story: four of a kind
  2. The rack as a machine with a shape
  3. Power is not a wall socket
  4. Measure the draw: a power-and-thermal readiness pass
  5. The factory that parks itself
  6. The off-grid interlock
  7. Every watt on purpose

Chapter 4 — Network Fabric: the Quad-25Gb Front-End, the Back-to-Back 100Gb, and Where CloudStack Ends

  1. The Story: the question on the wall
  2. Two fabrics, one estate
  3. The quad-25Gb front-end: the network CloudStack governs
  4. The back-to-back 100Gb: short, switchless, and honest about its shape
  5. Where CloudStack ends
  6. Bring up the fabric: a network-readiness pass
  7. Every wire on purpose

Chapter 5 — The Management Plane: the Conductor of the Estate, Built Two Honest Ways

  1. The Story: the conductor they already had
  2. What the conductor actually holds
  3. The caveat, stated plainly: one node is no HA
  4. Option A — the datacentre default: a Dell R660
  5. Option B — the low-power showcase: two Raspberry Pi 5s
  6. The HA aside: keepalived, HAProxy, and a replicated database
  7. A note on management traffic, and a warning about the keys
  8. Bring up the conductor — a management-node readiness pass
  9. The Mars test, applied to the conductor
  10. Built, not yet installed

Chapter 6 — CloudStack on the R7525s: KVM, and the GPU and NIC Made Schedulable

  1. The Story: the hosts that waited
  2. Installing against the conductor
  3. KVM on the R7525s
  4. A zone, a pod, a cluster
  5. Making the iron schedulable
  6. Bring the platform up — an install-and-passthrough readiness pass
  7. The Mars test, applied to the install
  8. An orchestra, ready to be told what to play

Chapter 7 — The Flex: Profiles A, B and C as Templates

  1. The Story: the cluster that left
  2. The template is the unit of the flex
  3. Switching is destroy-and-redeploy, not live-migrate
  4. Profile A — the four-node vLLM cluster
  5. Profile B — the two-node aggregated pair
  6. Profile C — four independent tenants
  7. Placement is your job, not the scheduler’s
  8. The flex as code
  9. Build the flex — a profiles-and-templates readiness pass
  10. Three machines, one rack

Chapter 8 — The RDMA Inference Fabric: RoCEv2 Reality-Checked, vLLM Across the Cluster

  1. The Story: the number nobody would say
  2. The link is RDMA, not merely a fast port
  3. The switchless pair is the easy case — say so
  4. The driver question: in-box first
  5. GPUDirect RDMA: assume it is not there
  6. vLLM across the pair: what distribution actually buys
  7. The measured reality: a method, not a number
  8. Build the fabric — an RDMA-and-vLLM readiness pass
  9. The fast wire, reckoned with

Chapter 9 — The Quality Brain: GLM-5.2 (and Kimi) on the 768GB R740XDs via llama.cpp

  1. The Story: the quiet join
  2. The tier, and why it is a separate machine
  3. Why this model: ownership you cannot have revoked
  4. What GLM-5.2 actually is
  5. The engine: llama.cpp, and the honest size problem
  6. How llama.cpp spends the one GPU
  7. Speed, told straight
  8. The line-up: one brain, or two, and which ones
  9. What this tier is for — and what it is emphatically not
  10. Settled, and VERIFY-on-live
  11. Build the brain — a quality-brain readiness pass
  12. The brain that answers the hard ones

Chapter 10 — The ARM Model Farm: Ampere Ultra, CPU-or-GPU

  1. The Story: the odd one in
  2. The tier nobody else covers
  3. Why ARM, on purpose
  4. The multi-architecture zone, made real
  5. CPU or GPU: the two ways this box serves
  6. The engines that run clean on arm64
  7. What this tier is for, and what it is not
  8. Settled, and VERIFY-on-live
  9. Build the farm — an ARM-farm readiness pass
  10. The third tier, in its place

Chapter 11 — Corpus and RAG: Retrieval Over Your Own Documents

  1. The Story: the one who asked what the data said
  2. Two ways to make a model know something
  3. The pipeline, end to end
  4. Ingest and chunk: cutting the corpus to size
  5. Embed: the second, smaller model
  6. Store: Qdrant, the vector index
  7. Retrieve and rerank: similarity is not relevance
  8. Ground: which brain answers, and from what
  9. What RAG is, and what it is not
  10. The sovereign half
  11. Build the pipeline — a RAG readiness pass
  12. Settled, and VERIFY-on-live
  13. The half that knows your documents
  14. The Story: a promise with a number waiting

Afterword - The estate is yours now

  1. About the author
  2. Stay in touch

Glossary

Appendix A — Bill of Materials

  1. The four tiers at a glance
  2. Tier 1 — Responsive / flex compute: Dell PowerEdge R7525 (×2)
  3. Tier 2 — Quality brain: Dell PowerEdge R740XD (×2)
  4. Tier 3 — ARM model farm: Ampere Ultra workstation (×1)
  5. Tier 4 — Management: two documented options
  6. The RDMA fabric
  7. On power and cooling

Get the free sample chapters

Click the buttons to get the free sample in PDF or EPUB, or read the sample online here

The Leanpub 60 Day 100% Happiness Guarantee

Within 60 days of purchase you can get a 100% refund on any Leanpub purchase, in two clicks.

See full terms...

Earn $8 on a $10 Purchase, and $16 on a $20 Purchase

We pay 80% royalties on purchases of $7.99 or more, and 80% royalties minus a 50 cent flat fee on purchases between $0.99 and $7.98. You earn $8 on a $10 sale, and $16 on a $20 sale. So, if we sell 5000 non-refunded copies of your book for $20, you'll earn $80,000.

(Yes, some authors have already earned much more than that on Leanpub.)

In fact, authors have earned over $15 million writing, publishing and selling on Leanpub.

Learn more about writing on Leanpub

Free Updates. DRM Free.

If you buy a Leanpub book, you get free updates for as long as the author updates the book! Many authors use Leanpub to publish their books in-progress, while they are writing them. All readers get free updates, regardless of when they bought the book or how much they paid (including free).

Most Leanpub books are available in PDF (for computers) and EPUB (for phones, tablets and Kindle). The formats that a book includes are shown at the top right corner of this page.

Finally, Leanpub books don't have any DRM copy-protection nonsense, so you can easily read them on any supported device.

Learn more about Leanpub's ebook formats and where to read them

Write and Publish on Leanpub

You can use Leanpub to easily write, publish and sell in-progress and completed ebooks and online courses!

Leanpub is a powerful platform for serious authors, combining a simple, elegant writing and publishing workflow with a store focused on selling in-progress ebooks.

Leanpub is a magical typewriter for authors: just write in plain text, and to publish your ebook, just click a button. (Or, if you are producing your ebook your own way, you can even upload your own PDF and/or EPUB files and then publish with one click!) It really is that easy.

Learn more about writing on Leanpub