Leanpub Header

Skip to main content

Local Intelligence

Running Large Language Models on Your MacBook with Apple Silicon

This book is 100% completeLast updated on 2026-07-13

Local Intelligence shows you how to run large language models entirely on your Mac with Apple Silicon. Learn to use tools like Ollama, MLX, and llama.cpp, understand quantization, and build real local AI applications with open-source code.

Minimum price

$19.00

$29.00

You pay

Author earns

$

Also available for 1 book credit with a Reader Membership

PDF
EPUB
116
Pages
About

About

About the Book

This book is a complete, practical guide to running large language models entirely on your own Mac. If you have an M1, M2, or M3 MacBook and want to deploy open-source LLMs locally for privacy, cost savings, or independence from cloud APIs, this book takes you from zero to mastery. You will learn the hardware architecture that makes Apple Silicon uniquely suited for this work, master every major inference framework (llama.cpp, Ollama, MLX, and Hugging Face Transformers), understand quantization strategies that fit billion-parameter models into your unified memory, and build real applications with local AI. Every technique described is fully open-source, reproducible on macOS, and grounded in real benchmarks and working code.

Bundles

Bundles that include this book

Author

About the Author

Steve Publications

Steve is a technology professional with more than 20 years of experience in software development, server infrastructure, cybersecurity, vulnerability research and reverse engineering. Throughout his career, he has designed, secured, analyzed and tested complex software and infrastructure, with a particular focus on understanding how systems fail and how they can be made more secure.

Outside of work, Steve enjoys sharing knowledge with the technology community. He collaborates with researchers, industry experts and technology professionals to write practical books covering software development, cybersecurity, cloud computing, networking, DevOps, artificial intelligence and enterprise technologies. His books focus on practical learning through clear explanations, real-world examples and hands-on exercises. With more than two decades of industry experience, his goal is to help IT professionals, students and technology enthusiasts build useful skills and stay current in a rapidly changing industry.

We believe readers deserve to know how our books are created. Most of our authors are not native English speakers, so we use AI to help translate, proofread manuscripts, fix grammar, improve sentence structure and make technical explanations easier to read. AI is used as an editing tool only. It does not replace the research, technical knowledge or hands-on experience behind our books. Some of our authors also prefer to remain anonymous for privacy or professional reasons. In those cases, we publish their work under a different name. The author's name may be different, but the quality of the content and our review process remain the same.

Every book is written, reviewed and maintained by experienced technology professionals, with contributions from our private technical community of more than 400 engineers and researchers from Ukraine, Belarus and Russia. We spend far more time validating technical accuracy and keeping our content up to date than generating text. We are always interested in working with experienced professionals who have deep expertise in a particular technology or domain. If you would like to publish a book with us or help review an existing manuscript, we'd love to hear from you. Send us a message describing your area of expertise. We are especially interested in niche technologies, specialized skills and emerging topics that are underrepresented in existing technical literature.

If you look through the contents of our books, you'll see practical examples, detailed explanations and material that is regularly updated. Our goal is to publish books that professionals can actually rely on, not low-effort AI-generated content. If you ever feel that one of our books does not meet that standard, Leanpub offers a 60-day money-back guarantee. Feel free to request a refund if you are not satisfied with your purchase.

Contents

Table of Contents

Running Large Language Models on Your MacBook with Apple Silicon

Introduction

  1. The Problem with Cloud APIs
  2. Why Apple Silicon Is Different
  3. What This Book Covers
  4. Who This Book Is For
  5. How to Use This Book

Chapter 1: The Local LLM Revolution

  1. Why Local? Privacy, Cost, and Independence
  2. The Cloud API Trap: Latency, Censorship, and Hidden Costs
  3. Apple Silicon Changes Everything
  4. What You Can Actually Run on a MacBook Today
  5. How This Book Is Structured

Chapter 2: Apple Silicon Architecture for ML Workloads

  1. Unified Memory Architecture Explained
  2. CPU, GPU, and Neural Engine: Who Does What?
  3. M1 vs M2 vs M3: The Progression for AI Workloads
  4. Thermal Design and Sustained Performance
  5. Benchmarking Your Mac’s ML Capability

Chapter 3: Model Selection – Finding the Right LLM

  1. The Open-Source LLM Landscape in 2025 and Beyond
  2. Model Families: Llama, Mistral, Qwen, Gemma, and Others
  3. Parameter Count vs Performance: The Sweet Spot for MacBooks
  4. GGUF Format and the Quantization Zoo
  5. Building Your Personal Model Library

Chapter 4: llama.cpp – The Foundation of Local Inference

  1. Installing llama.cpp on macOS
  2. Loading Your First GGUF Model
  3. Understanding Quantization: Q4_K_M, Q5_K_S, Q8_0, and Beyond
  4. Server Mode: REST API and Multi-User Access
  5. Advanced Features: LoRA Adapters, Embeddings, and Multimodal

Chapter 5: Ollama – Simplicity at Scale

  1. Installing Ollama on macOS
  2. The Library: Pulling and Managing Models
  3. Modelfiles: Customizing System Prompts and Parameters
  4. The Ollama API: Integrating into Applications
  5. Creating and Publishing Your Own Models

Chapter 6: MLX – Apple’s Native Machine Learning Framework

  1. What Is MLX and Why It Matters
  2. Installing and Setting Up the MLX Ecosystem
  3. Running Models with mlx-lm
  4. Fine-Tuning Models on Your Mac
  5. MLX vs llama.cpp vs Ollama: When to Use What

Chapter 7: Hugging Face Transformers on Apple Silicon

  1. The Hugging Face Ecosystem on macOS
  2. Bitsandbytes and 4-Bit/8-Bit Quantization
  3. Hugging Face Optimum for Apple Silicon
  4. PEFT: LoRA, QLoRA, and Parameter-Efficient Tuning
  5. Building a Complete Local Training Pipeline

Chapter 8: Performance Optimization and Tuning

  1. Memory Management: Unified Memory as Both Blessing and Constraint
  2. Context Window Optimization and KV Cache Strategies
  3. Batch Size, Prompt Processing, and Token Generation Speed
  4. Thermal Throttling: Keeping Your Mac Cool Under Load
  5. Profiling Tools and Performance Metrics

Chapter 9: Building Local Applications with LLMs

  1. The Local Application Stack
  2. Chat Interfaces and Web UIs: Open WebUI, Text Generation WebUI
  3. RAG Pipelines: Local Retrieval-Augmented Generation
  4. Tool Use and Function Calling with Local Models
  5. Building a Production-Grade Local AI Assistant

Chapter 10: Fine-Tuning and Customization at Home

  1. When to Fine-Tune vs When to Use Prompt Engineering
  2. Data Preparation for Local Fine-Tuning
  3. QLoRA Training on Apple Silicon with MLX
  4. Evaluating Your Fine-Tuned Model
  5. Deploying Custom Models in Production

Chapter 11: The Broader Ecosystem – Tools and Integrations

  1. Model Serving: vLLM, Text Generation Inference, and Local Alternatives
  2. Evaluation Frameworks: Measuring Your Model’s Quality
  3. Embeddings Locally: Sentence Transformers and Beyond
  4. Multimodal Models on Apple Silicon
  5. CI/CD for Local LLM Deployments

Chapter 12: Troubleshooting and Real-World Gotchas

  1. Out of Memory: Diagnosing and Solving RAM Exhaustion
  2. Slow Inference: Identifying Bottlenecks
  3. Model Compatibility and Format Conversion
  4. macOS-Specific Issues and Kernel Panics
  5. Recovery Strategies and Safe Shutdown

Conclusion: The Future of Personal AI

  1. What We Have Learned
  2. The Trajectory: Smaller Models, Bigger Chips
  3. The Philosophical Shift: From Cloud Dependency to Personal Sovereignty
  4. Your Next Steps

References

Get the free sample chapters

Click the buttons to get the free sample in PDF or EPUB, or read the sample online here

The Leanpub 60 Day 100% Happiness Guarantee

Within 60 days of purchase you can get a 100% refund on any Leanpub purchase, in two clicks.

See full terms...

Earn $8 on a $10 Purchase, and $16 on a $20 Purchase

We pay 80% royalties on purchases of $7.99 or more, and 80% royalties minus a 50 cent flat fee on purchases between $0.99 and $7.98. You earn $8 on a $10 sale, and $16 on a $20 sale. So, if we sell 5000 non-refunded copies of your book for $20, you'll earn $80,000.

(Yes, some authors have already earned much more than that on Leanpub.)

In fact, authors have earned over $15 million writing, publishing and selling on Leanpub.

Learn more about writing on Leanpub

Free Updates. DRM Free.

If you buy a Leanpub book, you get free updates for as long as the author updates the book! Many authors use Leanpub to publish their books in-progress, while they are writing them. All readers get free updates, regardless of when they bought the book or how much they paid (including free).

Most Leanpub books are available in PDF (for computers) and EPUB (for phones, tablets and Kindle). The formats that a book includes are shown at the top right corner of this page.

Finally, Leanpub books don't have any DRM copy-protection nonsense, so you can easily read them on any supported device.

Learn more about Leanpub's ebook formats and where to read them

Write and Publish on Leanpub

You can use Leanpub to easily write, publish and sell in-progress and completed ebooks and online courses!

Leanpub is a powerful platform for serious authors, combining a simple, elegant writing and publishing workflow with a store focused on selling in-progress ebooks.

Leanpub is a magical typewriter for authors: just write in plain text, and to publish your ebook, just click a button. (Or, if you are producing your ebook your own way, you can even upload your own PDF and/or EPUB files and then publish with one click!) It really is that easy.

Learn more about writing on Leanpub