Leanpub Header

Skip to main content

Surviving Distributed Systems: A Practical Field Guide

A Practical Field Guide to Resilience, Concurrency, and Fault Tolerance

Surviving Distributed Systems: A Practical Field Guide
This book is 100% completeLast updated on 2026-10-06

Happy-path architecture is a myth. When downstream dependencies degrade, naive retries trigger thundering herds, thread pools exhaust in seconds, and dual-writes silently corrupt financial balances. Surviving Distributed Systems is an unvarnished, production-tested field guide for senior backend engineers. Packed with architectural diagrams and modern Java 21 code, this book teaches you how to design, isolate, and harden high-concurrency systems so they survive catastrophic cascading failures.

Minimum price

$19.00

$29.00

You pay

Author earns

$

Also available for 1 book credit with a Reader Membership

Buying multiple copies for your team? See below for a discount!

About

About

About the Book

Distributed systems do not fail gracefully by default. They collapse under the weight of their own recovery mechanisms.

Every distributed architecture looks clean on a whiteboard. But production is hostile: downstream dependencies spike with 2-second tail latencies, network partitions sever transaction boundaries, and Kubernetes nodes vanish mid-flight.

When degradation strikes, standard patterns often become weapons against your own infrastructure. A naive 3-attempt retry loop transforms a minor 20% downstream failure into an instantaneous 300% load amplification. An unchecked queue turns into a fatal out-of-memory crash. A dual-write without transactional integrity leaves databases and event brokers permanently out of sync.

Surviving Distributed Systems: A Practical Field Guide bypasses academic abstractions to dissect how modern high-throughput microservices break under load—and how to engineer them to survive.

Through mathematical models, failure trace analysis, and battle-tested Java 21 implementations, this book breaks down the core mechanics of production resilience:

  • Cascading Failures & Traffic Amplification: Model queue behavior using Little’s Law and neutralize thundering herds with full jitter backoff curves.
  • Deterministic Concurrency & State Integrity: Prevent lost updates with fencing tokens, manage high-contention locking, and eliminate out-of-order state regression.
  • Resilient Ingestion & Isolation: Move beyond static rate limiting with domain-aware load shedding, reactive backpressure, bulkheads, and distributed deadline propagation.
  • Consistency Boundaries: Expose the fatal flaws of naive dual-writes and implement production-grade idempotent consumers with atomic conflict resolution.

No hand-waving, no trivial "Hello World" toys, and zero filler. Every pattern is paired with executable code and real-world trade-offs designed for mission-critical enterprise backends.

Equip your services to survive the cascade.

Team Discounts

Team Discounts

Get a team discount on this book!

  • Up to 3 members

    Minimum price
    $48.00
    Suggested price
    $74.00
  • Up to 5 members

    Minimum price
    $76.00
    Suggested price
    $116
  • Up to 10 members

    Minimum price
    $142
    Suggested price
    $217
  • Up to 25 members

    Minimum price
    $318
    Suggested price
    $485
  • Up to 50 members

    Minimum price
    $570
    Suggested price
    $870

Author

About the Author

Artur Ostapyshyn

Artur Ostapyshyn is a Senior Backend Engineer specializing in high-throughput distributed systems, fault-tolerant architectures, and concurrent services within modern Java ecosystems.

He spends his days designing resilient transactional backends, wrestling with distributed state consistency, and tuning services to survive unpredictable production failures. His engineering philosophy centers on pragmatism over buzzwords: building systems that are simple to reason about, mathematically sound under contention, and resilient enough to keep engineers asleep through on-call shifts.

When he isn’t designing concurrency boundaries or debugging distributed race conditions, he enjoys diving into system internals, sharing real-world failure patterns, and writing practical engineering guides.

Contents

Table of Contents

ABOUT THIS BOOK

  1. Who This Book Is For
  2. The Four Core Tenets of Resilience
  3. How Each Chapter Is Structured

INTRODUCTION

  1. The Myth of the Clean Architecture
  2. The Physics of Distributed Boundaries
  3. The Three States of Distributed Execution
  4. The Fallacy Spectrum
  5. How to Approach This Book

CHAPTER 1: The Fallacy of a Reliable Network: Timeouts, Retries, and Cascading Storms

  1. 1.1 The Low-Level Anatomy of Connection and Socket Timeouts
  2. 1.2 The Anatomy of Thread Starvation: How Outbound Slowness Kills Upstream Ingestion
  3. 1.3 Retry Storms: The Mathematics of Self-Inflicted Denial of Service
  4. 1.4 The Mathematics of Recovery: Exponential Backoff and Full Jitter
  5. 1.6 Bulkheading: Isolating Blast Radius
  6. 1.7 Distributed Deadlines and Time-Budget Propagation
  7. 1.8 Architectural Blueprint: The Resilient Outbound Invoker
  8. 1.9 Trade-offs: The Cost of Transport-Layer Resilience
  9. CHAPTER 1: PRODUCTION CHECKLIST

CHAPTER 2: Delivery Guarantees and the Myth of “Exactly-Once”

  1. 2.1 The Distributed Reality: Why You Only Get At-Least-Once
  2. 2.2 The Dual-Write Problem: Why Database Commits and Broker Publishes Diverge
  3. 2.3 The Transactional Outbox Pattern: Atomic Messaging Without Distributed Locks
  4. 2.4 Designing Truly Idempotent Consumers
  5. 2.5 Dead Letter Queues (DLQ) and Poison Pill Management
  6. 2.6 Out-of-Order Message Processing: The Replicated Sequence Vector
  7. 2.7 Trade-offs: The Architectural Tax of Asynchronous Reliability
  8. CHAPTER 2: PRODUCTION CHECKLIST

CHAPTER 3: Caching at Scale: Prevention, Invalidation, and System Collapse

  1. 3.1 Cache Topologies and Write Policies: The Consistency Spectrum
  2. 3.2 The Anatomy of Cache Disasters
  3. 3.3 Mitigating the Collapse: Probabilistic Early Expiration (XFetch)
  4. 3.4 Request Coalescing: The Singleflight Mutex
  5. 3.5 Java Reference Implementation: Resilient Cache Coordinator
  6. 3.6 Defending Against Cache Penetration: Bloom Filters and Negative Caching
  7. 3.7 The Multi-Master Replication Trap: Stale Read Races
  8. 3.8 Trade-offs: The Price of In-Memory Speed
  9. CHAPTER 3: PRODUCTION CHECKLIST

CHAPTER 4: Concurrency and Distributed State: Locking Without Deadlocks

  1. 4.1 The Lost Update Anomaly: Read-Modify-Write in Stateless Runtimes
  2. 4.2 Optimistic Concurrency Control (OCC): Mechanics and the Contention Cliff
  3. 4.3 Pessimistic Concurrency Control (PCC): Lock Contention and Deadlocks
  4. 4.4 Distributed Locks and the “Redlock” Fallacy: Stop-the-World and Clock Drift
  5. 4.5 The Ultimate Guard: Fencing Tokens
  6. 4.6 Java Reference Implementations
  7. 4.7 Trade-offs: The Spectrum of Concurrency Guarantees
  8. CHAPTER 4: PRODUCTION CHECKLIST

CHAPTER 5: Self-Preservation: Backpressure, Rate Limiting, and Graceful Degradation

  1. 5.1 The Unbounded Buffer Trap and the Linux OOM Killer
  2. 5.2 The Mechanics of Backpressure: Push vs. Pull Topologies
  3. 5.3 Distributed Rate Limiting: Algorithms and Topologies
  4. 5.4 Load Shedding: Why Dropping Traffic Is the Only Way to Survive
  5. 5.5 Graceful Degradation: Tiered Degradation Architectures
  6. 5.6 Java Reference Implementations
  7. 5.7 Trade-offs: The Compromises of Self-Preservation
  8. CHAPTER 5: PRODUCTION CHECKLIST

CONCLUSION: The Architecture of Failure

APPENDIX A: The Ultimate Production Readiness Checklist

APPENDIX B: Recommended Reading & Seminal Papers

Get the free sample chapters

Click the buttons to get the free sample in PDF or EPUB, or read the sample online here

The Leanpub 60 Day 100% Happiness Guarantee

Within 60 days of purchase you can get a 100% refund on any Leanpub purchase, in two clicks.

See full terms...

Earn $8 on a $10 Purchase, and $16 on a $20 Purchase

We pay 80% royalties on purchases of $7.99 or more, and 80% royalties minus a 50 cent flat fee on purchases between $0.99 and $7.98. You earn $8 on a $10 sale, and $16 on a $20 sale. So, if we sell 5000 non-refunded copies of your book for $20, you'll earn $80,000.

(Yes, some authors have already earned much more than that on Leanpub.)

In fact, authors have earned over $15 million writing, publishing and selling on Leanpub.

Learn more about writing on Leanpub

Free Updates. DRM Free.

If you buy a Leanpub book, you get free updates for as long as the author updates the book! Many authors use Leanpub to publish their books in-progress, while they are writing them. All readers get free updates, regardless of when they bought the book or how much they paid (including free).

Most Leanpub books are available in PDF (for computers) and EPUB (for phones, tablets and Kindle). The formats that a book includes are shown at the top right corner of this page.

Finally, Leanpub books don't have any DRM copy-protection nonsense, so you can easily read them on any supported device.

Learn more about Leanpub's ebook formats and where to read them

Write and Publish on Leanpub

You can use Leanpub to easily write, publish and sell in-progress and completed ebooks and online courses!

Leanpub is a powerful platform for serious authors, combining a simple, elegant writing and publishing workflow with a store focused on selling in-progress ebooks.

Leanpub is a magical typewriter for authors: just write in plain text, and to publish your ebook, just click a button. (Or, if you are producing your ebook your own way, you can even upload your own PDF and/or EPUB files and then publish with one click!) It really is that easy.

Learn more about writing on Leanpub