Designing, Building, and Operating Reliable Distributed Systems
Introduction
- What This Book Covers
- How This Book Is Structured
- Who Should Read This Book
- The Philosophy Behind This Book
Chapter 1: The Reality of Distributed Systems
- The Moment Everything Changes
- The Fundamental Tensions: Consistency, Availability, and Partition Tolerance
- Latency, Concurrency, and Partial Failure
- Why “Just Add a Message Broker” Is Not an Answer
- Communication Models: RPC, Messaging, Events, and Their Trade-offs
- The Operational Dimension: What Production Really Demands
- Introducing OrderFlow: Our Reference Platform
Chapter 2: Modern Java for Platform Engineering
- Why Java Matters for Distributed Systems
- Java 21 Features That Matter: Records, Sealed Classes, Pattern Matching, Virtual Threads
- JVM Fundamentals for Production: Memory Model, Garbage Collection, Tuning Basics
- Build Tooling: Maven, Modularization, and Dependency Management
- Concurrency in Java: Executors, Virtual Threads, Structured Concurrency
- Structured Logging and Observability-Ready Code
- Defensive Programming: Null Safety, Immutable Data, and Failure Handling
Chapter 3: Architectural Principles and Patterns
- From Monolith to Distributed: When and Why to Break Apart
- Domain-Driven Design: Bounded Contexts, Ubiquitous Language, and Strategic Design
- Event-Driven Architecture: Core Concepts and Anti-Patterns
- Microservices and Service-Oriented Architecture: Truths and Misconceptions
- Synchronous vs Asynchronous: Choosing the Right Communication Style
- Designing for Failure: Resilience Patterns and Graceful Degradation
Chapter 4: Spring Boot for Production Systems
- How Spring Boot Works: Auto-Configuration, Starters, and the Application Context
- Dependency Injection and the Component Model
- Configuration Management: Profiles, Externalization, and Hierarchies
- REST with Spring Web MVC and WebFlux
- Actuator: Health, Metrics, and Production Visibility
- Lifecycle Management and Graceful Shutdown
Chapter 5: Data Access and Persistence with Spring Data
- Spring Data JPA: Entities, Repositories, and Query Methods
- Transaction Management: Propagation, Isolation, and Distributed Considerations
- Database Migrations with Flyway
- Connection Pooling and Performance
- Choosing Persistence Technologies: Relational vs Document vs Specialized Stores
- Persistence Anti-Patterns and Pitfalls
Chapter 6: RabbitMQ Fundamentals
- RabbitMQ Architecture: Broker, Nodes, Clusters, and the Management Layer
- The AMQP Protocol: Concepts, Versioning, and Alternatives
- Connections, Channels, and Resource Management
- Queues, Exchanges, and Bindings: The Core Data Model
- Message Flow and Lifecycle
- RabbitMQ Management: UI, HTTP API, and CLI Tools
Chapter 7: Exchange Types, Routing, and Topology Design
- Direct Exchanges: Point-to-Point Routing
- Fanout Exchanges: Broadcasting and Multicast
- Topic Exchanges: Pattern-Based Routing and Hierarchies
- Headers Exchanges and Use Cases
- Designing Exchange and Queue Topologies for OrderFlow
- Virtual Hosts, Namespaces, and Multi-Tenancy
Chapter 8: Message Delivery Semantics and Reliability
- At-Least-Once, At-Most-Once, and Exactly-Once: What They Really Mean
- Publisher Confirms: Ensuring Messages Reach the Broker
- Consumer Acknowledgments: Manual vs Automatic, Negative Acknowledgments
- Message Durability: Persistent Messages and Queue Durability
- Transactional Publishing: AMQP Transactions and Trade-offs
- Implementing Reliable Delivery in OrderFlow
Chapter 9: Advanced Messaging Patterns with RabbitMQ
- Dead Letter Exchanges and Queues: Handling Failed Messages
- Message Timeouts and TTL-Based Expiration
- Delayed Message Processing with Plugins and Alternatives
- Priority Queues and Work Prioritization
- Queue Length Limits and Overflow Handling
- Message Inspection and Redelivery Control
Chapter 10: RabbitMQ Clustering, High Availability, and Scaling
- RabbitMQ Clustering: Architecture, Topologies, and Quorum
- HA Queues and Quorum Queues: Choosing the Right Model
- Scaling Out: Adding Nodes, Rebalancing, and Monitoring
- Broker Upgrades and Maintenance Procedures
- Disaster Recovery Strategies and Data Backup
- Performance Tuning: Memory, Disk, and Erlang Considerations
Chapter 11: Spring AMQP Fundamentals
- The Spring AMQP Abstraction: RabbitTemplate, ConnectionFactory, and Channel
- Configuring RabbitMQ Connections in Spring Boot
- Sending Messages: RabbitTemplate and Producer Configuration
- Receiving Messages: RabbitListener and Consumer Configuration
- Message Conversion and Serialization
- Error Handling Fundamentals: DefaultErrorHandler and Beyond
Chapter 12: Building Reliable Producers with Spring AMQP
- Publisher Confirms and Returns in Spring AMQP
- ConfirmsCallback and ReturnCallback: Handling Broker Feedback
- Retry Semantics: SimpleMessageConverter and RetryTemplate Integration
- Correlation IDs and Message Tracing
- Batching and Bulk Publishing
- Producer Performance and Throughput Optimization
Chapter 13: Building Reliable Consumers with Spring AMQP
- Acknowledgment Modes: AUTO, MANUAL, and BATCH
- Concurrency Tuning: Prefetch, Consumer Threads, and Queue Partitioning
- The DefaultMessageListenerContainer Lifecycle and Configuration
- Consumer Error Handling: Retry, Backoff, and DLQ Routing
- Idempotent Consumer Design
- Graceful Consumer Shutdown and Message In-Flight Handling
Chapter 14: Message Serialization, Schema Evolution, and Compatibility
- JSON Serialization with Jackson: Configuration and Customization
- Avro and Schema Registry for Contract-Governed Messages
- Protobuf and Performance-Sensitive Scenarios
- Schema Evolution Strategies: Backward, Forward, and Full Compatibility
- Versioning Messages and Supporting Multiple Consumers
- Message Schema Documentation and Governance
Chapter 15: Event-Driven Design and the Outbox Pattern
- Domain Events and Integration Events: Defining the Distinction
- The Transactional Outbox Pattern: Theory and Implementation
- Implementing Outbox with Spring Data and Polling
- CDC-Based Outbox: Debezium and Change Data Capture
- Event Publishing in OrderFlow: From Database to RabbitMQ
- Testing Event-Driven Components
Chapter 16: The Inbox Pattern and Idempotent Processing
- Why Idempotency Matters: Duplicate Detection and Processing
- The Inbox Pattern: Tracking Received Messages
- Implementing Idempotency Keys and State Stores
- Outbox-Inbox: End-to-End Reliable Communication
- Handling Partial Failures and Compensating Actions
- Idempotency in OrderFlow Services
Chapter 17: Sagas and Distributed Transactions
- Why Distributed Transactions Fail: Two-Phase Commit and Its Problems
- Saga Pattern: Choreography vs Orchestration
- Implementing Sagas with RabbitMQ and Spring Events
- Compensating Actions and Rollback Semantics
- Saga State Machines and Persistence
- OrderFlow Saga: Implementing an Order Fulfillment Workflow
Chapter 18: Message Ordering and Sequence Guarantees
- When Ordering Matters and When It Does Not
- FIFO Queues and Single-Consumer Constraints
- Partitioning by Key for Ordered Substreams
- Sequence Numbers and Gap Detection
- Implementing Order Guarantees in RabbitMQ
- Ordering Trade-offs: Throughput vs Determinism
Chapter 19: Concurrency, Backpressure, and Flow Control
- Understanding Throughput, Latency, and Saturation
- Backpressure Mechanisms: Prefetch, Queue Limits, and Rate Limiting
- Spring AMQP Flow Control and Consumer Throttling
- Circuit Breakers and Bulkheads: Protecting Services
- Implementing Flow Control in OrderFlow
- Capacity Planning and Load Testing
Chapter 20: Security, Authentication, and Authorization
- RabbitMQ Security: Users, VHosts, Permissions, and TLS
- Spring Security with OAuth2 and JWT
- Service-to-Service Authentication Patterns
- Secrets Management: Vault, Kubernetes Secrets, and Environment Config
- Encrypting Data in Transit and at Rest
- Security Testing and Compliance
Chapter 21: Observability — Logging, Metrics, and Tracing
- Structured Logging with Micrometer and Logback
- Metrics with Micrometer: Gauges, Counters, Timers, and Histograms
- Distributed Tracing with Micrometer Tracing and Brave/OpenTelemetry
- Correlating Logs, Metrics, and Traces
- RabbitMQ Metrics and Broker Visibility
- Implementing Observability in OrderFlow
Chapter 22: Containerization, Kubernetes, and Deployment
- Docker Images for Spring Boot Applications
- Docker Compose for Local Development and Integration Testing
- Kubernetes Deployments, Services, and ConfigMaps
- RabbitMQ on Kubernetes: Helm Charts and Operators
- Ingress, Networking, and Service Mesh Considerations
- Deployment Strategies: Rolling, Blue-Green, and Canary
Chapter 23: Testing Strategy for Distributed Systems
- Unit Testing Spring and RabbitMQ Components
- Integration Testing with Testcontainers: RabbitMQ, Databases, and More
- Contract Testing with Spring Cloud Contract
- End-to-End Testing in an Isolated Environment
- Testing Failure Scenarios: Network Partition, Broker Down, Slow Consumers
- Test Infrastructure and CI Pipelines
Chapter 24: CI/CD, Automation, and Release Engineering
- CI Pipeline Design: Build, Test, and Quality Gates
- Container Registry and Image Tagging Strategy
- CD Strategies and Automated Deployments
- Environment Promotion and Feature Flags
- Rollback Procedures and Release Management
- The OrderFlow CI/CD Pipeline
Chapter 25: Troubleshooting, Maintenance, and Long-Term Evolution
- Common Failure Modes and Their Symptoms
- RabbitMQ Troubleshooting: Connections, Queues, Memory, and Disk
- Spring Boot Troubleshooting: Memory Leaks, Thread Pools, and Timeouts
- Incident Response and Debugging Production Issues
- Dependency Management and Version Upgrades
- Technical Debt, Refactoring, and Architectural Evolution
Chapter 26: The Production OrderFlow Platform
- Full Architecture: Services, Events, and Flows
- Production Deployment: Complete Kubernetes and Infrastructure Configuration
- Operational Runbook: Monitoring, Scaling, and Incident Response
- Performance Characteristics and Capacity Profiles
- Evolutionary Roadmap: How the Architecture Could Grow
- Lessons Learned and Key Takeaways