Leanpub Header

Skip to main content

Filters

Category: "Site Reliability Engineering"

Site Reliability Engineering

  1. Observability-Driven Testing: Testing in Production
    Observability-Driven Testing: Testing in Production
    Logs, Metrics, Traces, and Quality in Live Systems
    Yuri Syuganov

    The deploy is not the finish line. Use logs, traces, metrics, canaries, and feature flags to keep testing after release — where it counts most.

  2. Self-Healing Infrastructure: Building Autonomous Cloud Systems with AI

    Most writing about AI and infrastructure stops at the demo. This book starts on the day the model is confidently wrong at 3 a.m. and an auditor asks who authorised the action it took. Seven working labs — MCP servers with real identity and audit, closed-loop remediation behind a reversibility gate, and autonomy that can be revoked — for platform engineers in safety-critical and regulated industries.