Leanpub Header

Skip to main content

Filters

Category: "Observability"

Observability

  1. eBPF for Systems Engineers
    eBPF for Systems Engineers
    From First Principles to Production Deployments
    Steve Publications

    Discover how eBPF is changing Linux systems engineering. Starting with the fundamentals, this book walks you through building, debugging and deploying real eBPF programs for observability, networking, security and performance. Learn practical techniques you can use with confidence in production across modern Linux environments.

  2. Self-Healing Infrastructure: Building Autonomous Cloud Systems with AI

    Most writing about AI and infrastructure stops at the demo. This book starts on the day the model is confidently wrong at 3 a.m. and an auditor asks who authorised the action it took. Seven working labs — MCP servers with real identity and audit, closed-loop remediation behind a reversibility gate, and autonomy that can be revoked — for platform engineers in safety-critical and regulated industries.

  3. Building Your LLM Twin
    Building Your LLM Twin
    A Production-Grade Guide to Engineering Large Language Models from Concept to Deployment
    Steve Publications

    What if you could build an AI that writes just like you? In Building Your LLM Twin, you'll follow a hands-on journey from raw writing data to a fully deployed AI that captures your unique voice. Learn to build, fine-tune and scale real-world LLM systems with practical code, proven techniques and the tools to take your project from concept to production.

  4. Building Reliable LLM Abstractions and API Wrappers
    Building Reliable LLM Abstractions and API Wrappers
    Architecture, Implementation, and Production Engineering
    Steve Publications

    A practical guide to building reliable LLM integrations that hold up in production. Learn how to design provider-agnostic APIs, handle failures, add observability and security, and keep your architecture maintainable. With a complete Python codebase, this book turns proven patterns into working software.

  5. Unified Observability with OpenObserve
    Unified Observability with OpenObserve
    From red flags to root cause across infrastructure, network, application, data and AI layers
    GitforGits | Asian Publishing House

    In this book, one estate runs through all fourteen chapters. There are eight services across two regions, a hybrid link to an older data centre, a storefront that real customers use, and one service backed by a language model. We use OpenObserve for that, bring all the signals into it, shape the data until long-term storage becomes affordable, and then spend the rest of the time diagnosing the estate when it breaks. There won't be a feature tour here. The features don't stand the test of time, and the screenshots are even worse. What you'll find is a method. First, scope the flag, orient before acting, form a hypothesis you can disprove, test it against evidence rather than instinct, repair, verify, and add the prevention that stops the same failure twice. That loop works whatever platform you use.

  6. Observability Tools
    Observability Tools
    Design & Comparison Guide
    Sudhanshu Jaiswal

    "Your systems are talking. Are you listening?"In a world where downtime costs millions and silent failures hide in plain sight, observability isn’t just a buzzword—it’s your superpower. But with hundreds of tools, frameworks, and metrics to choose from, how do you cut through the noise?Observability Tools: Design & Comparison Guide is your no-BS playbook to: ✅ Design observability pipelines tailored to AI/ML, microservices, and event-driven architectures. ✅ Compare the top 10 tools—from Datadog to Grafana Stack—and pick the perfect fit for your budget and scale. ✅ Avoid costly mistakes with red flags, alert fatigue fixes, and cost-saving strategies.

  7. Production Debugging with Sentry
    Production Debugging with Sentry
    A Practical SRE Playbook for Tracing, Logging, Profiling and Alerting Across Full-Stack Production Systems
    GitforGits | Asian Publishing House

    A brilliant engineer with no evidence is reduced to guessing, while an ordinary engineer with good evidence looks like a genius. Everything in this book is an attempt to get you to think differently. You don't need a huge team to do this. You don't need to get anyone's approval for a budget, and you don't have to rewrite any services. What you need is a system that tells you the truth about itself, and you need it before the incident rather than during it.As a team, we'll build that system together on one small storefront, and we'll do it the way real teams build, which is to say not perfectly but in order.

  8. OTLP on the Wire
    OTLP on the Wire
    Luciana Reynaud Ferreira

    Read OTLP from raw bytes to spans: Protobuf wire format, gRPC frames, HTTP payloads, PCAPs, failure modes, and telemetry cost—explained for engineers who need to debug the protocol, not merely configure it.

  9. Production LLMOps
    Production LLMOps
    Designing, Deploying, and Operating Large Language Model Systems at Scale
    Steve Publications

    Building an LLM demo is easy. Running one reliably at scale is not. Production LLMOps shows you how to design, deploy and operate real-world LLM systems, covering RAG, agents, fine-tuning, evaluation, CI/CD, observability, security and more, with practical code you can adapt for production.

  10. AI-Native Systems Engineering: Building, Running, and Securing Autonomous Software on Linux
    AI-Native Systems Engineering: Building, Running, and Securing Autonomous Software on Linux
    A Production-Focused Handbook for Designing Useful, Controllable, Observable, Reliable, Maintainable, Secure, and Production-Ready Autonomous Systems
    Steve Publications

    Build AI systems that do more than demo well. This hands-on guide shows you how to engineer autonomous software on Linux that is secure, observable, reliable and ready for production. From agents and memory to threat modeling, OpenTelemetry and Kubernetes, you’ll build a real system while learning what it takes to run AI you can actually trust.

  11. OpenTelemetry in Production
    OpenTelemetry in Production
    Designing, Deploying, and Operating Observability at Scale
    Steve Publications

    Take OpenTelemetry from theory to production. This practical guide shows you how to design, deploy and operate observability at scale, with real configurations, proven architecture patterns and hands-on guidance for performance, security, troubleshooting and cost. Built for engineers who need telemetry that works when it matters.