Leanpub Header

Skip to main content

Filters

Category: "Eval Engineering"

Eval Engineering

  1. Multi-Agent Failure Behaviors
    Multi-Agent Failure Behaviors
    How to recognize pathological behaviors of your agents and properly handle them to avoid hallucinations
    Luca Bianchi
    No Description Available
  2. Specification Engineering
    Specification Engineering
    The Complete Guide to Principles, Strategies, Methods, and Practical Application
    Steve Publications

    Specification Engineering turns complex ideas into clear, workable specifications. This practical guide takes you from core principles to advanced methods for defining, validating, managing and evolving requirements across products, software, systems and services. Built for real-world use, it gives you the tools to create specifications that people can understand and act on.

  3. Python AI Programming, Second Edition
    Python AI Programming, Second Edition
    Kickstart developing AI-ready apps with RAG, DSPy, MCP, agents, evals, observability and open-source models
    GitforGits | Asian Publishing House

    Vectors, embeddings, retrieval, agents, and evaluation are all built from first principles inside the chapter that needs them. No mathematics. No machine learning background. No prior AI experience and no framework knowledge is required. We build with plain Python and small, single-purpose libraries.  

  4. Failure-First AI Agents
    Failure-First AI Agents
    A Practical Field Guide to Evaluating RAG, Tool Use, Memory, and Multi-Agent Systems
    storymaker

    A hands-on, failure-first guide to evaluating LLM agents, RAG, tool use, grounding, and production release gates—with executable Python examples and tests.

  5. Modern Code Review in the AI Era
    Modern Code Review in the AI Era
    A Comprehensive Guide to Building Better Software Through Collaborative Review and Intelligent Automation
    Steve Publications

    Great software is built through great reviews. This book shows you how to spot hidden bugs, improve design, strengthen security, and review AI-generated code with confidence. Packed with practical examples and proven techniques, it helps you write better software in a world where humans and AI build code together.

  6. Evaluation Engineering for AI Systems
    Evaluation Engineering for AI Systems
    Building Reliable Evaluations from Benchmarks to Production
    Steve Publications

    Evaluation Engineering for AI Systems is a practical guide to building reliable AI evaluations for production. Learn how to measure performance, compare models, detect regressions and create evaluation pipelines using real-world examples, modern Python and proven industry practices.