Leanpub Header

Skip to main content

Filters

Category: "Eval Engineering"

Eval Engineering

  1. Multi-Agent AI Systems The Complete Handbook for Building Intelligent  Scalable, and Autonomous Agent Teams

    What happens when AI agents stop working alone and start working as a team?Artificial Intelligence is entering a new phase in which intelligent systems can do more than respond to individual instructions. Multiple specialized agents can collaborate, divide complex tasks, communicate with one another, use tools, evaluate results, and coordinate their actions toward a shared objective.Multi-Agent AI Systems: The Complete Handbook for Building Intelligent, Scalable, and Autonomous Agent Teams provides a practical roadmap for understanding this emerging paradigm.The book begins with the fundamentals of multi-agent systems and explains why collaboration between specialized agents can be valuable for complex workflows. Readers will learn about hierarchical, peer-to-peer, and hybrid architectures, along with roles such as manager, planner, worker, critic, and supervisor agents.It then moves into modern frameworks and technologies, including CrewAI, AutoGen, LangGraph, MetaGPT, LLMs, vector databases, and agent memory systems. Practical chapters explain how to design agent teams, decompose tasks, establish communication protocols, manage shared memory, coordinate workflows, integrate tools and APIs, and recover from failures.Readers will also explore advanced concepts such as dynamic replanning, parallel execution, swarm intelligence, agent debates, self-organizing systems, human-in-the-loop workflows, and multimodal agents.The book goes beyond experimentation and addresses the challenges of deploying multi-agent systems in real environments. Cloud deployment, Docker, Kubernetes, monitoring, logging, scaling, evaluation, benchmarking, testing, and cost optimization are included.Real-world applications demonstrate how agent teams can support software development, research, customer service, content creation, and business operations.Equally important, the book examines AI safety, privacy, security, transparency, governance, alignment, and responsible AI development.Whether you are a student discovering agentic AI, a developer building your first agent team, a researcher exploring collaborative intelligence, or a professional preparing for the next generation of AI applications, this book provides a foundation for moving from individual AI agents toward coordinated intelligent systems.Understand the architecture. Design the team. Build the agents. Coordinate the intelligence.

  2. Silent Wiring
    Silent Wiring
    The Hidden Failure Mode in AI-Generated Code
    Tom Killi

    Your AI-generated code passes 18,000 tests and reports healthy — while entire data pipelines silently produce nothing. Silent Wiring names the failure mode nobody's tooling catches, and shows you how to find it before your users do.

  3. Governance Before Decision
    Governance Before Decision
    A Field Guide to Proof, Permission, and Human-Controlled AI Execution
    Kari Hyotyla

    AI can decide. That does not mean it is authorized to act. Governance Before Decision examines the missing control layer between intelligent systems and real-world consequences — where authority, scope, verification, and execution must meet.

  4. Agentic  AI: The Rise of Autonomous Goal Driven Intelligence

    What happens when AI moves beyond answering questions and begins to pursue goals?Agentic AI: The Rise of Autonomous, Goal-Driven Intelligence explores the emerging world of intelligent systems that can perceive, reason, plan, learn, collaborate, and act within defined environments.From the foundations of autonomous agents and multi-agent systems to robotics, finance, logistics, security, ethics, and future research, this book provides a structured journey through the rapidly evolving landscape of Agentic AI.Discover how goal-driven intelligence works, how autonomous agents make decisions, how multiple agents collaborate, and why safety, trust, accountability, and human oversight are essential to the future of autonomous systems.Understand the agents. Explore the architectures. Examine the applications. Think about the future.

  5. Mastering Artificial Intelligence & Machine Learning System Desig
    Mastering Artificial Intelligence & Machine Learning System Desig
    A Complete Interview Guide with Frameworks, Case Studies & Insider Strategies
    Anshuman Mishra

    Master AI/ML System Design with a Structured Interview FrameworkLearn how to approach complex ML system design interviews using a 20-step framework, architecture blueprints, real-world case studies, scalability strategies, data and model pipelines, deployment patterns, monitoring, and interview-ready templates.

  6. The Conversational AI Revolution: Building  Intelligent Chatbots and Voice Assistants

    Build Smarter Chatbots and Voice AssistantsExplore Conversational AI from fundamentals to real-world applications. Learn NLP, intent and context management, intelligent chatbots, voice assistants, speech technologies, multimodal interfaces, APIs, deployment, business integration, analytics, ethics, LLMs, and Generative AI.

  7. Cybersecurity in Artificial Intelligence: Attacks Defenses and Real World Application

    Secure AI. Understand the Threats. Build More Trustworthy Intelligent Systems.Explore adversarial attacks, data poisoning, model theft, AI-enabled cyber threats, secure AI development, adversarial defenses, trustworthy AI, MLOps security, governance, privacy, red teaming, and the future of AI cybersecurity through practical concepts and real-world case studies.

  8. Intelligent Machines: How AI is  Revolutionizing Robotics

    Explore the Future of AI-Powered RoboticsDiscover how Artificial Intelligence is transforming robots into intelligent machines capable of learning, seeing, communicating, navigating, and making decisions. Explore Machine Learning, Computer Vision, NLP, Deep Learning, autonomous systems, smart factories, healthcare robots, drones, humanoids, agricultural robotics, ethics, and future trends.

  9. Evaluating Local and Small Language Models
    Evaluating Local and Small Language Models
    A Practical Engineering Guide to Benchmarking, Optimization, and Production Deployment
    Steve Publications

    A practical guide to evaluating small and local language models in the real world. Learn how to benchmark reliably, compare models across hardware and runtimes, optimize performance and avoid misleading results. With practical code and rigorous methods, you can build your own evaluation setup and make confident deployment decisions.

  10. Semantic Search from Scratch
    Semantic Search from Scratch
    Build a Working Semantic Search Engine with Pure Python and NumPy
    AhmedAdawy

    Build a fully functional semantic search engine from first principles using pure Python and NumPy—no heavy AI frameworks, no vector databases, just pure intuition and mathematics.

  11. Specification Engineering
    Specification Engineering
    The Complete Guide to Principles, Strategies, Methods, and Practical Application
    Steve Publications

    Specification Engineering turns complex ideas into clear, workable specifications. This practical guide takes you from core principles to advanced methods for defining, validating, managing and evolving requirements across products, software, systems and services. Built for real-world use, it gives you the tools to create specifications that people can understand and act on.

  12. Python AI Programming, Second Edition
    Python AI Programming, Second Edition
    Kickstart developing AI-ready apps with RAG, DSPy, MCP, agents, evals, observability and open-source models
    GitforGits | Asian Publishing House

    Vectors, embeddings, retrieval, agents, and evaluation are all built from first principles inside the chapter that needs them. No mathematics. No machine learning background. No prior AI experience and no framework knowledge is required. We build with plain Python and small, single-purpose libraries.  

  13. Multi-Agent Failure Behaviors
    Multi-Agent Failure Behaviors
    How to recognize pathological behaviors of your agents and properly handle them to avoid hallucinations
    Luca Bianchi

    Eight agents, zero coordination, one identical error. That's not a coincidence. It's a signature. Learn to read the five families of failure behaviors in multi-agent systems, reproduce each one in TypeScript, and turn every bias into a CI assertion before it ships.

  14. Failure-First AI Agents
    Failure-First AI Agents
    A Practical Field Guide to Evaluating RAG, Tool Use, Memory, and Multi-Agent Systems
    storymaker

    A hands-on, failure-first guide to evaluating LLM agents, RAG, tool use, grounding, and production release gates—with executable Python examples and tests.

  15. Modern Code Review in the AI Era
    Modern Code Review in the AI Era
    A Comprehensive Guide to Building Better Software Through Collaborative Review and Intelligent Automation
    Steve Publications

    Great software is built through great reviews. This book shows you how to spot hidden bugs, improve design, strengthen security, and review AI-generated code with confidence. Packed with practical examples and proven techniques, it helps you write better software in a world where humans and AI build code together.