Leanpub Header

Skip to main content

Computer Use Agents: A Practical Guide to Building AI That Controls Your Desktop

From First Prototype to Production-Ready Systems

This book is 100% completeLast updated on 2026-07-23

Build AI agents that can see screens, reason about tasks and control desktop applications and web browsers. Starting with a simple prototype, you'll learn how to create reliable production-ready systems that do real work on a computer.

Minimum price

$19.00

$29.00

You pay

Author earns

$

Also available for 1 book credit with a Reader Membership

PDF
EPUB
WEB
APP
204
Pages
About

About

About the Book

This book teaches you how to build computer use agents that perceive screens, reason about tasks, and execute actions across desktop applications and web browsers. Starting from a working prototype in under thirty minutes, you will incrementally add perception pipelines, planning systems, memory, tool integrations, safety guardrails, and production infrastructure. Every chapter builds on the previous one with production-quality code, real-world case studies, and honest discussion of trade-offs and failure modes. If you are a software engineer who wants to move beyond chatbots and build agents that actually do work on a computer, this is the book for you.

Bundle

Bundles that include this book

Author

About the Author

Steve Publications

Steve is a technology professional with more than 20 years of experience in software development, server infrastructure, cybersecurity, vulnerability research and reverse engineering. Throughout his career, he has designed, secured, analyzed and tested complex software and infrastructure, with a particular focus on understanding how systems fail and how they can be made more secure.

Outside of work, Steve enjoys sharing knowledge with the technology community. He collaborates with researchers, industry experts and technology professionals to write practical books covering software development, cybersecurity, cloud computing, networking, DevOps, artificial intelligence and enterprise technologies. His books focus on practical learning through clear explanations, real-world examples and hands-on exercises. With more than two decades of industry experience, his goal is to help IT professionals, students and technology enthusiasts build useful skills and stay current in a rapidly changing industry.

We believe readers deserve to know how our books are created. Most of our authors are not native English speakers, so we use AI to help translate, proofread manuscripts, fix grammar, improve sentence structure and make technical explanations easier to read. AI is used as an editing tool only. It does not replace the research, technical knowledge or hands-on experience behind our books. Some of our authors also prefer to remain anonymous for privacy or professional reasons. In those cases, we publish their work under a different name. The author's name may be different, but the quality of the content and our review process remain the same.

Every book is written, reviewed and maintained by experienced technology professionals, with contributions from our private technical community of more than 400 engineers and researchers from Ukraine, Belarus and Russia. We spend far more time validating technical accuracy and keeping our content up to date than generating text. We are always interested in working with experienced professionals who have deep expertise in a particular technology or domain. If you would like to publish a book with us or help review an existing manuscript, we'd love to hear from you. Send us a message describing your area of expertise. We are especially interested in niche technologies, specialized skills and emerging topics that are underrepresented in existing technical literature.

If you look through the contents of our books, you'll see practical examples, detailed explanations and material that is regularly updated. Our goal is to publish books that professionals can actually rely on, not low-effort AI-generated content. If you ever feel that one of our books does not meet that standard, Leanpub offers a 60-day money-back guarantee. Feel free to request a refund if you are not satisfied with your purchase.

Contents

Table of Contents

From First Prototype to Production-Ready Systems

Introduction: The Agent on Your Screen

  1. What Is a Computer Use Agent?
  2. Why Now: The Convergence That Made This Possible
  3. The Promise and the Peril
  4. How This Book Is Structured
  5. Your First CUA in 30 Seconds

Chapter 1: Foundations: LLMs, Agents, and the Control Loop

  1. The Agent Abstraction: Perception, Thought, Action
  2. Why LLMs Are Different From Previous AI
  3. The Control Loop: Observe, Decide, Act, Repeat
  4. When Prompting Is Enough and When It Is Not
  5. Designing the Minimal CUA Architecture

Chapter 2: Your First Agent: A Working Prototype in Python

  1. Environment Setup: Tools, Libraries, and Dependencies
  2. Capturing the Screen: Your First Perception Pipeline
  3. Sending Screenshots to an LLM
  4. Parsing Actions From Model Output
  5. Executing Clicks, Keystrokes, and Navigation
  6. Running the Prototype End to End
  7. What This Prototype Gets Right
  8. What This Prototype Gets Wrong

Chapter 3: Seeing the Interface: Multimodal Perception for UI Understanding

  1. Screenshots Alone Are Not Enough
  2. Optical Character Recognition: Tesseract, PaddleOCR, and Beyond
  3. Accessibility Trees and DOM Extraction
  4. Hybrid Perception: Combining Vision, Text, and Structure
  5. Object Detection for UI Elements
  6. Choosing Your Perception Stack

Chapter 4: Planning: From Single Steps to Multi-Horizon Reasoning

  1. The One-Step Problem: Why Naive Agents Fail on Complex Tasks
  2. Chain of Thought for Task Decomposition
  3. Planning Ahead: Tree Search and MCTS for Agents
  4. Hierarchical Planning: High-Level Goals and Low-Level Actions
  5. Integrating Planning Into the Control Loop

Chapter 5: Memory: Giving Your Agent Context and Continuity

  1. Why Stateless Agents Cannot Do Real Work
  2. Short-Term Memory: Conversation History and Context Windows
  3. Episodic Memory: Recording Past Trajectories
  4. Long-Term Memory: Vector Stores and Retrieval-Augmented Agents
  5. Procedural Memory: Learning Skills and Routines
  6. Integrating Memory Into the Agent Architecture

Chapter 6: Tools and APIs: Extending Beyond Mouse and Keyboard

  1. The Limits of Pure GUI Automation
  2. Structured Tool Calling: OpenAI Functions, Claude Tools, and MCP
  3. Designing a Tool Schema Your Agent Can Use Reliably
  4. Browser APIs: Direct DOM Manipulation vs Clicking
  5. File System, Clipboard, and System-Level Tools
  6. The Model Context Protocol (MCP)
  7. When to Use Tools versus UI Automation

Chapter 7: Action Execution: Reliable Control of Desktop and Browser

  1. The Action Space: What Can Your Agent Do?
  2. Mouse and Keyboard: pyautogui, Input Events, and Pitfalls
  3. Browser Automation: Playwright, Selenium, and Puppeteer Compared
  4. Accessibility APIs: Direct Element Control
  5. Handling Timing, Loading States, and Race Conditions
  6. Error Recovery When Actions Fail

Chapter 8: Prompt Engineering for Agents: Designing Robust System Prompts

  1. Why Agent Prompts Are Different From Chat Prompts
  2. The Anatomy of a Production System Prompt
  3. Defining the Action Schema Clearly
  4. Few-Shot Examples That Actually Help
  5. Constraining Behavior: Guardrails in the Prompt
  6. Testing and Iterating on Prompts

Chapter 9: State Management: Tracking Progress Through Complex Tasks

  1. The State Problem: How Does the Agent Know Where It Is?
  2. Task Graphs and State Machines
  3. Progress Tracking: Checklists, Milestones, and Completion Criteria
  4. Self-Correction When the Agent Loses Track
  5. Persistence: Saving and Restoring Long-Running Tasks

Chapter 10: Safety and Guardrails: Preventing Catastrophic Mistakes

  1. What Can Go Wrong: A Taxonomy of Agent Failures
  2. Confirmation Flows and Human Oversight
  3. Sandboxing: VMs, Containers, and Isolated Environments
  4. Action Restrictions: Whitelists, Blacklists, and Risk Scoring
  5. Rate Limiting, Budget Caps, and Emergency Stops
  6. Adversarial Considerations and Prompt Injection

Chapter 11: Evaluation: Measuring Whether Your Agent Actually Works

  1. Why “It Looks Cool” Is Not an Evaluation Metric
  2. Task Success Rate: Defining What “Done” Means
  3. Automated Evaluation Harnesses and Test Suites
  4. Human Evaluation: When You Need a Person in the Loop
  5. Efficiency Metrics: Steps, Time, and Cost
  6. Existing Benchmarks: WebArena, OSWorld, and Others

Chapter 12: Debugging and Observability: Understanding Agent Behavior

  1. Why Agents Are Hard to Debug
  2. Structured Logging for Every Step
  3. Tracing Decisions: From Observation to Action
  4. Replay and Recording: Watching the Agent Run Backwards
  5. Visualization Dashboards for Agent Behavior
  6. Common Failure Patterns and How to Diagnose Them

Chapter 13: Optimization: Speed, Cost, and Reliability at Scale

  1. The Three Constraints: Latency, Cost, Reliability
  2. Model Selection: When to Use Small Models vs Large Ones
  3. Caching Strategies for Perception and Planning
  4. Parallelizing Independent Subtasks
  5. Infrastructure: GPUs, Edge Devices, and Cloud Costs

Chapter 14: Security: Protecting Your Agent and Your Users

  1. The Attack Surface of an Agent System
  2. API Key Management and Secrets Handling
  3. Data Privacy: Screenshots, Logs, and User Content
  4. Secure Execution: Sandboxing, Least Privilege, and Isolation
  5. Transport Security and Integrity Verification
  6. Threat Modeling Your CUA Deployment

Chapter 15: Deployment: From Local Script to Production Service

  1. Packaging Your Agent for Deployment
  2. Containerization with Docker
  3. Orchestration: Kubernetes, Serverless, and Managed Services
  4. Multi-Tenancy: Supporting Multiple Users Safely
  5. Monitoring, Alerting, and Incident Response
  6. CI/CD for Agent Systems

Chapter 16: Case Studies: Real-World CUA Implementations

  1. Automated Software Testing: A QA Agent
  2. Customer Support: Agents That Navigate Legacy Systems
  3. Research Assistant: Cross-Tab Data Collection and Synthesis
  4. Personal Productivity: Email, Calendar, and File Management
  5. Lessons Learned Across Deployments

Conclusion: The Next Frontier: Where CUA Technology Is Heading

  1. What We Have Built: A Recap of the Journey
  2. Emerging Paradigms: World Models and Embodied Agents
  3. Agentic Workflows: When Multiple Agents Collaborate
  4. The Human-Agent Interface: Designing for Partnership
  5. Your Next Steps as a CUA Builder

References

Get the free sample chapters

Click the buttons to get the free sample in PDF or EPUB, or read the sample online here

The Leanpub 60 Day 100% Happiness Guarantee

Within 60 days of purchase you can get a 100% refund on any Leanpub purchase, in two clicks.

See full terms...

Earn $8 on a $10 Purchase, and $16 on a $20 Purchase

We pay 80% royalties on purchases of $7.99 or more, and 80% royalties minus a 50 cent flat fee on purchases between $0.99 and $7.98. You earn $8 on a $10 sale, and $16 on a $20 sale. So, if we sell 5000 non-refunded copies of your book for $20, you'll earn $80,000.

(Yes, some authors have already earned much more than that on Leanpub.)

In fact, authors have earned over $15 million writing, publishing and selling on Leanpub.

Learn more about writing on Leanpub

Free Updates. DRM Free.

If you buy a Leanpub book, you get free updates for as long as the author updates the book! Many authors use Leanpub to publish their books in-progress, while they are writing them. All readers get free updates, regardless of when they bought the book or how much they paid (including free).

Most Leanpub books are available in PDF (for computers) and EPUB (for phones, tablets and Kindle). The formats that a book includes are shown at the top right corner of this page.

Finally, Leanpub books don't have any DRM copy-protection nonsense, so you can easily read them on any supported device.

Learn more about Leanpub's ebook formats and where to read them

Write and Publish on Leanpub

You can use Leanpub to easily write, publish and sell in-progress and completed ebooks and online courses!

Leanpub is a powerful platform for serious authors, combining a simple, elegant writing and publishing workflow with a store focused on selling in-progress ebooks.

Leanpub is a magical typewriter for authors: just write in plain text, and to publish your ebook, just click a button. (Or, if you are producing your ebook your own way, you can even upload your own PDF and/or EPUB files and then publish with one click!) It really is that easy.

Learn more about writing on Leanpub