Leanpub Header

Skip to main content

Production-Grade Python, LangChain & LangGraph

A deployment-first handbook for developers who know Python syntax but have not shipped Python

Production-Grade Python, LangChain & LangGraph
This book is 100% completeLast updated on 2026-09-20

AI engineering is becoming one of the most valuable and in-demand areas of software development, and Python, LangChain, and LangGraph are core skills for building the systems behind it. Learn how to turn basic Python knowledge into production-grade backend and agentic AI applications—and move toward an engineering niche centered on building and controlling AI rather than competing with it.

Minimum price

$9.99

$14.99

You pay

Author earns

$

Also available for 1 book credit with a Reader Membership

PDF
EPUB
WEB
APP
About

About

About the Book

Learn the engineering skills behind one of software’s most valuable emerging fields

AI engineering is rapidly becoming one of the most commercially valuable areas of software development.

Companies are moving beyond simple chatbot prototypes and looking for engineers who can build reliable AI-powered applications, agentic workflows, retrieval systems, model integrations, and production infrastructure. That creates strong demand for developers who understand both software engineering fundamentals and modern AI application frameworks.

Python sits at the center of this ecosystem. It is the dominant language across AI, machine learning, data engineering, model tooling, and an increasing share of agent development. LangChain and LangGraph add the abstractions needed to build applications around large language models, tools, retrieval, state, and increasingly autonomous workflows.

These skills also occupy an interesting position in the changing software market. AI can automate more routine implementation work, but organizations still need engineers who can design the systems around AI: integration, orchestration, reliability, evaluation, security, deployment, observability, and production control. Learning to build the systems that use AI is therefore a compelling way to move toward the part of software engineering that AI itself is helping to create.

Production-Grade Python, LangChain & LangGraph is designed to take you there.

The book assumes that you already know basic Python syntax. Instead of teaching loops, functions, and classes again, it focuses on the engineering knowledge required to turn Python code into real production software.

You will learn:

  • How Python actually executes code through CPython, bytecode, modules, and processes
  • Modern project structure with pyproject.toml
  • Virtual environments, dependency management, locking, packaging, wheels, and reproducible builds
  • FastAPI and production API design
  • Configuration and secrets management
  • Async programming and concurrency
  • Database access, transactions, migrations, and persistence
  • Caching and external service integration
  • Error handling, retries, timeouts, and idempotency
  • Structured logging, metrics, tracing, and health checks
  • Unit, integration, API, and system testing
  • Containers, CI/CD, and production deployment
  • Security practices for backend and AI applications

The second half applies those engineering foundations to modern AI engineering with LangChain and LangGraph.

You will learn how to build:

  • LLM-powered applications
  • Structured model interactions
  • Tool-using agents
  • Retrieval-Augmented Generation systems
  • Stateful workflows
  • Conditional and parallel execution
  • Persistent agent state
  • Durable workflows
  • Human-in-the-loop approval
  • Streaming applications
  • Long-running agent processes
  • Multi-step and multi-agent systems
  • Testable and observable AI workflows

The emphasis throughout is on production engineering, not demos.

Calling an LLM is easy. Building a system around it that remains secure, reliable, observable, recoverable, and maintainable is where much of the real engineering value lies.

A complete production-style support-agent application brings the concepts together using FastAPI, SQLAlchemy, LangChain, LangGraph, persistence, authentication, observability, human approval, and safe execution of irreversible actions.

The book also includes a companion code repository containing the complete application, runnable examples, tests, deployment configuration, and every code sample used throughout the manuscript.

If you already know basic Python and want to position yourself for the growing field of AI engineering, this book gives you the production foundations and agent-development skills needed to move beyond prototypes and build systems organizations can actually depend on.

Author

About the Author

Fiodar Sazanavets

Fiodar Sazanavets is a senior software engineer specializing in AI systems, distributed applications, and data-intensive software. A former Microsoft engineer and fourt-time Microsoft MVP, he has more than a decade of professional experience designing and building production systems across a wide range of industries.

His core areas of expertise include Python, .NET, cloud technologies, distributed systems, and modern AI engineering. Throughout his career, he has worked on systems ranging from railway passenger information platforms and distributed IoT clusters to e-commerce applications and financial transaction-processing systems. He has also led engineering teams, mentored developers, and helped organizations translate complex technical requirements into reliable software that solves real business problems.

Alongside his engineering work, Fiodar is passionate about technical education. He writes books, creates online courses, mentors developers, and regularly publishes practical software engineering content. His teaching focuses on the skills developers need to build real-world systems: architecture, reliability, cloud engineering, AI, and production software development.

He writes about software engineering and AI at fiodar.substack.com/.

The Leanpub Podcast

Launch

Launch Video

Podcast

Podcast Episode

Subscribe on YouTube

Clips

Clips

Contents

Table of Contents

About This Handbook

What This Book Is About

How to Use This Handbook

  1. Part I - Runtime, packaging, and deployment

1. What “compiling and deploying Python” actually means

  1. 1.1 Source, bytecode, wheels, and executables are different things
  2. 1.2 What compileall is for - and what it is not for
  3. 1.3 The three main deployment units
  4. 1.5 Build once, deploy the same artifact

2. A production-ready Python project skeleton

  1. 2.1 pyproject.toml is the project control plane
  2. 2.2 Virtual environments are process isolation for dependencies
  3. 2.3 Lock application dependencies
  4. 2.4 Imports should reflect package boundaries

3. Building and installing release artifacts

  1. 3.1 Why a wheel is useful even when you deploy a container
  2. 3.2 Editable installs are for development, not production
  3. 3.3 Pin the runtime as well as Python packages
  4. 3.4 Do not build with production secrets

4. Configuration, settings, and secrets

  1. 4.1 Validate settings at process startup
  2. 4.2 Never log secrets
  3. 4.3 Prefer identity to long-lived secrets where possible

5. Processes, ASGI, and serving an HTTP application

  1. 5.1 Understand the process model
  2. 5.2 Containers often favor one server process per container
  3. 5.3 Use lifespan for startup and shutdown resources
  4. 5.4 Graceful shutdown is part of correctness

6. Health checks, readiness, and startup

  1. 6.1 Startup probes protect slow initialization
  2. 6.2 Do not run destructive migrations on every replica startup

7. Containers: the default service deployment pattern

  1. 7.1 A production Dockerfile should optimize reproducibility and attack surface
  2. 7.2 Build dependencies should not remain in the runtime image
  3. 7.3 Run as a non-root user
  4. 7.4 Treat the filesystem as ephemeral
  5. 7.5 Signals and PID 1 matter

8. Deployment choices: VM, containers, serverless, Kubernetes

  1. 8.1 Virtual machine deployment
  2. 8.2 Managed containers
  3. 8.3 Serverless functions
  4. 8.4 Kubernetes fundamentals for Python services
  5. Part II - Production Python engineering

9. Architecture and dependency direction

  1. 9.1 Dependency injection can be simple in Python
  2. 9.2 Keep import-time work minimal

10. Types, schemas, dataclasses, and Pydantic

  1. 10.1 Use type hints to make contracts explicit
  2. 10.2 Use Pydantic at untrusted boundaries
  3. 10.3 Validation is not authorization
  4. 10.4 Avoid Any as an architecture strategy

11. Errors and exception design

  1. 11.1 Exceptions should communicate semantic failure
  2. 11.2 Do not catch Exception just to continue
  3. 11.3 Preserve causes
  4. 11.4 Expected failures should not become 500s

12. Structured logging

  1. 12.1 Logs are event records, not print statements
  2. 12.2 Correlation IDs connect the request path
  3. 12.3 Log at the right level
  4. 12.4 Redact sensitive content

13. Testing: a production pyramid that actually helps

  1. 13.1 Running tests in Python
  2. 13.2 Unit tests isolate logic and run fast
  3. 13.3 Fixtures reduce repeated test setup
  4. 13.4 Integration tests verify infrastructure contracts
  5. 13.5 End-to-end tests are few but valuable
  6. 13.6 Test behavior, not implementation trivia
  7. 13.7 Use mocks carefully
  8. 13.8 Tests should be deterministic
  9. 13.9 LLM tests require a different strategy

14. Static checks, formatting, and developer automation

  1. 14.1 Automate style so humans review design
  2. 14.2 Make the same commands run locally and in CI
  3. 14.3 Treat warnings as decisions

15. Concurrency: async, threads, processes, and the GIL

  1. 15.1 Concurrency and parallelism are not the same thing
  2. 15.2 Understanding the event loop
  3. 15.3 async def helps only when the call chain is asynchronous
  4. 15.4 Do not sprinkle async everywhere
  5. 15.5 Concurrent I/O can reduce latency
  6. 15.6 Do not create unlimited concurrency
  7. 15.7 Threads are useful, especially for blocking I/O
  8. 15.8 Understand the GIL
  9. 15.9 Processes provide stronger CPU parallelism
  10. 15.10 Timeouts are mandatory for network I/O
  11. 15.11 Cancellation is a correctness concern
  12. 15.12 Use processes or job workers for heavy CPU tasks
  13. 15.13 Choosing between async, threads, and processes

16. Outbound HTTP: resilience without retry storms

  1. 16.1 Centralize clients
  2. 16.2 Retry only failures that are plausibly transient
  3. 16.3 Exponential backoff and jitter
  4. 16.4 Idempotency makes safe retries possible
  5. 16.5 Circuit breakers are about protecting capacity

17. PostgreSQL and SQLAlchemy fundamentals

  1. 17.1 A database connection is a scarce pooled resource
  2. 17.2 Make transaction boundaries explicit
  3. 17.3 commit() is not a detail to scatter everywhere
  4. 17.4 Avoid the N+1 query problem
  5. 17.5 Constraints belong in the database too
  6. 17.6 Migrations are versioned code

18. Safe database migrations and zero-downtime releases

  1. 18.1 Rolling deployments create mixed-version windows
  2. 18.2 Dangerous migration patterns
  3. 18.3 Rollback is not always “run the down migration”

19. Background jobs, queues, and message processing

  1. 19.1 Do not hide durable work in an in-process task
  2. 19.2 Assume at-least-once delivery unless proven otherwise
  3. 19.3 Dead-letter queues preserve failures for investigation
  4. 19.4 Event schemas are APIs

20. Caching with Redis and application caches

  1. 20.1 Cache only when you know the source of truth
  2. 20.2 Common patterns
  3. 20.3 Cache invalidation requires a clear policy
  4. 20.4 Cache failure should usually not become data loss
  5. 20.5 In-memory caches are per process

21. API design fundamentals

  1. 21.1 Stable external contracts deserve explicit schemas
  2. 21.2 Pagination must be designed before tables become huge
  3. 21.3 Version only when necessary, but plan compatibility
  4. 21.4 Error responses should be machine-readable

22. Authentication, authorization, and service security

  1. 22.1 Authentication answers who; authorization answers whether
  2. 22.2 Validate tokens fully
  3. 22.3 The principle of least privilege applies to application identities
  4. 22.4 Protect against SSRF
  5. 22.5 Dependency and supply-chain security

23. Observability: logs, metrics, and traces

  1. 23.1 The three signals answer different questions
  2. 23.2 Start with service-level indicators
  3. 23.3 Histograms beat average latency
  4. 23.4 Distributed tracing turns a request into a timeline
  5. 23.5 OpenTelemetry is a useful vendor-neutral foundation

24. Reliability patterns

  1. 24.1 Define failure domains
  2. 24.2 Bulkheads isolate scarce resources
  3. 24.3 Backpressure is better than collapse
  4. 24.4 Graceful degradation must be explicit

25. Performance and profiling

  1. 25.1 Measure before optimizing
  2. 25.2 Benchmark realistic workloads
  3. 25.3 LLM latency often dominates application latency

26. CI/CD for a Python service

  1. 26.1 CI should recreate the project from a clean state
  2. 26.2 Example GitHub Actions skeleton
  3. 26.3 Deployments need verification and rollback
  4. 26.4 Feature flags decouple deploy from release

27. Production checklist before adding AI

  1. Part III - LangChain fundamentals

28. Where LangChain fits in a production architecture

  1. 28.1 LangChain is an integration and agent framework, not the whole application
  2. 28.2 LangChain, LangGraph, and LangSmith solve different problems
  3. 28.3 Provider integrations are separate packages

29. Models: the fundamental reasoning/generation interface

  1. 29.1 Initialize a chat model
  2. 29.2 Core invocation styles
  3. 29.3 Configure hard budgets around model calls

30. Messages and conversational context

  1. 30.1 Messages are the canonical context unit
  2. 30.2 Message roles have semantic meaning
  3. 30.3 Tool messages complete tool-call conversations
  4. 30.4 Context is a scarce resource

31. Prompt templates and runnable composition

  1. 31.1 Prompts are versioned application logic
  2. 31.2 Compose runnables for deterministic pipelines
  3. 31.3 Prefer structured output over parsing prose

32. Tools: giving a model controlled capabilities

  1. 32.1 A tool is an API contract exposed to the model
  2. 32.2 Tool descriptions should be precise
  3. 32.3 Keep dangerous capabilities narrow
  4. 32.4 Authorization belongs inside the tool boundary too
  5. 32.5 ToolRuntime exposes trusted execution context

33. The agent loop with create_agent

  1. 33.1 An agent is a controlled loop, not “the LLM thinks forever”
  2. 33.2 Minimal agent
  3. 33.3 Put hard stop conditions around autonomous loops
  4. 33.4 Agent state can hold more than messages

34. Structured output from agents

  1. 34.1 Application code should consume typed results
  2. 34.2 Validate semantics after schema validation

35. Middleware: cross-cutting control around agents

  1. 35.1 Middleware is where production policy often belongs
  2. 35.2 Keep deterministic policy deterministic
  3. 35.3 Summarization middleware manages long threads

36. Retrieval-Augmented Generation (RAG) fundamentals

  1. 36.1 RAG separates knowledge retrieval from generation
  2. 36.2 LangChain’s core retrieval abstractions
  3. 36.3 Chunking is an information-retrieval decision
  4. 36.4 Metadata is essential
  5. 36.5 Retrieval authorization must happen before content reaches the model
  6. 36.6 Retrieval quality and generation quality are separate

37. A minimal production-shaped RAG service

  1. 37.1 Indexing path and query path are different workloads
  2. 37.2 Keep provenance with retrieved content
  3. 37.3 Protect against prompt injection in documents

38. Short-term and long-term memory

  1. 38.1 Short-term memory is thread-scoped state
  2. 38.2 Production checkpointers must be durable and shared
  3. 38.3 Long-term memory is cross-thread data
  4. 38.4 Memory is not a substitute for systems of record

39. Testing LangChain applications

  1. 39.1 Test prompt construction without a live model
  2. 39.2 Test tools as ordinary functions
  3. 39.3 Test the tool schema exposed to the model
  4. 39.4 Use fake model responses for orchestration tests

40. LangSmith observability and evaluation

  1. 40.1 Agent traces need more detail than normal HTTP traces
  2. 40.2 Offline evaluation protects releases
  3. 40.3 Online evaluation watches real traffic
  4. 40.4 LLM-as-judge is a measurement tool, not ground truth
  5. Part IV - LangGraph fundamentals and production patterns

41. Why LangGraph exists

  1. 41.1 Use a graph when the workflow topology matters
  2. 41.2 The core model is state + nodes + edges
  3. 41.3 “Compile” means validate and build the graph runtime

42. State schemas

  1. 42.1 TypedDict is a common lightweight state shape
  2. 42.2 Use state for workflow data, not every dependency
  3. 42.3 Keep state small
  4. 42.4 Separate input and internal state when useful

43. Reducers: how concurrent state updates are merged

  1. 43.1 Default updates replace values
  2. 43.2 Reducers define accumulation semantics
  3. 43.3 Reducers are critical under parallelism
  4. 43.4 Message state uses specialized message reducers

44. Nodes and edges

  1. 44.1 Nodes should do one meaningful step
  2. 44.2 Fixed edges express deterministic sequencing
  3. 44.3 Conditional edges express routing

45. Command: update state and choose control flow together

  1. 45.1 Command is useful when routing belongs to the node result
  2. 45.2 Keep routing destinations constrained

46. Loops and evaluator-reviser workflows

  1. 46.1 Graphs naturally model iteration
  2. 46.2 Add an iteration budget
  3. 46.3 Self-critique is not independent verification

47. Parallel branches and fan-out

  1. 47.1 Parallelize truly independent work
  2. 47.2 Send supports dynamic fan-out
  3. 47.3 Apply concurrency limits

48. Persistence and checkpointers

  1. 48.1 Checkpoints save graph state at execution steps
  2. 48.2 thread_id is the durable execution identity
  3. 48.3 Use production persistence for production workflows
  4. 48.4 Persistence changes your data-retention obligations

49. Durable execution, determinism, and idempotency

  1. 49.1 Resuming a workflow can replay code
  2. 49.2 Make side-effecting operations idempotent
  3. 49.3 Persist important external results
  4. 49.4 Exactly-once is usually an end-to-end property, not a broker setting

50. Retry policies in graphs

  1. 50.1 Retry at the smallest sensible unit
  2. 50.2 Separate transient from permanent errors
  3. 50.3 Keep retry budgets below workflow budgets

51. Interrupts and human-in-the-loop

  1. 51.1 Some decisions should pause rather than guess
  2. 51.2 Resumption must be authenticated and authorized
  3. 51.3 Approval UI should show grounded evidence

52. Streaming graphs

  1. 52.1 Streaming improves perceived latency and transparency
  2. 52.2 Streaming does not reduce total work by itself
  3. 52.3 Use protocol-level streaming designed for proxies

53. Subgraphs

  1. 53.1 Subgraphs are workflow modules
  2. 53.2 Decide how subgraph state should persist
  3. 53.3 Keep contracts narrow

54. Multi-agent systems

  1. 54.1 Multiple agents are not automatically better than one
  2. 54.2 A supervisor pattern
  3. 54.3 Security boundaries can justify subagents

55. A production LangGraph support workflow

  1. 55.1 Requirements
  2. 55.2 State
  3. 55.3 Topology
  4. 55.4 Deterministic policy before model quality
  5. 55.5 Keep external actions after approval
  6. 55.6 Persistence and streaming complete the production shape

56. Deploying LangChain/LangGraph applications

  1. 56.1 You can host the graph inside your own API
  2. 56.2 Agent Server / LangGraph deployment is another option
  3. 56.3 Local development and production are different

Part VI - End-to-end reference implementation

57. The final project layout

  1. 57.1 Application startup
  2. 57.2 Startup should verify configuration, not every external dependency

58. Exposing a graph through an authenticated API

  1. 58.1 The API owns identity
  2. 58.2 Authorize thread ownership before state access

59. A safe mutation tool

60. Local infrastructure and the production release pipeline

  1. 60.1 Use Docker Compose for local dependencies
  2. 60.2 The release pipeline should verify more than Python code
  3. 60.3 Build once and promote the same artifact
  4. 60.4 Verify both system health and AI behavior
  5. 60.5 Use different evaluation depth at different stages

Final words

The Leanpub 60 Day 100% Happiness Guarantee

Within 60 days of purchase you can get a 100% refund on any Leanpub purchase, in two clicks.

See full terms...

Earn $8 on a $10 Purchase, and $16 on a $20 Purchase

We pay 80% royalties on purchases of $7.99 or more, and 80% royalties minus a 50 cent flat fee on purchases between $0.99 and $7.98. You earn $8 on a $10 sale, and $16 on a $20 sale. So, if we sell 5000 non-refunded copies of your book for $20, you'll earn $80,000.

(Yes, some authors have already earned much more than that on Leanpub.)

In fact, authors have earned over $15 million writing, publishing and selling on Leanpub.

Learn more about writing on Leanpub

Free Updates. DRM Free.

If you buy a Leanpub book, you get free updates for as long as the author updates the book! Many authors use Leanpub to publish their books in-progress, while they are writing them. All readers get free updates, regardless of when they bought the book or how much they paid (including free).

Most Leanpub books are available in PDF (for computers) and EPUB (for phones, tablets and Kindle). The formats that a book includes are shown at the top right corner of this page.

Finally, Leanpub books don't have any DRM copy-protection nonsense, so you can easily read them on any supported device.

Learn more about Leanpub's ebook formats and where to read them

Write and Publish on Leanpub

You can use Leanpub to easily write, publish and sell in-progress and completed ebooks and online courses!

Leanpub is a powerful platform for serious authors, combining a simple, elegant writing and publishing workflow with a store focused on selling in-progress ebooks.

Leanpub is a magical typewriter for authors: just write in plain text, and to publish your ebook, just click a button. (Or, if you are producing your ebook your own way, you can even upload your own PDF and/or EPUB files and then publish with one click!) It really is that easy.

Learn more about writing on Leanpub