About This Handbook
What This Book Is About
How to Use This Handbook
- Part I - Runtime, packaging, and deployment
1. What “compiling and deploying Python” actually means
- 1.1 Source, bytecode, wheels, and executables are different things
- 1.2 What
compileallis for - and what it is not for - 1.3 The three main deployment units
- 1.5 Build once, deploy the same artifact
2. A production-ready Python project skeleton
- 2.1
pyproject.tomlis the project control plane - 2.2 Virtual environments are process isolation for dependencies
- 2.3 Lock application dependencies
- 2.4 Imports should reflect package boundaries
3. Building and installing release artifacts
- 3.1 Why a wheel is useful even when you deploy a container
- 3.2 Editable installs are for development, not production
- 3.3 Pin the runtime as well as Python packages
- 3.4 Do not build with production secrets
4. Configuration, settings, and secrets
- 4.1 Validate settings at process startup
- 4.2 Never log secrets
- 4.3 Prefer identity to long-lived secrets where possible
5. Processes, ASGI, and serving an HTTP application
- 5.1 Understand the process model
- 5.2 Containers often favor one server process per container
- 5.3 Use lifespan for startup and shutdown resources
- 5.4 Graceful shutdown is part of correctness
6. Health checks, readiness, and startup
- 6.1 Startup probes protect slow initialization
- 6.2 Do not run destructive migrations on every replica startup
7. Containers: the default service deployment pattern
- 7.1 A production Dockerfile should optimize reproducibility and attack surface
- 7.2 Build dependencies should not remain in the runtime image
- 7.3 Run as a non-root user
- 7.4 Treat the filesystem as ephemeral
- 7.5 Signals and PID 1 matter
8. Deployment choices: VM, containers, serverless, Kubernetes
- 8.1 Virtual machine deployment
- 8.2 Managed containers
- 8.3 Serverless functions
- 8.4 Kubernetes fundamentals for Python services
- Part II - Production Python engineering
9. Architecture and dependency direction
- 9.1 Dependency injection can be simple in Python
- 9.2 Keep import-time work minimal
10. Types, schemas, dataclasses, and Pydantic
- 10.1 Use type hints to make contracts explicit
- 10.2 Use Pydantic at untrusted boundaries
- 10.3 Validation is not authorization
- 10.4 Avoid
Anyas an architecture strategy
11. Errors and exception design
- 11.1 Exceptions should communicate semantic failure
- 11.2 Do not catch
Exceptionjust to continue - 11.3 Preserve causes
- 11.4 Expected failures should not become 500s
12. Structured logging
- 12.1 Logs are event records, not print statements
- 12.2 Correlation IDs connect the request path
- 12.3 Log at the right level
- 12.4 Redact sensitive content
13. Testing: a production pyramid that actually helps
- 13.1 Running tests in Python
- 13.2 Unit tests isolate logic and run fast
- 13.3 Fixtures reduce repeated test setup
- 13.4 Integration tests verify infrastructure contracts
- 13.5 End-to-end tests are few but valuable
- 13.6 Test behavior, not implementation trivia
- 13.7 Use mocks carefully
- 13.8 Tests should be deterministic
- 13.9 LLM tests require a different strategy
14. Static checks, formatting, and developer automation
- 14.1 Automate style so humans review design
- 14.2 Make the same commands run locally and in CI
- 14.3 Treat warnings as decisions
15. Concurrency: async, threads, processes, and the GIL
- 15.1 Concurrency and parallelism are not the same thing
- 15.2 Understanding the event loop
- 15.3
async defhelps only when the call chain is asynchronous - 15.4 Do not sprinkle
asynceverywhere - 15.5 Concurrent I/O can reduce latency
- 15.6 Do not create unlimited concurrency
- 15.7 Threads are useful, especially for blocking I/O
- 15.8 Understand the GIL
- 15.9 Processes provide stronger CPU parallelism
- 15.10 Timeouts are mandatory for network I/O
- 15.11 Cancellation is a correctness concern
- 15.12 Use processes or job workers for heavy CPU tasks
- 15.13 Choosing between async, threads, and processes
16. Outbound HTTP: resilience without retry storms
- 16.1 Centralize clients
- 16.2 Retry only failures that are plausibly transient
- 16.3 Exponential backoff and jitter
- 16.4 Idempotency makes safe retries possible
- 16.5 Circuit breakers are about protecting capacity
17. PostgreSQL and SQLAlchemy fundamentals
- 17.1 A database connection is a scarce pooled resource
- 17.2 Make transaction boundaries explicit
- 17.3
commit()is not a detail to scatter everywhere - 17.4 Avoid the N+1 query problem
- 17.5 Constraints belong in the database too
- 17.6 Migrations are versioned code
18. Safe database migrations and zero-downtime releases
- 18.1 Rolling deployments create mixed-version windows
- 18.2 Dangerous migration patterns
- 18.3 Rollback is not always “run the down migration”
19. Background jobs, queues, and message processing
- 19.1 Do not hide durable work in an in-process task
- 19.2 Assume at-least-once delivery unless proven otherwise
- 19.3 Dead-letter queues preserve failures for investigation
- 19.4 Event schemas are APIs
20. Caching with Redis and application caches
- 20.1 Cache only when you know the source of truth
- 20.2 Common patterns
- 20.3 Cache invalidation requires a clear policy
- 20.4 Cache failure should usually not become data loss
- 20.5 In-memory caches are per process
21. API design fundamentals
- 21.1 Stable external contracts deserve explicit schemas
- 21.2 Pagination must be designed before tables become huge
- 21.3 Version only when necessary, but plan compatibility
- 21.4 Error responses should be machine-readable
22. Authentication, authorization, and service security
- 22.1 Authentication answers who; authorization answers whether
- 22.2 Validate tokens fully
- 22.3 The principle of least privilege applies to application identities
- 22.4 Protect against SSRF
- 22.5 Dependency and supply-chain security
23. Observability: logs, metrics, and traces
- 23.1 The three signals answer different questions
- 23.2 Start with service-level indicators
- 23.3 Histograms beat average latency
- 23.4 Distributed tracing turns a request into a timeline
- 23.5 OpenTelemetry is a useful vendor-neutral foundation
24. Reliability patterns
- 24.1 Define failure domains
- 24.2 Bulkheads isolate scarce resources
- 24.3 Backpressure is better than collapse
- 24.4 Graceful degradation must be explicit
25. Performance and profiling
- 25.1 Measure before optimizing
- 25.2 Benchmark realistic workloads
- 25.3 LLM latency often dominates application latency
26. CI/CD for a Python service
- 26.1 CI should recreate the project from a clean state
- 26.2 Example GitHub Actions skeleton
- 26.3 Deployments need verification and rollback
- 26.4 Feature flags decouple deploy from release
27. Production checklist before adding AI
- Part III - LangChain fundamentals
28. Where LangChain fits in a production architecture
- 28.1 LangChain is an integration and agent framework, not the whole application
- 28.2 LangChain, LangGraph, and LangSmith solve different problems
- 28.3 Provider integrations are separate packages
29. Models: the fundamental reasoning/generation interface
- 29.1 Initialize a chat model
- 29.2 Core invocation styles
- 29.3 Configure hard budgets around model calls
30. Messages and conversational context
- 30.1 Messages are the canonical context unit
- 30.2 Message roles have semantic meaning
- 30.3 Tool messages complete tool-call conversations
- 30.4 Context is a scarce resource
31. Prompt templates and runnable composition
- 31.1 Prompts are versioned application logic
- 31.2 Compose runnables for deterministic pipelines
- 31.3 Prefer structured output over parsing prose
32. Tools: giving a model controlled capabilities
- 32.1 A tool is an API contract exposed to the model
- 32.2 Tool descriptions should be precise
- 32.3 Keep dangerous capabilities narrow
- 32.4 Authorization belongs inside the tool boundary too
- 32.5
ToolRuntimeexposes trusted execution context
33. The agent loop with create_agent
- 33.1 An agent is a controlled loop, not “the LLM thinks forever”
- 33.2 Minimal agent
- 33.3 Put hard stop conditions around autonomous loops
- 33.4 Agent state can hold more than messages
34. Structured output from agents
- 34.1 Application code should consume typed results
- 34.2 Validate semantics after schema validation
35. Middleware: cross-cutting control around agents
- 35.1 Middleware is where production policy often belongs
- 35.2 Keep deterministic policy deterministic
- 35.3 Summarization middleware manages long threads
36. Retrieval-Augmented Generation (RAG) fundamentals
- 36.1 RAG separates knowledge retrieval from generation
- 36.2 LangChain’s core retrieval abstractions
- 36.3 Chunking is an information-retrieval decision
- 36.4 Metadata is essential
- 36.5 Retrieval authorization must happen before content reaches the model
- 36.6 Retrieval quality and generation quality are separate
37. A minimal production-shaped RAG service
- 37.1 Indexing path and query path are different workloads
- 37.2 Keep provenance with retrieved content
- 37.3 Protect against prompt injection in documents
38. Short-term and long-term memory
- 38.1 Short-term memory is thread-scoped state
- 38.2 Production checkpointers must be durable and shared
- 38.3 Long-term memory is cross-thread data
- 38.4 Memory is not a substitute for systems of record
39. Testing LangChain applications
- 39.1 Test prompt construction without a live model
- 39.2 Test tools as ordinary functions
- 39.3 Test the tool schema exposed to the model
- 39.4 Use fake model responses for orchestration tests
40. LangSmith observability and evaluation
- 40.1 Agent traces need more detail than normal HTTP traces
- 40.2 Offline evaluation protects releases
- 40.3 Online evaluation watches real traffic
- 40.4 LLM-as-judge is a measurement tool, not ground truth
- Part IV - LangGraph fundamentals and production patterns
41. Why LangGraph exists
- 41.1 Use a graph when the workflow topology matters
- 41.2 The core model is state + nodes + edges
- 41.3 “Compile” means validate and build the graph runtime
42. State schemas
- 42.1
TypedDictis a common lightweight state shape - 42.2 Use state for workflow data, not every dependency
- 42.3 Keep state small
- 42.4 Separate input and internal state when useful
43. Reducers: how concurrent state updates are merged
- 43.1 Default updates replace values
- 43.2 Reducers define accumulation semantics
- 43.3 Reducers are critical under parallelism
- 43.4 Message state uses specialized message reducers
44. Nodes and edges
- 44.1 Nodes should do one meaningful step
- 44.2 Fixed edges express deterministic sequencing
- 44.3 Conditional edges express routing
45. Command: update state and choose control flow together
- 45.1
Commandis useful when routing belongs to the node result - 45.2 Keep routing destinations constrained
46. Loops and evaluator-reviser workflows
- 46.1 Graphs naturally model iteration
- 46.2 Add an iteration budget
- 46.3 Self-critique is not independent verification
47. Parallel branches and fan-out
- 47.1 Parallelize truly independent work
- 47.2
Sendsupports dynamic fan-out - 47.3 Apply concurrency limits
48. Persistence and checkpointers
- 48.1 Checkpoints save graph state at execution steps
- 48.2
thread_idis the durable execution identity - 48.3 Use production persistence for production workflows
- 48.4 Persistence changes your data-retention obligations
49. Durable execution, determinism, and idempotency
- 49.1 Resuming a workflow can replay code
- 49.2 Make side-effecting operations idempotent
- 49.3 Persist important external results
- 49.4 Exactly-once is usually an end-to-end property, not a broker setting
50. Retry policies in graphs
- 50.1 Retry at the smallest sensible unit
- 50.2 Separate transient from permanent errors
- 50.3 Keep retry budgets below workflow budgets
51. Interrupts and human-in-the-loop
- 51.1 Some decisions should pause rather than guess
- 51.2 Resumption must be authenticated and authorized
- 51.3 Approval UI should show grounded evidence
52. Streaming graphs
- 52.1 Streaming improves perceived latency and transparency
- 52.2 Streaming does not reduce total work by itself
- 52.3 Use protocol-level streaming designed for proxies
53. Subgraphs
- 53.1 Subgraphs are workflow modules
- 53.2 Decide how subgraph state should persist
- 53.3 Keep contracts narrow
54. Multi-agent systems
- 54.1 Multiple agents are not automatically better than one
- 54.2 A supervisor pattern
- 54.3 Security boundaries can justify subagents
55. A production LangGraph support workflow
- 55.1 Requirements
- 55.2 State
- 55.3 Topology
- 55.4 Deterministic policy before model quality
- 55.5 Keep external actions after approval
- 55.6 Persistence and streaming complete the production shape
56. Deploying LangChain/LangGraph applications
- 56.1 You can host the graph inside your own API
- 56.2 Agent Server / LangGraph deployment is another option
- 56.3 Local development and production are different
Part VI - End-to-end reference implementation
57. The final project layout
- 57.1 Application startup
- 57.2 Startup should verify configuration, not every external dependency
58. Exposing a graph through an authenticated API
- 58.1 The API owns identity
- 58.2 Authorize thread ownership before state access
59. A safe mutation tool
60. Local infrastructure and the production release pipeline
- 60.1 Use Docker Compose for local dependencies
- 60.2 The release pipeline should verify more than Python code
- 60.3 Build once and promote the same artifact
- 60.4 Verify both system health and AI behavior
- 60.5 Use different evaluation depth at different stages