The Databricks platform and data-engineering playbook for the engineers who own pipelines, govern catalogs, and keep workloads on schedule. Sixteen chapters on Unity Catalog, Lakeflow, identity, observability, and performance. Azure examples; concepts mapped to AWS and GCP.
RAG, Agent Bricks, the Multi-Agent Supervisor with MCP, Lakebase, MLflow 3, Lakehouse Monitoring, Feature Store, Vector Search. Every AI surface Databricks shipped at GA in 2025 and 2026, taught by a practitioner, current to 2026. What you will learn - Build RAG pipelines with Vector Search, embedding models, and citation grounding- Ship Agent Bricks for classification and information extraction- Orchestrate specialist agents with the Multi-Agent Supervisor and MCP- Use Lakebase as the operational Postgres layer for AI apps and agents- Detect data and model drift with Lakehouse Monitoring; wire alerts to retraining- Manage the ML lifecycle with MLflow 3 and the UC Model Registry- Govern features across training and serving with Feature Store (offline + online)- Serve foundation and custom models with AI Gateway controls Who this book is for Data engineers, ML engineers, and AI/ML architects who know PySpark and the Databricks platform and now need to ship production AI. Volume 3 is the recommended prerequisite. Table of Contents 1. Databricks SQL in Production. Warehouses, materialized views, three latency signals (admission, compilation, execution), the full dashboard backend wiring.2. External BI: Tableau, Power BI, dbt. Performance tips that take a dashboard from sluggish to instant, dbt configuration at incremental scale, the seam between BI and the lakehouse.3. AI/BI Dashboards. Anatomy of a Lakeview dashboard, draft vs published flow, the Dashboard Agent's reliable patterns, the five-grant permission model.4. Genie: Natural-Language Analytics. Grounding sources, the priority rule, the SQL Genie actually writes, the questions Genie answers cleanly versus the ones that confuse it.5. AI SQL Functions. ai_query, ai_parse_document, ai_extract for PDFs and HTML, univariate forecasts, the daily cost math for production AI SQL pipelines.6. Model Serving. Endpoints, the three fields that decide capacity and cost, the chat-completion payload, the five moving pieces of a production recommender.7. Foundation Models. Five major providers, the External Models config, the vendor-swap pattern (Claude to Gemini in hours, not weeks), the three habits that keep swap cost low.8. Vector Search and RAG. Six delta-sync arguments, three chunking strategies compared, the RAG function your app imports, end-to-end answer evaluation with traces.9. MLflow 3 and UC Model Registry. Versions, aliases, tags (and what each is not for), five tracking calls and what each one writes, the experiment-to-production lifecycle.10. Feature Store. Why SDP is the right producer, the six-file project layout, four parity-failure classes between offline and online stores and what causes each.11. MLOps as a Practice. Seven sources every incident reads from, three deploy patterns (canary, shadow, blue-green), three retrain strategies, five golden signals for an ML endpoint.12. Lakehouse Monitoring: Drift Detection. Six monitor parameters, the loop from drift alert to retraining, what to do when the baseline table is missing.13. Distributed Deep Learning. Three signals that force distributed training, picking the flavor (data, model, hybrid) from the bottleneck, four pieces of GPU memory worked out for a 7B model.14. Agent Bricks. Declarative classification and information-extraction agents, eval-set ingredients, the pre-compute pattern that makes small seed sets work.15. Multi-Agent Supervisor and MCP. The supervisor build, synthetic-turn evaluation, three real conversations end to end, the auth-passthrough chain across child agents.16. Lakebase: Operational Postgres for AI. Five alternatives compared, sub-10ms reads for AI apps, the lineage from Delta source through SDP into Postgres and onward to the endpoint.17. Capstone: Retail Intelligence App. Ten stages, each anchored to an earlier chapter. The smoke test that confirms every stage of the platform is reachable, the new-data path through the recommender.18. Certification and What's Next. The certification paths that actually map to the book, and the reading list the on-call team uses when something breaks.
Data platform interviews test CDC, streaming, warehouse modeling, and trade-offs two levels deep — no generic system-design book covers it. This one does: 50 full mock interviews, 122 rapid-fire Q&A, and the framework that turns a vague prompt into a hire.
Explore data before you trust it. Data Exploration and Visualisation with R shows how to uncover structure, diagnose data-quality problems, challenge misleading patterns, and turn imperfect observations into credible evidence through thoughtful, reproducible visual analysis.
Stop only using databases — start understanding them. A living study of PostgreSQL internals for backend engineers: how pages, indexes, and WAL actually work, written clearly while learning, not after mastery.
Learn pandas, the most widely used data analysis library.
What I love about Dask is that it won't make us start from scratch. It doesn't ask us to learn a new language, adopt a new way of thinking, or rewrite five years of careful analysis into something completely different. It just takes the tools we've already got and quietly teaches them to work in parallel. The dataframe remains a dataframe. The array is still an array. The difference is that the work now spreads across eight cores, or eight hundred, and we barely had to change how we think.
AI can analyze the data. But who decides what the data means—and what to do next?Modern organizations have more information than ever, yet making good decisions remains difficult.AI-Powered Smart Decision Making shows how to combine AI, data, analytics, and human judgment to turn information into insight and insight into action.Discover a practical framework for making better decisions across strategy, marketing, sales, operations, finance, risk, and beyond.AI is not the decision. AI improves the decision process. Table of ContentsIntroduction—The New Era of AI-Powered Decision MakingChapter 1 — The Foundations of Better Business DecisionsChapter 2 — From Intuition to Data-Driven Decision MakingChapter 3 — Understanding Data for Better DecisionsChapter 4 — Business Analytics: Turning Data into InsightChapter 5 — How AI Changes Business Decision MakingChapter 6 — Building an AI-Powered Decision FrameworkChapter 7 — Asking the Right Questions: AI, Data & Decision ProblemsChapter 8 — Predictive Decision MakingChapter 9 — Prescriptive Analytics and Decision OptimizationChapter 10 — AI for Strategic Business DecisionsChapter 11 — AI for Marketing and Customer DecisionsChapter 12 — AI for Sales, Operations and Supply ChainChapter 13 — AI for Financial and Risk DecisionsChapter 14 — Human + AI: The New Decision-Making PartnershipChapter 15 — Bias, Uncertainty and the Limits of AI DecisionsChapter 16 — Trustworthy, Responsible and Explainable AI DecisionsChapter 17 — Measuring Decision Quality and Business ImpactChapter 18 — Building an AI-Powered Decision CultureChapter 19 — Designing an AI-Powered Decision SystemChapter 20 — The Future of Smart Decision MakingConclusion — From Data to Better DecisionsDisclaimer
The ultimate technical blueprint for developers and AI engineers to strip away execution overhead, leverage hardware-accelerated NumPy/PyTorch operations, and scale Python code effortlessly
Running ClickHouse at petabyte scale takes more than knowing SQL. This practical guide shows experienced engineers how to design, deploy and operate production clusters, from data modeling and ingestion to Kubernetes, query tuning, observability, disaster recovery and cost control, with real-world patterns and concrete configurations throughout.
Your data warehouse holds the answers — LLMs can help you ask better questions. This practical guide shows data professionals how to build text-to-SQL pipelines, enrich warehouse data with AI, and bring semantic search to their existing tables.
We'll be working on a platform made up of seventy-nine recipes together. It starts off as a simple task, printing a line, but by the last chapter it covers extraction, warehousing, containers, machine learning and incident response. You can't just throw away examples in your work, and you shouldn't be doing that in your examples either. You don't need to be an Airflow expert to get started. What you're really learning here isn't a tool. It's all about making sure work is repeatable, observable and safe to rerun.
Build real-time data systems that can keep up with the demands of production. Streaming at Scale takes you under the hood of Apache Flink, covering state, event time, exactly-once processing, performance, Kubernetes and more. With runnable examples and practical lessons, it shows how to build systems that are fast, reliable and ready to scale.
DuckDB is changing how developers think about analytics. This practical guide takes you from its architecture and SQL capabilities to performance tuning, cloud storage and production deployments. Learn how DuckDB works, where it shines and how to build fast, flexible analytical systems around it.
Passage 1 — Chapter 9, "Data Security": a definition that sets the technical tone immediately A hacker is someone skilled at finding undocumented techniques and loopholes in the tangled architecture of complex information systems; by intent, a "white hat" looks for such loopholes in order to strengthen the system, while a "black hat" uses them to steal or to cause harm. Passage 2 — Chapter 14, "Metadata Management": an image that explains the whole topic in one paragraph A vivid illustration: an enormous document archive with no index at all — the shelves are full, but there is no way to find out what sits on them short of examining every single box by hand. That is exactly what an organization looks like when it has piled up mountains of data without also taking care of metadata — the information physically exists, but it cannot be used systematically, since the only way to find what is needed is to already know where it sits. Knowledge about data is always scattered: in a large company, one person carries the structure of a single database in their head, another the rules of a single integration, a third the history of a single metric, and nobody holds the complete picture. Passage 3 — Chapter 15, "Data Quality Management, Part 1": why "quality" is an empty word without a yardstick A postal address missing an apartment number works perfectly well for a mass catalog mailing and works terribly for a courier who has to knock on the right door — and in both cases it is the very same row in the very same database, only the yardstick applied to it differs. Passage 4 — Chapter 21, "Organizational Change Management": the book's closing summary Data governance, the coordinating hub of eleven knowledge areas this book opened with, stays an empty frame until specific people come to value the new way of working with data through their own experience — which is exactly why a book that began by mapping the circle of disciplines around that hub fittingly closes not with another technique or tool, but with a conversation about the person without whose deliberate participation no structure ever becomes a practice.