Running ClickHouse at petabyte scale takes more than knowing SQL. This practical guide shows experienced engineers how to design, deploy and operate production clusters, from data modeling and ingestion to Kubernetes, query tuning, observability, disaster recovery and cost control, with real-world patterns and concrete configurations throughout.
Data platform interviews test CDC, streaming, warehouse modeling, and trade-offs two levels deep — no generic system-design book covers it. This one does: 50 full mock interviews, 122 rapid-fire Q&A, and the framework that turns a vague prompt into a hire.
Discover how to make Python data processing faster, leaner and ready to scale with Polars. Starting with the basics, this practical guide takes you all the way to production-grade pipelines for massive datasets, with clear explanations of how Polars works, why it is fast and how to get the best performance from it.
Discover how to build fast, elegant and reliable software with Julia. Starting from the basics, this book guides you through real-world projects, performance optimization and production-ready techniques. Whether you work in data science, engineering or research, you'll gain the skills to write Julia code with confidence.
Stop only using databases — start understanding them. A living study of PostgreSQL internals for backend engineers: how pages, indexes, and WAL actually work, written clearly while learning, not after mastery.
Explore data before you trust it. Data Exploration and Visualisation with R shows how to uncover structure, diagnose data-quality problems, challenge misleading patterns, and turn imperfect observations into credible evidence through thoughtful, reproducible visual analysis.
Learn pandas, the most widely used data analysis library.
What I love about Dask is that it won't make us start from scratch. It doesn't ask us to learn a new language, adopt a new way of thinking, or rewrite five years of careful analysis into something completely different. It just takes the tools we've already got and quietly teaches them to work in parallel. The dataframe remains a dataframe. The array is still an array. The difference is that the work now spreads across eight cores, or eight hundred, and we barely had to change how we think.
AI can analyze the data. But who decides what the data means—and what to do next?Modern organizations have more information than ever, yet making good decisions remains difficult.AI-Powered Smart Decision Making shows how to combine AI, data, analytics, and human judgment to turn information into insight and insight into action.Discover a practical framework for making better decisions across strategy, marketing, sales, operations, finance, risk, and beyond.AI is not the decision. AI improves the decision process. Table of ContentsIntroduction—The New Era of AI-Powered Decision MakingChapter 1 — The Foundations of Better Business DecisionsChapter 2 — From Intuition to Data-Driven Decision MakingChapter 3 — Understanding Data for Better DecisionsChapter 4 — Business Analytics: Turning Data into InsightChapter 5 — How AI Changes Business Decision MakingChapter 6 — Building an AI-Powered Decision FrameworkChapter 7 — Asking the Right Questions: AI, Data & Decision ProblemsChapter 8 — Predictive Decision MakingChapter 9 — Prescriptive Analytics and Decision OptimizationChapter 10 — AI for Strategic Business DecisionsChapter 11 — AI for Marketing and Customer DecisionsChapter 12 — AI for Sales, Operations and Supply ChainChapter 13 — AI for Financial and Risk DecisionsChapter 14 — Human + AI: The New Decision-Making PartnershipChapter 15 — Bias, Uncertainty and the Limits of AI DecisionsChapter 16 — Trustworthy, Responsible and Explainable AI DecisionsChapter 17 — Measuring Decision Quality and Business ImpactChapter 18 — Building an AI-Powered Decision CultureChapter 19 — Designing an AI-Powered Decision SystemChapter 20 — The Future of Smart Decision MakingConclusion — From Data to Better DecisionsDisclaimer
The ultimate technical blueprint for developers and AI engineers to strip away execution overhead, leverage hardware-accelerated NumPy/PyTorch operations, and scale Python code effortlessly
Your data warehouse holds the answers — LLMs can help you ask better questions. This practical guide shows data professionals how to build text-to-SQL pipelines, enrich warehouse data with AI, and bring semantic search to their existing tables.
We'll be working on a platform made up of seventy-nine recipes together. It starts off as a simple task, printing a line, but by the last chapter it covers extraction, warehousing, containers, machine learning and incident response. You can't just throw away examples in your work, and you shouldn't be doing that in your examples either. You don't need to be an Airflow expert to get started. What you're really learning here isn't a tool. It's all about making sure work is repeatable, observable and safe to rerun.
Build real-time data systems that can keep up with the demands of production. Streaming at Scale takes you under the hood of Apache Flink, covering state, event time, exactly-once processing, performance, Kubernetes and more. With runnable examples and practical lessons, it shows how to build systems that are fast, reliable and ready to scale.
DuckDB is changing how developers think about analytics. This practical guide takes you from its architecture and SQL capabilities to performance tuning, cloud storage and production deployments. Learn how DuckDB works, where it shines and how to build fast, flexible analytical systems around it.
Passage 1 — Chapter 9, "Data Security": a definition that sets the technical tone immediately A hacker is someone skilled at finding undocumented techniques and loopholes in the tangled architecture of complex information systems; by intent, a "white hat" looks for such loopholes in order to strengthen the system, while a "black hat" uses them to steal or to cause harm. Passage 2 — Chapter 14, "Metadata Management": an image that explains the whole topic in one paragraph A vivid illustration: an enormous document archive with no index at all — the shelves are full, but there is no way to find out what sits on them short of examining every single box by hand. That is exactly what an organization looks like when it has piled up mountains of data without also taking care of metadata — the information physically exists, but it cannot be used systematically, since the only way to find what is needed is to already know where it sits. Knowledge about data is always scattered: in a large company, one person carries the structure of a single database in their head, another the rules of a single integration, a third the history of a single metric, and nobody holds the complete picture. Passage 3 — Chapter 15, "Data Quality Management, Part 1": why "quality" is an empty word without a yardstick A postal address missing an apartment number works perfectly well for a mass catalog mailing and works terribly for a courier who has to knock on the right door — and in both cases it is the very same row in the very same database, only the yardstick applied to it differs. Passage 4 — Chapter 21, "Organizational Change Management": the book's closing summary Data governance, the coordinating hub of eleven knowledge areas this book opened with, stays an empty frame until specific people come to value the new way of working with data through their own experience — which is exactly why a book that began by mapping the circle of disciplines around that hub fittingly closes not with another technique or tool, but with a conversation about the person without whose deliberate participation no structure ever becomes a practice.