Leanpub Header

Skip to main content

Filters

Category: "Data Engineering"

Data Engineering

  1. Operating Petabyte-Scale ClickHouse Clusters
    Operating Petabyte-Scale ClickHouse Clusters
    A Production Engineering Guide from Architecture to Operations
    Steve Publications

    Running ClickHouse at petabyte scale takes more than knowing SQL. This practical guide shows experienced engineers how to design, deploy and operate production clusters, from data modeling and ingestion to Kubernetes, query tuning, observability, disaster recovery and cost control, with real-world patterns and concrete configurations throughout.

  2. Cracking the System Design Interview for Data Engineers.

    Data platform interviews test CDC, streaming, warehouse modeling, and trade-offs two levels deep — no generic system-design book covers it. This one does: 50 full mock interviews, 122 rapid-fire Q&A, and the framework that turns a vague prompt into a hire.

  3. High-Performance Data Processing with Polars: From Beginner to Advanced
    High-Performance Data Processing with Polars: From Beginner to Advanced
    A Practical Guide to Fast, Scalable, Memory-Efficient Data Analytics in Python
    Steve Publications

    Discover how to make Python data processing faster, leaner and ready to scale with Polars. Starting with the basics, this practical guide takes you all the way to production-grade pipelines for massive datasets, with clear explanations of how Polars works, why it is fast and how to get the best performance from it.

  4. Mastering Julia Programming
    Mastering Julia Programming
    From First Program to Production Systems
    Steve Publications

    Discover how to build fast, elegant and reliable software with Julia. Starting from the basics, this book guides you through real-world projects, performance optimization and production-ready techniques. Whether you work in data science, engineering or research, you'll gain the skills to write Julia code with confidence.

  5. Database Engine Level Thinking
    Database Engine Level Thinking
    Building a mental model of PostgreSQL internals — written while learning
    Kareem Ashraf

    Stop only using databases — start understanding them. A living study of PostgreSQL internals for backend engineers: how pages, indexes, and WAL actually work, written clearly while learning, not after mastery.

  6. Data Exploration and Visualisation with R: From Data Quality to Credible Evidence

    Explore data before you trust it. Data Exploration and Visualisation with R shows how to uncover structure, diagnose data-quality problems, challenge misleading patterns, and turn imperfect observations into credible evidence through thoughtful, reproducible visual analysis.

  7. Deep Analysis with Pandas
    Deep Analysis with Pandas
    Transforming and Visualizing Data for Insights
    Joram Mutenge

    Learn pandas, the most widely used data analysis library.

  8. Parallel Python with Dask, Second Edition
    Parallel Python with Dask, Second Edition
    Scale pandas, NumPy, and Xarray across cores, clusters, and GPUs for terabyte-scale analytics and machine learning
    GitforGits | Asian Publishing House

    What I love about Dask is that it won't make us start from scratch. It doesn't ask us to learn a new language, adopt a new way of thinking, or rewrite five years of careful analysis into something completely different. It just takes the tools we've already got and quietly teaches them to work in parallel. The dataframe remains a dataframe. The array is still an array. The difference is that the work now spreads across eight cores, or eight hundred, and we barely had to change how we think.

  9. AI-Powered Smart Decision Making
    AI-Powered Smart Decision Making
    A Practical Guide to Using AI, Data & Analytics for Smarter Business Decisions
    Mohammad Kamrul Hassan

    AI can analyze the data. But who decides what the data means—and what to do next?Modern organizations have more information than ever, yet making good decisions remains difficult.AI-Powered Smart Decision Making shows how to combine AI, data, analytics, and human judgment to turn information into insight and insight into action.Discover a practical framework for making better decisions across strategy, marketing, sales, operations, finance, risk, and beyond.AI is not the decision. AI improves the decision process. Table of ContentsIntroduction—The New Era of AI-Powered Decision MakingChapter 1 — The Foundations of Better Business DecisionsChapter 2 — From Intuition to Data-Driven Decision MakingChapter 3 — Understanding Data for Better DecisionsChapter 4 — Business Analytics: Turning Data into InsightChapter 5 — How AI Changes Business Decision MakingChapter 6 — Building an AI-Powered Decision FrameworkChapter 7 — Asking the Right Questions: AI, Data & Decision ProblemsChapter 8 — Predictive Decision MakingChapter 9 — Prescriptive Analytics and Decision OptimizationChapter 10 — AI for Strategic Business DecisionsChapter 11 — AI for Marketing and Customer DecisionsChapter 12 — AI for Sales, Operations and Supply ChainChapter 13 — AI for Financial and Risk DecisionsChapter 14 — Human + AI: The New Decision-Making PartnershipChapter 15 — Bias, Uncertainty and the Limits of AI DecisionsChapter 16 — Trustworthy, Responsible and Explainable AI DecisionsChapter 17 — Measuring Decision Quality and Business ImpactChapter 18 — Building an AI-Powered Decision CultureChapter 19 — Designing an AI-Powered Decision SystemChapter 20 — The Future of Smart Decision MakingConclusion — From Data to Better DecisionsDisclaimer

  10. High-Performance Python for AI & Data Engineering
    High-Performance Python for AI & Data Engineering
    The Ultimate Loop, Vectorization & Memory Optimization Blueprint
    AhmedAdawy

    The ultimate technical blueprint for developers and AI engineers to strip away execution overhead, leverage hardware-accelerated NumPy/PyTorch operations, and scale Python code effortlessly

  11. Data Warehouses & LLMs

    Your data warehouse holds the answers — LLMs can help you ask better questions. This practical guide shows data professionals how to build text-to-SQL pipelines, enrich warehouse data with AI, and bring semantic search to their existing tables.

  12. Apache Airflow Cookbook
    Apache Airflow Cookbook
    Handy solutions to build, containerize and troubleshoot production ETL, ELT, MLOps and AIOps pipelines
    GitforGits | Asian Publishing House

    We'll be working on a platform made up of seventy-nine recipes together. It starts off as a simple task, printing a line, but by the last chapter it covers extraction, warehousing, containers, machine learning and incident response. You can't just throw away examples in your work, and you shouldn't be doing that in your examples either. You don't need to be an Airflow expert to get started. What you're really learning here isn't a tool. It's all about making sure work is repeatable, observable and safe to rerun.

  13. Streaming at Scale
    Streaming at Scale
    Distributed Stream Processing with Apache Flink
    Steve Publications

    Build real-time data systems that can keep up with the demands of production. Streaming at Scale takes you under the hood of Apache Flink, covering state, event time, exactly-once processing, performance, Kubernetes and more. With runnable examples and practical lessons, it shows how to build systems that are fast, reliable and ready to scale.

  14. DuckDB and the Rise of Embedded Analytical Databases
    DuckDB and the Rise of Embedded Analytical Databases
    A Comprehensive Guide to High-Performance Local Analytics
    Steve Publications

    DuckDB is changing how developers think about analytics. This practical guide takes you from its architecture and SQL capabilities to performance tuning, cloud storage and production deployments. Learn how DuckDB works, where it shines and how to build fast, flexible analytical systems around it.

  15. Data Management Engineering
    Data Management Engineering
    Architecture, Quality, Implementation
    Andrii Bogdanovych

    Passage 1 — Chapter 9, "Data Security": a definition that sets the technical tone immediately A hacker is someone skilled at finding undocumented techniques and loopholes in the tangled architecture of complex information systems; by intent, a "white hat" looks for such loopholes in order to strengthen the system, while a "black hat" uses them to steal or to cause harm. Passage 2 — Chapter 14, "Metadata Management": an image that explains the whole topic in one paragraph A vivid illustration: an enormous document archive with no index at all — the shelves are full, but there is no way to find out what sits on them short of examining every single box by hand. That is exactly what an organization looks like when it has piled up mountains of data without also taking care of metadata — the information physically exists, but it cannot be used systematically, since the only way to find what is needed is to already know where it sits. Knowledge about data is always scattered: in a large company, one person carries the structure of a single database in their head, another the rules of a single integration, a third the history of a single metric, and nobody holds the complete picture. Passage 3 — Chapter 15, "Data Quality Management, Part 1": why "quality" is an empty word without a yardstick A postal address missing an apartment number works perfectly well for a mass catalog mailing and works terribly for a courier who has to knock on the right door — and in both cases it is the very same row in the very same database, only the yardstick applied to it differs. Passage 4 — Chapter 21, "Organizational Change Management": the book's closing summary Data governance, the coordinating hub of eleven knowledge areas this book opened with, stays an empty frame until specific people come to value the new way of working with data through their own experience — which is exactly why a book that began by mapping the circle of disciplines around that hub fittingly closes not with another technique or tool, but with a conversation about the person without whose deliberate participation no structure ever becomes a practice.