Leanpub Header

Skip to main content

Forecasting Metrics That Don’t Lie

Choosing, Testing, and Monitoring Forecast Metrics for Demand, Inventory, and Risk Decisions

Forecasting Metrics That Don’t Lie
This book is 100% completeLast updated on 2026-09-19

A model that wins on MAPE can lose money on the shelf. A 15% accuracy gain can hide a 12% bias. An interval that looks tight can be wrong a third of the time.

Forecast metrics are decision instruments, not neutral truth detectors — and most teams are still choosing them by habit. This book shows where MAE, MAPE, sMAPE, WAPE, MASE, RMSSE, CRPS, pinball loss, energy scores, bias measures and cost-based metrics help, and where they quietly mislead.

Twelve chapters take you from elicitation and baselines through intermittent demand, probabilistic and hierarchical scoring, and the accuracy–utility gap, to a production evaluation system with monitoring, drift detection, alerting and governance. Every metric arrives with its assumptions, every criticism carries a remedy, and a retrospective Walmart M5 case study runs throughout, with Python listings and reproducibility checks.

For demand planners, forecasting data scientists, ML engineers and analytics leads who need evaluation choices they can defend.

Measure what matters.

Minimum price

$64.95

$70.00

You pay

Author earns

$

Also available for 2 book credits with a Reader Membership

PDF
About

About

About the Book

Forecast metrics are decision instruments, not neutral truth detectors. A model that
wins on MAPE can lose money on the shelf. A 15% accuracy improvement can hide a 12%
bias. A reconciled hierarchy can look coherent and still be worse at every level that
matters. This book is about the gap between the number on the scorecard and the
decision it is supposed to support — and how to close it.

Rather than cataloguing formulas, it works one question at a time: what does this
metric actually elicit, what does it reward, where does it break, and what should you
use instead? Every metric is defined with its assumptions stated, every criticism
carries a remedy, and every chapter opens with a failure that really happens in
production.

What the book covers

Foundations of honest evaluation. Why metrics mislead, the three types of drift,
and a decision-centred taxonomy that replaces the usual alphabet soup. Temporal
validation and why train/test splits lie in time series. Naive baselines as the floor
your model must beat, and statistical tests for whether a difference is real at all.
Absolute error metrics (MAE, MSE, RMSE) treated through elicitation — the median
versus the mean, the objective mismatch problem, and ranking reversals across metrics.

Scale, bias, and intermittent demand. MAPE’s fatal asymmetry, why sMAPE made it
worse, and where WAPE and MAAPE actually help. Scaled metrics — MASE, RMSSE, WRMSSE —
including the denominator trap that quietly invalidates published comparisons, plus
RMSSE-B, an author proposal for scaling against the business benchmark rather than a
statistical one. Bias metrics, tracking signals, the bias–accuracy decomposition, and
Forecast Value Added for judging whether human overrides earn their keep. A full
treatment of intermittent demand: ADI and the Syntetos–Boylan classification, Periods
in Stock, lead-time demand, SPEC, and Croston/SBA/TSB.

Distributions, dependence, and decisions. Calibration and sharpness, PIT
histograms and reliability diagrams, pinball loss, CRPS, interval and Winkler scores,
log score, Kupiec and Christoffersen tests, and conformal prediction’s coverage
guarantee under exchangeability. Multivariate and hierarchical scoring with energy and
variogram scores, coherence metrics, temporal hierarchies and copula-based evaluation.
Then the decision layer: the accuracy–utility gap, ranked probability score,
information ratio, newsvendor cost and regret, service levels and fill rate.

Production evaluation systems. Drift and failure modes, monitoring with PSI and
CUSUM, forecast stability and revision volatility, alerting thresholds, triage from
alarm to diagnosis, and retraining triggers. The book closes by assembling everything
into a system: metric bundles, scorecards, domain-specific bundles for supply chain,
energy and finance, evaluation governance, and a maturity model.

How it teaches

A retrospective Walmart M5 case study runs through the whole book, so the same data
is re-examined as the metrics get more demanding. Controlled examples isolate single
failure modes. Python listings are written to be read as well as run, and the
companion supplies code, data-acquisition instructions and reproducibility checks —
you obtain the source data separately under its own terms. The case study’s
cohort-selection limitations are stated explicitly rather than glossed over.

A metric specification sheet sits in the front matter as the book’s single sign
and denominator contract; every later chapter is written against it, so a disagreement
about a sign has one place to be settled.

Who it is for

Demand planners, forecasting data scientists, ML engineers, analytics leads and
researchers who need evaluation choices that are explicit, defensible and survivable
under review. Formulas are given in full, but the book is written for people who have
to justify a metric to a business, not only to a journal.

Measure what matters.

Author

About the Author

Valery Manokhin

Valery Manokhin (PhD, MBA, CQF) is a data scienstist, machine learning researcher and book author. He earned his PhD in ML at Royal Holloway, University of London, under Prof. Vladimir Vovk, the creator of Conformal Prediction, and holds an MBA from Warwick and an MSc in Computational Statistics & Machine Learning from UCL.

His technical books — Mastering Modern Time Series Forecasting and Applied Conformal Prediction among them — are used by data scientists, ML engineers, and researchers in over 100+ countries.

In parallel, he produces English editions of classic Russian mathematics textbooks, beginning with Kiselev’s Arithmetic and now Algebra, Part I. The aim is straightforward: give English-speaking students access to the books that shaped generations of Russian mathematicians, in editions that match the originals’ rigour.

Contents

Table of Contents

TABLE OF CONTENTS Forecasting Metrics That Don't Lie Choosing, Testing, and Monitoring Forecast Metrics for Demand, Inventory, and Risk Decisions Valery Manokhin PhD, MBA, CQF FRONT MATTER • About This Book • Reader Guide • Metric Specification Sheet — the book's sign and denominator contract PART IFOUNDATIONS Foundations of Honest Evaluation CHAPTER 1 Why Forecasting Metrics Lie • Why Forecasting Metrics Matter • The Generalisation Paradox • A Decision-Centred Taxonomy of Forecasting Metrics • How the M-Competitions Shaped Metric Practice • The Three Types of Drift • How Metrics Mislead: Pitfalls, Failure Modes, and Mitigations • Aligning Metrics with Business Goals • Our Running Example: Walmart M5 • Chapter Summary and Looking Ahead CHAPTER 2 The Rules of the Game — Evaluation Design and Baselines • The Forecast That Won on Paper and Lost in Production • Temporal Validation: Why Train/Test Splits Lie in Time Series • Naive Baselines: The Floor Your Model Must Beat • Statistical Testing: Is the Difference Real? • Evaluation Pitfalls and How to Avoid Them • Designing Your Evaluation: A Practitioner Checklist • Running Example: M5 Evaluation Setup • Summary CHAPTER 3 Absolute Error Metrics — MAE, MSE, and RMSE • The Objective Mismatch • Errors, Losses, and Elicited Functionals • MAE: The Median Metric • MSE and RMSE: The Mean Metric • The Objective Mismatch Problem • Ranking Non-Preservation: Ranking Reversals Across Metrics • Python Implementation • The Four Questions: Decision Table for Absolute Error Metrics • Summary PART IISCALE AND BIAS Scale, Bias, and Intermittent Demand CHAPTER 4 Percentage Metrics — MAPE, sMAPE, WAPE, and Why They Lie • How MAPE Picked the Wrong Model • MAPE: The Fatal Flaw • sMAPE: The “Fix” That Made Things Worse • WAPE: The Least-Bad Percentage Metric • MAAPE, Log Accuracy Ratio, and Other Alternatives • Python Implementation • Decision Framework: When (If Ever) to Use Percentage Metrics • What to Use • Summary CHAPTER 5 Scaled Metrics — MASE, RMSSE, and the Benchmarking Revolution • Why MAE = 10 Tells You Nothing • MASE: Definition, Properties, and the Denominator Trap • RMSSE: From MASE to Squared Scaling • WRMSSE: The M5 Competition Metric • RMSSE-B: Scaling Against the Business Benchmark — An Author Proposal • WSPL: Weighted Scaled Pinball Loss • Reconciliation vs. Accuracy: The M5 Lesson • Other Relative and Skill Metrics • Python Implementation • When Scaled Metrics Fail • The Four Questions: Decision Table for Scaled Metrics • What to Use • Summary CHAPTER 6 Bias and Forecast Value Added • The 15% Accuracy That Hid a 12% Bias • Bias Metrics: Measuring Systematic Direction • The Bias–Accuracy Decomposition • Forecast Value Added (FVA) • Detecting and Correcting Bias • The Four Questions: Decision Table for Bias and Forecast Value Added • What to Use • Summary CHAPTER 7 Intermittent Demand Metrics • The Spare Parts Disaster • What Makes Demand Intermittent? • Why Standard Metrics Fail on Intermittent Demand • Periods in Stock (PIS) • Lead-Time Demand: Scoring the Quantity the Decision Actually Uses • MAAPE: A Bounded Alternative to MAPE • Matching Metrics to Demand Classification • Forecasting Methods for Intermittent Demand • A Complete Evaluation Framework for Intermittent Demand • SPEC: Stock-Keeping Oriented Prediction Error Costs • Python Implementation • What to Use • Summary PART IIIDISTRIBUTIONS AND DECISIONS Distributions, Dependence, and Decisions CHAPTER 8 Probabilistic Forecast Evaluation • The Overconfident Interval • Why Evaluate Probabilistic Forecasts? • Calibration and Sharpness • Pinball Loss (Quantile Score) • CRPS: The Continuous Ranked Probability Score • Interval Score • Log Score (Ignorance Score) • Formal Calibration Tests • Skill Scores for Probabilistic Forecasts • Conformal Prediction: Guarantees Under Exchangeability • Comparing Probabilistic Forecasts • Python Implementation • What to Use • Summary CHAPTER 9 Multivariate and Hierarchical Scoring • The Reconciliation Illusion • Why Score Multivariate Forecasts? • Energy Score • Variogram Score • Discrimination Ability of Multivariate Scores • Coherence Metrics for Hierarchical Forecasts • Temporal Hierarchies • Weighted Scores Across Portfolios • Copula-Based Evaluation • What to Use CHAPTER 10 Decision and Utility Metrics • The Accurate Forecast That Lost Money • The Accuracy–Utility Gap • Ranked Probability Score (RPS) • Information Ratio • Cost-Based Evaluation • Service Level Metrics • Directional and Classification Metrics • Bridging Accuracy and Utility • What to Use PART IVPRODUCTION Production Evaluation Systems CHAPTER 11 Production Reliability — Stability, Drift, and Monitoring • The Model That Rotted Silently • Production Change and Failure Modes • Monitoring Metrics • Forecast Stability • Alerting and Thresholds • From Alarm to Diagnosis: The Triage Step • Retraining Triggers • Python Implementation • What to Use CHAPTER 12 Building Your Evaluation System • From Metrics to Systems • The Metric Bundle Framework • Metric Scorecards • Domain-Specific Evaluation Bundles • Evaluation Governance • The Evaluation Maturity Model • Python: Complete Evaluation Pipeline • Practitioner Checklist • The Book in One Page BACK MATTER • References • Code, Data, Errata, and Acknowledgments • Index Measure what matters.

The Leanpub 60 Day 100% Happiness Guarantee

Within 60 days of purchase you can get a 100% refund on any Leanpub purchase, in two clicks.

See full terms...

Earn $8 on a $10 Purchase, and $16 on a $20 Purchase

We pay 80% royalties on purchases of $7.99 or more, and 80% royalties minus a 50 cent flat fee on purchases between $0.99 and $7.98. You earn $8 on a $10 sale, and $16 on a $20 sale. So, if we sell 5000 non-refunded copies of your book for $20, you'll earn $80,000.

(Yes, some authors have already earned much more than that on Leanpub.)

In fact, authors have earned over $15 million writing, publishing and selling on Leanpub.

Learn more about writing on Leanpub

Free Updates. DRM Free.

If you buy a Leanpub book, you get free updates for as long as the author updates the book! Many authors use Leanpub to publish their books in-progress, while they are writing them. All readers get free updates, regardless of when they bought the book or how much they paid (including free).

Most Leanpub books are available in PDF (for computers) and EPUB (for phones, tablets and Kindle). The formats that a book includes are shown at the top right corner of this page.

Finally, Leanpub books don't have any DRM copy-protection nonsense, so you can easily read them on any supported device.

Learn more about Leanpub's ebook formats and where to read them

Write and Publish on Leanpub

You can use Leanpub to easily write, publish and sell in-progress and completed ebooks and online courses!

Leanpub is a powerful platform for serious authors, combining a simple, elegant writing and publishing workflow with a store focused on selling in-progress ebooks.

Leanpub is a magical typewriter for authors: just write in plain text, and to publish your ebook, just click a button. (Or, if you are producing your ebook your own way, you can even upload your own PDF and/or EPUB files and then publish with one click!) It really is that easy.

Learn more about writing on Leanpub