Leanpub Header

Skip to main content

Machine Learning for Market Prediction

Feature Engineering and Backtesting Without Lookahead Bias in Python.

Machine Learning for Market Prediction
This book is 100% completeLast updated on 2026-10-10

Your Sharpe ratio is 2.8. Your model is wrong.

Not because the code is broken, but because somewhere between the raw data and the equity curve, the future leaked into the past. A centered rolling window. A normalization fit on the full sample. A fundamentals join on the wrong date. Each one runs clean, and each one produces a number that disappears the week real money replaces history.

Machine Learning for Market Prediction is a quant reviewer’s playbook for building a pipeline you can actually trust: point in time data, leakage-free features, triple barrier labels, embargoed walk-forward validation, realistic costs, and SHAP audits that trace a suspicious feature to its root cause.

It will not make your model more profitable. It will tell you the truth about it.

Distrust your backtest. Then prove it right.

Minimum price

$29.00

$49.00

You pay

Author earns

$

Also available for 2 book credits with a Reader Membership

PDF
About

About

About the Book

Learn to eliminate silent backtest inflation with confidence using measured, verified engineering, guided by a pragmatic quant researcher's playbook for point in time data, leakage-free features, and honest validation.

Key Features

  • Master core reliability disciplines used across real reviewed backtests, from data sourcing to walk-forward validation
  • Learn to eliminate silent leakage, including lookahead bias, survivorship bias, label leakage, and search-driven overfitting
  • Build provable validity without assuming correctness, using deflated Sharpe ratios, permutation testing, and SHAP-based audits

Book Description

Most traders ship a backtest once it produces a strong Sharpe ratio, unaware the pipeline never checked whether that number is real. This book closes that gap, providing the practical knowledge of point in time data sourcing, leakage-free feature engineering, and honest validation needed to build a model with measured correctness, not assumed correctness.

You will begin by sourcing and cleaning market data at the standard real backtests require, joining fundamentals on release date rather than report date, and reconstructing point in time universe membership that includes securities later delisted. You'll gain a clear framework for engineering out silent leakage at its source, understanding where centered rolling windows, full sample normalization, and untested labeling introduce risk, and how each cost compounds, or gets eliminated, under a properly embargoed pipeline.

The book walks through the full arc of a production research process: feature construction from price, volume, and sentiment data that never touches the future, triple barrier labeling scaled to genuine volatility, walk-forward validation with an embargo sized to the labeling horizon, and a full backtest rebuilt with realistic execution and cost. It addresses evaluation rigor, including deflated Sharpe ratios and regime-conditional stress testing, and closes with a full case study from raw data to a validated model.

Whether you're building your first trading model or auditing your fourth, this book provides actionable techniques for securing correctness that holds under live capital.

What You Will Learn

  • Define lookahead bias, survivorship bias, and overfitting precisely and distinguish them from ordinary noise
  • Build real point in time data pipelines without relying on assumed-safe vendor joins
  • Deconstruct backtest performance into signal, cost, and search-inflated components
  • Evaluate fixed horizon and triple barrier labeling before committing to a target
  • Verify correctness against deflated Sharpe ratios and permutation testing before trusting a result
  • Recognize and eliminate silent leakage sources without guessing at root cause
  • Know when a model's edge is proven versus merely untested
  • Design validation for a noisy, regime-shifting live market environment

Who This Book Is For

This book is for Python developers, quantitative researchers, and self-taught traders building or hardening a production trading model. It's useful for engineers moving into research-driven roles where validation guarantees grow more complex. No prior quantitative finance experience is assumed.

Table of Contents

  1. Why Most Trading Models Fail Before Deployment
  2. Sourcing and Cleaning Market Data
  3. Feature Engineering from Price Data
  4. Feature Engineering from Alternative Data
  5. Labeling the Prediction Target
  6. Choosing and Training the Model
  7. Backtesting Without Fooling Yourself
  8. Evaluating Model Performance Honestly
  9. Feature Importance and Model Interpretability
  10. Case Study: From Raw Data to a Validated Model; Appendices

Author

About the Author

Weston Ashgrove

Weston Ashgrove is a C++ quant trading developer who builds the parts of a trading system that cannot be allowed to be wrong: order books, matching logic, and the concurrency and memory-layout work underneath them. He has a simple rule for latency-critical code: nothing is true until it has been profiled, fuzzed, and replayed. That rule runs through this book, from the first order book benchmark to byte-identical crash recovery. He writes for engineers who ship and debug these systems, and he skips the theory that never survives contact with real order flow.https://x.com/dxled_dc

Get the free sample chapters

Click the buttons to get the free sample in PDF or EPUB, or read the sample online here

The Leanpub 60 Day 100% Happiness Guarantee

Within 60 days of purchase you can get a 100% refund on any Leanpub purchase, in two clicks.

See full terms...

Earn $8 on a $10 Purchase, and $16 on a $20 Purchase

We pay 80% royalties on purchases of $7.99 or more, and 80% royalties minus a 50 cent flat fee on purchases between $0.99 and $7.98. You earn $8 on a $10 sale, and $16 on a $20 sale. So, if we sell 5000 non-refunded copies of your book for $20, you'll earn $80,000.

(Yes, some authors have already earned much more than that on Leanpub.)

In fact, authors have earned over $15 million writing, publishing and selling on Leanpub.

Learn more about writing on Leanpub

Free Updates. DRM Free.

If you buy a Leanpub book, you get free updates for as long as the author updates the book! Many authors use Leanpub to publish their books in-progress, while they are writing them. All readers get free updates, regardless of when they bought the book or how much they paid (including free).

Most Leanpub books are available in PDF (for computers) and EPUB (for phones, tablets and Kindle). The formats that a book includes are shown at the top right corner of this page.

Finally, Leanpub books don't have any DRM copy-protection nonsense, so you can easily read them on any supported device.

Learn more about Leanpub's ebook formats and where to read them

Write and Publish on Leanpub

You can use Leanpub to easily write, publish and sell in-progress and completed ebooks and online courses!

Leanpub is a powerful platform for serious authors, combining a simple, elegant writing and publishing workflow with a store focused on selling in-progress ebooks.

Leanpub is a magical typewriter for authors: just write in plain text, and to publish your ebook, just click a button. (Or, if you are producing your ebook your own way, you can even upload your own PDF and/or EPUB files and then publish with one click!) It really is that easy.

Learn more about writing on Leanpub