Leanpub Header

Skip to main content

A-to-Z Transformer Walkthrough: “The capital of France is ___?”

A-to-Z Transformer Walkthrough: “The capital of France is ___?”
This book is 100% completeLast updated on 2026-08-23

Ever wondered what actually happens inside a Transformer? Follow “A-to-Z Transformer Walkthrough: “The capital of France is ___?”” through a complete miniature decoder-only Transformer—from embeddings, Q/K/V and multi-head attention to FFNs, logits, softmax, and finally “Paris.” Every major step is explained with real numbers and reproducible Python code.

Minimum price

$19.00

$29.00

You pay

Author earns

$

Also available for 1 book credit with a Reader Membership

PDF
About

About

About the Book

Transformers power many of today’s most capable AI systems—but understanding what actually happens inside a Transformer can be surprisingly difficult. Diagrams and equations often explain the individual components, yet it can still be hard to see how all those pieces work together to turn a sequence of words into a next-token prediction.

A-to-Z Transformer Walkthrough takes a different approach.

Using the simple prompt “The capital of France is ___?”, this book follows a deliberately small, transparent decoder-only Transformer from beginning to end. Rather than hiding the calculations behind high-level deep-learning libraries, it exposes the numbers and mathematics at each major stage.

You will follow the journey through:

  • tokenization and vocabulary construction
  • token embeddings
  • sinusoidal positional encoding
  • Query, Key, and Value (Q/K/V) projections
  • scaled dot-product attention
  • causal masking
  • multi-head self-attention
  • residual connections and LayerNorm
  • feed-forward networks (FFNs)
  • multiple Transformer layers
  • the final hidden representation
  • the language-model head
  • logits and softmax probabilities
  • next-token prediction
  • backpropagation and training with Adam

The walkthrough begins with randomly initialized weights, allowing you to see why an untrained Transformer does not magically know that the answer is Paris. The same miniature model is then actually trained, showing how gradient-based learning changes its parameters until Paris becomes the overwhelmingly likely next-token prediction.

The book also discusses what this tiny experiment does not prove, including memorization, shortcut learning, and the limitations of training on very small datasets.

Complete reproducible Python code is included so that you can run the model yourself and inspect the calculations.

This is not intended to be a production implementation of GPT. It is an educational, simplified GPT-style decoder-only Transformer designed for one purpose:

to make the machinery inside a Transformer tangible, traceable, and easier to understand—one step and one number at a time.

Share this book

Author

About the Author

Sajeewa Pemasinghe

Dr. Sajeewa Pemasinghe holds a PhD in Computational Modeling and Simulations from Wayne State University, USA, an MSc in IT (awarded for best performance) from the Sri Lanka Institute of Information Technology (SLIIT), and a BSc Hons degree from the University of Kelaniya. With over ten years of experience teaching at the undergraduate and diploma levels, Dr. Pemasinghe has delivered lectures on applications of IT, Bioinformatics and Computational Chemistry in Sri Lanka, the United States, and Australia. His expertise spans Artificial Intelligence, Robotics, IoT, and Bioinformatics, with a focus on integrating machine learning and robotics to develop efficient, low-cost solutions that improve quality of life.

Dr. Pemasinghe’s research centres on creating smart technologies to address challenges in agriculture, food technology, environmental conservation, and healthcare services. He is skilled in complex modelling techniques, from discrete events to molecular mechanical and quantum mechanical simulations and is committed to mentoring students in AI and robotics research. An IEEE member, Dr. Pemasinghe contributes actively to the academic community through publications and conference presentations.

The Leanpub 60 Day 100% Happiness Guarantee

Within 60 days of purchase you can get a 100% refund on any Leanpub purchase, in two clicks.

See full terms...

Earn $8 on a $10 Purchase, and $16 on a $20 Purchase

We pay 80% royalties on purchases of $7.99 or more, and 80% royalties minus a 50 cent flat fee on purchases between $0.99 and $7.98. You earn $8 on a $10 sale, and $16 on a $20 sale. So, if we sell 5000 non-refunded copies of your book for $20, you'll earn $80,000.

(Yes, some authors have already earned much more than that on Leanpub.)

In fact, authors have earned over $15 million writing, publishing and selling on Leanpub.

Learn more about writing on Leanpub

Free Updates. DRM Free.

If you buy a Leanpub book, you get free updates for as long as the author updates the book! Many authors use Leanpub to publish their books in-progress, while they are writing them. All readers get free updates, regardless of when they bought the book or how much they paid (including free).

Most Leanpub books are available in PDF (for computers) and EPUB (for phones, tablets and Kindle). The formats that a book includes are shown at the top right corner of this page.

Finally, Leanpub books don't have any DRM copy-protection nonsense, so you can easily read them on any supported device.

Learn more about Leanpub's ebook formats and where to read them

Write and Publish on Leanpub

You can use Leanpub to easily write, publish and sell in-progress and completed ebooks and online courses!

Leanpub is a powerful platform for serious authors, combining a simple, elegant writing and publishing workflow with a store focused on selling in-progress ebooks.

Leanpub is a magical typewriter for authors: just write in plain text, and to publish your ebook, just click a button. (Or, if you are producing your ebook your own way, you can even upload your own PDF and/or EPUB files and then publish with one click!) It really is that easy.

Learn more about writing on Leanpub