LM from First Principles
- A Notebook-Driven Guide to Building and Deploying Language Models
Chapter 1: The Data Pipeline — Text, Tensors, and GPU
- What You Will Learn
- Setup
- Getting the Corpus
- Building the Vocabulary
- The Lookup Tables
- From List to Tensor
- Train and Validation Split
- Moving to the GPU
- Building Batches
- A First Model: The Bigram
- Training
- Evaluation: Loss and Perplexity
- The Greedy Collapse
- How This Connects to Modern LMs and Inference Engineering
- Chapter Summary
- What’s Next
- The Journey
- Review Questions
Chapter 2: Vectors — The Geometry Behind Language Models
- What You Will Learn
- Setup
- 1. Revisit Chapter 1: What Was That Row of 95 Numbers?
- 2. Scalar, Vector, Matrix, Tensor
- 3. Rank, Shape, Number of Elements, and Dtype
- 4. Reading
(B, T, C) - 5. Coordinates, Magnitude, and Direction
- 6. Vector Addition and Scalar Multiplication
- 7. The Dot Product
- 8. Distance and Similarity
- 9. Matrices as Collections of Vectors
- 10. Matrix Multiplication as Transformation
- 11. Row Vectors, Column Vectors, and Orientation
- 12. Reshaping and Flattening
- 13. High-Dimensional Spaces
- 14. PyTorch Lab — Becoming Fluent in Vector Shapes
- 15. Research Connections: How Papers Enter This Chapter
- 16. Where These Vector Operations Reappear
- 17. What Chapter 2 Deliberately Does Not Explain
- Chapter Summary
- What’s Next
- Review Questions
Interlude — Can the Whole Sequence Be One Vector?
- Bigram: only one position contributes
- Flattening works
- But flattening is brittle
- Shared transformations remove the slot-specific wiring
- But the rows still cannot communicate
Chapter 3: Embeddings — From IDs to Learned Representations
- What You Will Learn
- Setup
- 1. IDs Are Indices, Not Representations
- 2.
nn.Embedding: A Learnable Lookup Table - 3. Why a Lookup Table? The One-Hot Connection
- 4. Embedding Lookup Produces Row-Local Gradients
- 5. From One Token to a Sequence:
(B, T)→(B, T, C) - 6. Building the Minimal Model
- 7. The Core Experiment: Training 2D Character Embeddings
- 8. Reading the Scatter Plot
- 9. Cosine Similarity and Nearest Neighbours
- 10. Movement Analysis
- 11. Scaling to Higher Dimensions: C=16
- 12. Three Types of Embedding
- 13. What Token Embeddings Still Do Not Know
- Research Connections
- What This Chapter Deliberately Does Not Explain
- Chapter 3 — Summary
- The Road Ahead
- Review Questions