Building a PDF Reader from the ISO 32000-2 Specification
Introduction: Why Build a PDF Parser?
- The Ubiquity and Hidden Complexity of PDF
- What This Book Will Build
- How to Read This Book: Code Structure and Conventions
- Getting Started: Project Setup and Toolchain
Chapter 1: History, Design Goals, and Architecture of PDF
- PostScript Heritage and the Birth of Acrobat
- Design Principles: Device Independence, Portability, Interactivity
- Version Timeline from PDF 1.0 to PDF 2.0
- The Object Model: A Foundation for Everything
- File Structure at a Glance: Headers, Bodies, Cross-References, Trailers
- Building the File I/O Layer
Chapter 2: Lexical Structure and the Tokenizer
- The PDF Character Set and Whitespace Rules
- Delimiters, Brackets, and Parentheses
- Numeric Tokens: Integers and Real Numbers
- Names, Literal Strings, Hexadecimal Strings, Booleans, Null
- Comments and Escaping Rules
- Building the Tokenizer: Complete Implementation
Chapter 3: Direct Objects and Recursive-Descent Parsing
- The Parser Architecture: From Tokens to Object Trees
- Parsing Numbers and Simple Values
- Parsing Literals Strings with Escape Sequences
- Parsing Hexadecimal Strings and Length Validation
- Parsing Names and Key Resolution
- Parsing Dictionaries: Keys, Values, and Nested Structures
- Parsing Arrays and Recursive Descent Patterns
- Putting It All Together: A Simple Test
Chapter 4: Indirect Objects, Object Numbering, and Streams
- Why Indirect Objects Matter: Referencing and Random Access
- The obj-endobj Syntax and Object Numbers
- Stream Objects: Headers, Lengths, and Data Boundaries
- Basic Stream Parsing and the /Length Entry
- Building an Object Table for Indirect References
- Handling Cross-References Between Objects
Chapter 5: Cross-Reference Tables, Streams, and the Trailer
- The %%EOF Marker and Backward Search
- The Trailer Dictionary: Root, Size, Encrypt, Prev
- Traditional Cross-Reference Tables: Entry Formats and Types
- Parsing the xref Section Byte by Byte
- Cross-Reference Streams (PDF 1.5+): W, Index, Filter Entries
- Incremental Updates and Multiple xref Sections
Chapter 6: Stream Filters and Decompression
- The Filter Chain: Sequential Decompression
- ASCIIHexDecode and ASCII85Decode Implementations
- LZWDecode: TIFF Class F Algorithm Details
- FlateDecode: zlib Integration and Raw Deflate
- RunLengthDecode and CCITTFaxDecode
- DCTDecode, JBIG2Decode, JPXDecode Overview
- The Crypt Filter and Encryption Basics
- Putting It All Together: Filter Chain Application
Chapter 7: Document Catalog and Page Tree Navigation
- The Catalog Dictionary: Type, Pages, Names, Outlines
- The Pages Object: Kids, Count, Parent Relationships
- Traversing the Page Tree Recursively
- Page Objects: MediaBox, CropBox, Rotate, Resources
- Resource Dictionaries: Fonts, XObjects, ColorSpaces, Patterns
- Building a Document Navigation API
Chapter 8: Content Streams and Graphics Operators
- Content Stream Syntax and Operator Tokens
- The Current Transformation Matrix (CTM)
- Path Construction Operators: moveto, lineto, curveto
- Path Painting: Stroke, Fill, Clip Operations
- Text State and Positioning Operators
- Graphics State Saving and Restoring: q and Q
- Marked Content and Tagged PDF Basics
- Parsing Content Streams
- Inline Images in Content Streams
Chapter 9: Coordinate Systems and Page Geometry
- User Space vs. Device Space
- The Transformation Matrix Pipeline
- Clipping Paths and the Clip Operator Family
- Tiling Patterns and Pattern Paint
- Shadings: Axial, Radial, and Coons Mesh
- Page Boxes: MediaBox, CropBox, BleedBox, TrimBox, ArtBox
Chapter 10: Text Rendering and Fonts
- The Text Rendering Model: ShowText Operators
- Font Dictionaries: Type, Subtype, Encoding
- Standard Type 1 Fonts and Built-in Encodings
- TrueType and OpenType Font Embedding
- CID Fonts and CMap Parsing
- Character Spacing, Word Spacing, Leading
- Extracting Text from Content Streams
Chapter 11: Images and Color Spaces
- Image XObjects: Width, Height, BitsPerComponent
- Image Masking and Soft Masks
- Device Color Spaces: RGB, CMYK, Gray
- ICC-Based Color Spaces and Profile Handling
- Special Color Spaces: Separation, DeviceN, Indexed
- Loading and Decoding Image Data
Chapter 12: Annotations, Links, and Interactivity
- Annotation Dictionaries: Type, Subtype, Rect
- Link Annotations and URI Actions
- Text and Highlight Annotations
- Destination Types: XYZ, Fit, FitH, FitV, FitR, FitB
- Outlines (Bookmarks): Tree Structure and Navigation
- Named Destinations and the Names Dictionary
Chapter 13: Forms, AcroForms, and Interactive Fields
- The AcroForm Dictionary and Field Hierarchy
- Field Types: Text, Button, Choice, Signature
- Field Appearance Streams and Rendering
- Field States: ReadOnly, Required, Export Values
- Radio Button Groups and Checkboxes
- Form Submission Actions
Chapter 14: Security, Encryption, and Digital Signatures
- The Encrypt Dictionary and Security Handlers
- Standard Security Handler: User/Owner Passwords
- RC4 Encryption Algorithm Implementation
- AES Encryption (AESV2, AESV3)
- Integrating Encryption into the Core Architecture
- Permission Flags and Document Restrictions
- Digital Signatures: Byte Range, Certificate Chains
Chapter 15: Advanced Features in PDF 2.0
- What Changed in PDF 2.0
- Enhanced Transparency and Blending
- Improved ICC Color Management
- New Filter Capabilities and Requirements
- Structural Improvements: Object Streams Refinements
- Removed Features: XFA Deprecation
- UTF-8 Support and Text String Improvements
- Portfolio and Collection Enhancements
Chapter 16: Robustness, Performance, and Real-World Issues
- Malformed Document Recovery Strategies
- Fuzz Testing and Security Hardening
- Memory Management: Caching Indirect Objects
- Lazy Loading and On-Demand Decoding
- Random Access Patterns and File I/O Optimization
- Interoperability Issues Across PDF Generators
Conclusion: A Complete Parser, An Open Format
- What We Built: The Complete Implementation
- Lessons from Building a Standards-Compliant Parser
- Extending the Reader: Rendering and Beyond
- The Future of PDF in a Digital World