From Binary Bytes to Complete Understanding
Chapter 1: Why Reverse Engineer Storage Formats
- The Data Archaeology Problem
- Recovery and Forensics: When Backup Fails
- Interoperability and Lock-In Escape
- Security Auditing of Storage
- Performance Debugging at the Byte Level
- A Roadmap for the Book
Chapter 2: Reading Binary: First Principles
- Hexadecimal Notation and Byte-Level Thinking
- Endianness and Multi-Byte Numbers
- Encoding Schemes: ASCII, UTF-8 and Binary Encodings
- The Anatomy of a Structured Binary File
- Essential Command-Line Tools
- A First Binary Inspection
- Developing the Right Mindset
Chapter 3: Simple File Formats: Headers, Magic Numbers and Metadata
- Designing a Minimal File Format
- Magic Numbers and Format Identification
- Headers and Global Metadata
- Fixed-Width vs Variable-Length Records
- Implementing a Writer in C
- Parsing the Format in Python
- Understanding the Byte Layout
- Why This Matters
Chapter 4: Serialization and Data Encoding
- Type Systems and Storage Representation
- Integer and Floating-Point Encoding
- Strings, Text and Collation
- Timestamps and Temporal Types
- Null Handling and Optionality
- Case Study: SQLite Record Format (varint, serial types)
Chapter 5: Page-Oriented Storage
- Why Pages and Not Records
- Page Headers and Metadata
- Checksums and Corruption Detection
- Page Types and Classification
- Building a Page Viewer Tool
- Case Study: SQLite Page Layout
Chapter 6: Record Layouts and Slot Directories
- The Slot Directory Pattern
- Variable-Length Fields and Backward Layout
- Overflow Pages and Chain Pointers
- Header Overhead Trade-Offs
- Implementing a Record Decoder
- Case Study: InnoDB Record Formats
Chapter 7: Free-Space Management
- In-Page Free Space: Slabs and Linked Lists
- Cross-Page Allocation: Extents and Bitmaps
- Detecting Allocation Structures in Unknown Formats
- Fragmentation: Causes and Symptoms
- Implementing a Free-Space Analyzer
- Case Study: PostgreSQL Free Space Map
Chapter 8: B-Trees and Index Structures
- B-Tree Fundamentals and Design Trade-Offs
- Page Layout: Root, Interior, Leaf
- Traversing an Unknown B-Tree from Raw Bytes
- Split and Merge Dynamics
- Building a B-Tree Visualizer
Chapter 9: Alternative Index Structures: LSM-Trees, Hash Indexes and More
- Log-Structured Merge Trees: Architecture
- SSTable Format and Block Structure
- Bloom Filters and Skip Lists
- Hash Indexes and Direct-Address Tables
- Tries and Radix Trees
- Case Study: LevelDB File Format
Chapter 10: Write-Ahead Logging and Journaling
- The Write-Ahead Logging Guarantee
- Log Record Formats and Structure
- LSN Ordering and Recovery Semantics
- Redo and Undo: Two Phases
- Parsing a WAL from Raw Bytes
Chapter 11: Transactions and ACID Internals
- ACID Properties as Storage Requirements
- Transaction IDs and Commit Ordering
- Implementing Atomicity with Logs
- Durability and fsync Semantics
- Detecting Transaction State in Storage
Chapter 12: Concurrency Control: Locking and MVCC
- Locking: Tables, Pages, Rows
- Deadlock Detection and Prevention
- MVCC: Multi-Version Architecture
- Version Chains and Timestamps
- Visibility Maps and Snapshot Resolution
- Case Study: PostgreSQL MVCC Implementation
Chapter 13: Snapshots, Isolation Levels and Consistent Reads
- Snapshot Semantics and Read Views
- Isolation Levels: From Read Committed to Serializable
- Phantoms and Range Queries
- Serializable Snapshot Isolation
- Detecting Snapshot Mechanisms in Storage
Chapter 14: Crash Recovery: From Failure to Consistency
- What Can Go Wrong During a Crash
- The ARIES Recovery Algorithm
- Analysis Phase: Building the Transaction Table
- Redo Phase: Reapplying Changes
- Undo Phase: Rolling Back Uncommitted Work
- Checkpointing Strategies
- Manual Recovery Techniques
Chapter 15: Caching and Buffer Pools
- The Buffer Pool Architecture
- Page Lifecycle: Read, Dirty, Write-Back
- Eviction Policies: LRU and Variants
- Clean vs Dirty Page Semantics
- Observing Cache Behavior on Disk
Chapter 16: Compaction and Fragmentation
- Why Fragmentation Happens
- In-Place vs Append-Only Compaction
- LSM-Tree Compaction Strategies
- Tombstones and Garbage Collection
- Measuring Fragmentation in Storage
- Case Study: RocksDB Compaction
Chapter 17: Compression in Storage Engines
- Why Compress and When Not To
- Page-Level vs Columnar Compression
- Dictionary Encoding and Direct Coding
- Run-Length and Delta Encoding
- LZ Algorithms and Block Compression
- Decompressing Unknown Compressed Pages
Chapter 18: Corruption Detection and Forensic Analysis
- Types of Corruption: Logical vs Physical
- Checksum Failures and Torn Pages
- Detecting Index-Data Inconsistencies
- Orphaned Pages and Lost Records
- Forensic Imaging and Analysis Techniques
Chapter 19: Tools of the Trade
- Hex Editors and Binary Viewers
- Command-Line Binary Analysis Tools
- Debuggers for Storage Investigation
- System Call Tracing with strace and bpftrace
- Building Custom Analysis Scripts
- Setting Up a Reverse Engineering Lab
Chapter 20: Dynamic Instrumentation and Runtime Analysis
- Tracing Page I/O Operations
- LD_PRELOAD and Library Injection
- eBPF for System-Level Observability
- Debug Builds and Internal Logging
- Correlating Runtime Behavior with Storage
Chapter 21: Fuzzing, Differential Analysis and Black-Box Experimentation
- Fuzzing Storage Engines and Parsers
- Differential Analysis of File States
- Controlled Mutation Experiments
- Statistical Pattern Recognition
- Black-Box Protocol Discovery
Chapter 22: The Reverse Engineering Workflow
- Phase 1: Initial Reconnaissance
- Phase 2: Hypothesis Formation
- Phase 3: Controlled Experimentation
- Phase 4: Decoding Structures
- Phase 5: Validation and Documentation
- Phase 6: Implementing a Compatible Parser
- The Hypothesis Journal Pattern
Chapter 23: Case Study: End-to-End SQLite Reverse Engineering
- Initial Inspection of a Fresh Database
- Discovering the Database Header
- Mapping the Page Structure
- Decoding B-Tree Pages
- Understanding Record Serialization
- Uncovering the WAL Format
- Lessons and Techniques Demonstrated
Chapter 24: Case Studies: Other Formats and Techniques
- InnoDB Tablespace and Page Structure
- PostgreSQL Heap Tuples and Catalog
- LevelDB SSTable and Manifest
- Redis RDB Persistence Format
- Common Patterns Across Systems
Chapter 25: Ethics, Legality and Best Practices
- Copyright and Reverse Engineering Law
- DMCA and Circumvention Questions
- Interoperability and Fair Use
- Data Privacy and Sensitivity
- Responsible Disclosure of Findings
- Best Practices for Professional Work