Leanpub Header

Skip to main content

Reverse Engineering Databases and Storage Formats

From Binary Bytes to Complete Understanding

Reverse Engineering Databases and Storage Formats
This book is 100% completeLast updated on 2026-09-07

What really happens inside a database file? Learn to read binary data like a map, uncover hidden structures and turn raw bytes into a working understanding of how storage engines work. With hands-on examples from SQLite, PostgreSQL, InnoDB and more, this book shows you how to reverse engineer unfamiliar formats from the ground up.

Minimum price

$25.00

$35.00

You pay

Author earns

$

Also available for 1 book credit with a Reader Membership

PDF
EPUB
WEB
APP
191
Pages
About

About

About the Book

This book teaches you how to systematically reverse engineer any database or binary storage format from scratch. Starting from the fundamentals of hexadecimal analysis and binary file structures, you will progress through page-oriented storage, B-tree indexes, write-ahead logging, multi-version concurrency control, crash recovery and advanced dynamic instrumentation techniques. Every concept is illustrated with complete, runnable code examples in C, Python, Rust and Go, using real-world database formats including SQLite, PostgreSQL, InnoDB, LevelDB and Redis as case studies. By the end, you will be able to take an unfamiliar binary database file, discover its structure through controlled experimentation, document its format and implement a compatible reader or analysis tool.

Author

About the Author

Steve Publications

Steve is a technology professional with more than 20 years of experience in software development, server infrastructure, cybersecurity, vulnerability research and reverse engineering. Throughout his career, he has designed, secured, analyzed and tested complex software and infrastructure, with a particular focus on understanding how systems fail and how they can be made more secure.

Outside of work, Steve enjoys sharing knowledge with the technology community. He collaborates with researchers, industry experts and technology professionals to write practical books covering software development, cybersecurity, cloud computing, networking, DevOps, artificial intelligence and enterprise technologies. His books focus on practical learning through clear explanations, real-world examples and hands-on exercises. With more than two decades of industry experience, his goal is to help IT professionals, students and technology enthusiasts build useful skills and stay current in a rapidly changing industry.

We believe readers deserve to know how our books are created. Most of our authors are not native English speakers, so we use AI to help translate, proofread manuscripts, fix grammar, improve sentence structure and make technical explanations easier to read. AI is used as an editing tool only. It does not replace the research, technical knowledge or hands-on experience behind our books. Some of our authors also prefer to remain anonymous for privacy or professional reasons. In those cases, we publish their work under a different name. The author's name may be different, but the quality of the content and our review process remain the same.

Every book is written, reviewed and maintained by experienced technology professionals, with contributions from our private technical community of more than 420 engineers and researchers. We spend far more time validating technical accuracy and keeping our content up to date than generating text. We are always interested in working with experienced professionals who have deep expertise in a particular technology or domain. If you would like to publish a book with us or help review an existing manuscript, we'd love to hear from you. Send us a message describing your area of expertise. We are especially interested in niche technologies, specialized skills and emerging topics that are underrepresented in existing technical literature.

If you look through the contents of our books, you'll see practical examples, detailed explanations and material that is regularly updated. Our goal is to publish books that professionals can actually rely on, not low-effort AI-generated content. If you ever feel that one of our books does not meet that standard, Leanpub offers a 60-day money-back guarantee. Feel free to request a refund if you are not satisfied with your purchase.

Contents

Table of Contents

From Binary Bytes to Complete Understanding

Chapter 1: Why Reverse Engineer Storage Formats

  1. The Data Archaeology Problem
  2. Recovery and Forensics: When Backup Fails
  3. Interoperability and Lock-In Escape
  4. Security Auditing of Storage
  5. Performance Debugging at the Byte Level
  6. A Roadmap for the Book

Chapter 2: Reading Binary: First Principles

  1. Hexadecimal Notation and Byte-Level Thinking
  2. Endianness and Multi-Byte Numbers
  3. Encoding Schemes: ASCII, UTF-8 and Binary Encodings
  4. The Anatomy of a Structured Binary File
  5. Essential Command-Line Tools
  6. A First Binary Inspection
  7. Developing the Right Mindset

Chapter 3: Simple File Formats: Headers, Magic Numbers and Metadata

  1. Designing a Minimal File Format
  2. Magic Numbers and Format Identification
  3. Headers and Global Metadata
  4. Fixed-Width vs Variable-Length Records
  5. Implementing a Writer in C
  6. Parsing the Format in Python
  7. Understanding the Byte Layout
  8. Why This Matters

Chapter 4: Serialization and Data Encoding

  1. Type Systems and Storage Representation
  2. Integer and Floating-Point Encoding
  3. Strings, Text and Collation
  4. Timestamps and Temporal Types
  5. Null Handling and Optionality
  6. Case Study: SQLite Record Format (varint, serial types)

Chapter 5: Page-Oriented Storage

  1. Why Pages and Not Records
  2. Page Headers and Metadata
  3. Checksums and Corruption Detection
  4. Page Types and Classification
  5. Building a Page Viewer Tool
  6. Case Study: SQLite Page Layout

Chapter 6: Record Layouts and Slot Directories

  1. The Slot Directory Pattern
  2. Variable-Length Fields and Backward Layout
  3. Overflow Pages and Chain Pointers
  4. Header Overhead Trade-Offs
  5. Implementing a Record Decoder
  6. Case Study: InnoDB Record Formats

Chapter 7: Free-Space Management

  1. In-Page Free Space: Slabs and Linked Lists
  2. Cross-Page Allocation: Extents and Bitmaps
  3. Detecting Allocation Structures in Unknown Formats
  4. Fragmentation: Causes and Symptoms
  5. Implementing a Free-Space Analyzer
  6. Case Study: PostgreSQL Free Space Map

Chapter 8: B-Trees and Index Structures

  1. B-Tree Fundamentals and Design Trade-Offs
  2. Page Layout: Root, Interior, Leaf
  3. Traversing an Unknown B-Tree from Raw Bytes
  4. Split and Merge Dynamics
  5. Building a B-Tree Visualizer

Chapter 9: Alternative Index Structures: LSM-Trees, Hash Indexes and More

  1. Log-Structured Merge Trees: Architecture
  2. SSTable Format and Block Structure
  3. Bloom Filters and Skip Lists
  4. Hash Indexes and Direct-Address Tables
  5. Tries and Radix Trees
  6. Case Study: LevelDB File Format

Chapter 10: Write-Ahead Logging and Journaling

  1. The Write-Ahead Logging Guarantee
  2. Log Record Formats and Structure
  3. LSN Ordering and Recovery Semantics
  4. Redo and Undo: Two Phases
  5. Parsing a WAL from Raw Bytes

Chapter 11: Transactions and ACID Internals

  1. ACID Properties as Storage Requirements
  2. Transaction IDs and Commit Ordering
  3. Implementing Atomicity with Logs
  4. Durability and fsync Semantics
  5. Detecting Transaction State in Storage

Chapter 12: Concurrency Control: Locking and MVCC

  1. Locking: Tables, Pages, Rows
  2. Deadlock Detection and Prevention
  3. MVCC: Multi-Version Architecture
  4. Version Chains and Timestamps
  5. Visibility Maps and Snapshot Resolution
  6. Case Study: PostgreSQL MVCC Implementation

Chapter 13: Snapshots, Isolation Levels and Consistent Reads

  1. Snapshot Semantics and Read Views
  2. Isolation Levels: From Read Committed to Serializable
  3. Phantoms and Range Queries
  4. Serializable Snapshot Isolation
  5. Detecting Snapshot Mechanisms in Storage

Chapter 14: Crash Recovery: From Failure to Consistency

  1. What Can Go Wrong During a Crash
  2. The ARIES Recovery Algorithm
  3. Analysis Phase: Building the Transaction Table
  4. Redo Phase: Reapplying Changes
  5. Undo Phase: Rolling Back Uncommitted Work
  6. Checkpointing Strategies
  7. Manual Recovery Techniques

Chapter 15: Caching and Buffer Pools

  1. The Buffer Pool Architecture
  2. Page Lifecycle: Read, Dirty, Write-Back
  3. Eviction Policies: LRU and Variants
  4. Clean vs Dirty Page Semantics
  5. Observing Cache Behavior on Disk

Chapter 16: Compaction and Fragmentation

  1. Why Fragmentation Happens
  2. In-Place vs Append-Only Compaction
  3. LSM-Tree Compaction Strategies
  4. Tombstones and Garbage Collection
  5. Measuring Fragmentation in Storage
  6. Case Study: RocksDB Compaction

Chapter 17: Compression in Storage Engines

  1. Why Compress and When Not To
  2. Page-Level vs Columnar Compression
  3. Dictionary Encoding and Direct Coding
  4. Run-Length and Delta Encoding
  5. LZ Algorithms and Block Compression
  6. Decompressing Unknown Compressed Pages

Chapter 18: Corruption Detection and Forensic Analysis

  1. Types of Corruption: Logical vs Physical
  2. Checksum Failures and Torn Pages
  3. Detecting Index-Data Inconsistencies
  4. Orphaned Pages and Lost Records
  5. Forensic Imaging and Analysis Techniques

Chapter 19: Tools of the Trade

  1. Hex Editors and Binary Viewers
  2. Command-Line Binary Analysis Tools
  3. Debuggers for Storage Investigation
  4. System Call Tracing with strace and bpftrace
  5. Building Custom Analysis Scripts
  6. Setting Up a Reverse Engineering Lab

Chapter 20: Dynamic Instrumentation and Runtime Analysis

  1. Tracing Page I/O Operations
  2. LD_PRELOAD and Library Injection
  3. eBPF for System-Level Observability
  4. Debug Builds and Internal Logging
  5. Correlating Runtime Behavior with Storage

Chapter 21: Fuzzing, Differential Analysis and Black-Box Experimentation

  1. Fuzzing Storage Engines and Parsers
  2. Differential Analysis of File States
  3. Controlled Mutation Experiments
  4. Statistical Pattern Recognition
  5. Black-Box Protocol Discovery

Chapter 22: The Reverse Engineering Workflow

  1. Phase 1: Initial Reconnaissance
  2. Phase 2: Hypothesis Formation
  3. Phase 3: Controlled Experimentation
  4. Phase 4: Decoding Structures
  5. Phase 5: Validation and Documentation
  6. Phase 6: Implementing a Compatible Parser
  7. The Hypothesis Journal Pattern

Chapter 23: Case Study: End-to-End SQLite Reverse Engineering

  1. Initial Inspection of a Fresh Database
  2. Discovering the Database Header
  3. Mapping the Page Structure
  4. Decoding B-Tree Pages
  5. Understanding Record Serialization
  6. Uncovering the WAL Format
  7. Lessons and Techniques Demonstrated

Chapter 24: Case Studies: Other Formats and Techniques

  1. InnoDB Tablespace and Page Structure
  2. PostgreSQL Heap Tuples and Catalog
  3. LevelDB SSTable and Manifest
  4. Redis RDB Persistence Format
  5. Common Patterns Across Systems

Chapter 25: Ethics, Legality and Best Practices

  1. Copyright and Reverse Engineering Law
  2. DMCA and Circumvention Questions
  3. Interoperability and Fair Use
  4. Data Privacy and Sensitivity
  5. Responsible Disclosure of Findings
  6. Best Practices for Professional Work

Conclusion: Beyond the Known Formats

References

Get the free sample chapters

Click the buttons to get the free sample in PDF or EPUB, or read the sample online here

The Leanpub 60 Day 100% Happiness Guarantee

Within 60 days of purchase you can get a 100% refund on any Leanpub purchase, in two clicks.

See full terms...

Earn $8 on a $10 Purchase, and $16 on a $20 Purchase

We pay 80% royalties on purchases of $7.99 or more, and 80% royalties minus a 50 cent flat fee on purchases between $0.99 and $7.98. You earn $8 on a $10 sale, and $16 on a $20 sale. So, if we sell 5000 non-refunded copies of your book for $20, you'll earn $80,000.

(Yes, some authors have already earned much more than that on Leanpub.)

In fact, authors have earned over $15 million writing, publishing and selling on Leanpub.

Learn more about writing on Leanpub

Free Updates. DRM Free.

If you buy a Leanpub book, you get free updates for as long as the author updates the book! Many authors use Leanpub to publish their books in-progress, while they are writing them. All readers get free updates, regardless of when they bought the book or how much they paid (including free).

Most Leanpub books are available in PDF (for computers) and EPUB (for phones, tablets and Kindle). The formats that a book includes are shown at the top right corner of this page.

Finally, Leanpub books don't have any DRM copy-protection nonsense, so you can easily read them on any supported device.

Learn more about Leanpub's ebook formats and where to read them

Write and Publish on Leanpub

You can use Leanpub to easily write, publish and sell in-progress and completed ebooks and online courses!

Leanpub is a powerful platform for serious authors, combining a simple, elegant writing and publishing workflow with a store focused on selling in-progress ebooks.

Leanpub is a magical typewriter for authors: just write in plain text, and to publish your ebook, just click a button. (Or, if you are producing your ebook your own way, you can even upload your own PDF and/or EPUB files and then publish with one click!) It really is that easy.

Learn more about writing on Leanpub