Leanpub Header

Skip to main content

The Market Is Not a List

Building answerable markets from identity, evidence, relationships, and change

This book is 100% completeLast updated on 2026-07-19

Markets are not static rows. This book shows how to turn fragmented public and private evidence into an answerable, versioned map of who exists, who fits, why now, and who to contact—without hiding uncertainty, provenance, or change.

Minimum price

$19.00

$29.00

You pay

Author earns

$

Also available for 1 book credit with a Reader Membership

PDF
EPUB
WEB
APP
About

About

About the Book

The Market Is Not a List explains how to build go-to-market data systems that answer four commercial questions: who exists, who fits, why now, and who to contact. It replaces static lead lists with a versioned evidence model spanning identity resolution, discovery, qualification, relationships, signals, contacts, provenance, rights, and change. The book follows the full production stack—from source acquisition and entity graphs through scoring, orchestration, quality measurement, and agentic workflows—and shows why trustworthy market intelligence is a maintained data product, not a one-time export. It is written for founders, data and product leaders, revenue operations teams, and engineers building commercial intelligence systems.

Author

About the Author

Contents

Table of Contents

Chapter 1 — The Market Is Not a List

  1. What This Book Means by “Market”
  2. One Commercial Request, Four Data Problems
  3. Discovery Is Not Enrichment
  4. Fit Is a Conclusion, Not a Field
  5. Durable Fit Is Not the Same as “Why Now?”
  6. Company, Account, Entity, Location and Person
  7. Evidence, Inference and Maintenance
  8. From Six CRM Rows to a Commercial Market
  9. The CRM Is Operating Memory, Not the Market Boundary
  10. Different Questions Require Different Measures
  11. The Market Is a Maintained Judgment
  12. Notes

Chapter 2 — The Anatomy of a B2B Data Record

  1. A Practical Model, Not a Perfect Ontology
  2. Every Claim Needs the Same Basic Questions
  3. Three Independent Axes
  4. Identity: Which Real Object Is This?
  5. Firmographics: Useful Simplifications
  6. Capability: What Can the Organization Actually Do?
  7. Relationships: Direction, Type, Scope and Time
  8. Events: What Changed?
  9. Behavioral Data: What Activity Was Observed?
  10. People: Person, Employment, Role and Channel
  11. Engagement: What Has the Seller Done?
  12. Evidence: The Cross-Cutting Layer
  13. Where Familiar Vendor Labels Fit
  14. Reconstructing the Medora Account
  15. The Company Row Is a Materialized View Over Claims
  16. Notes

Chapter 3 — The GTM Data Product Landscape

  1. A Product Has Four Coordinates
  2. Part A — Construct the Universe
  3. Part B — Make the Universe Intelligent and Maintain It
  4. Part C — Use the Intelligence
  5. Comparing the Product Shapes
  6. Cross-Cutting Requirements
  7. From Product Map to Product Wedge
  8. Notes

Chapter 4 — Inside the Data Factory

  1. The Factory Produces Claims, Not Rows
  2. The Factory Is a Loop, Not a Waterfall
  3. Phase I — Specify and Acquire
  4. Phase II — Interpret and Resolve
  5. Phase III — Validate and Publish
  6. Phase IV — Monitor and Correct
  7. The Hidden Workforce
  8. What Should Be Deterministic and What Should Be Semantic?
  9. The Finished Record
  10. Notes

Chapter 5 — What Good Data Means

  1. Begin With the Decision and the Unit
  2. The Quality Map
  3. Five Denominators That Must Not Be Confused
  4. Coverage: Where the Dataset Can See
  5. Precision: How Much Published Data Is Right?
  6. Recall: What Did the System Fail to Find?
  7. Quality Evaluation Has Its Own Quality
  8. Precision and Recall Are a Policy Choice
  9. Confidence, Calibration and Ranking
  10. Freshness: How Quickly Does Evidence Enter the Product?
  11. Temporal Validity: Is the Fact Still True?
  12. Completeness: Are the Necessary Pieces Present?
  13. Matchability: Can the Customer Connect the Record?
  14. Uniqueness and Consistency
  15. Provenance: Why Should the Customer Believe It?
  16. Actionability: Can Someone Use the Data?
  17. Economic Quality: What Does Each Useful Record Cost?
  18. Different Products Need Different Quality Cards
  19. The Buyer’s Quality Card
  20. Good Data Is a Portfolio of Tradeoffs
  21. Notes

Chapter 6 — Crawling the Business Web

  1. Crawling, Scraping and Extraction Are Not the Same Thing
  2. From a Known Company to an Open Web
  3. A Short Timeline of the Crawling Era
  4. The Business Web Is a Collection of Partial Witnesses
  5. The Crawl Frontier Is a Product Decision
  6. The DiscoverOrg–ZoomInfo Case: Depth Meets Breadth
  7. Web Coverage Is Not Market Coverage
  8. Failure Modes Unique to Crawling
  9. A Worked Crawl: Finding the Missing Distributor
  10. Crawling Is Not the Same as Licensing
  11. Refresh Is a Scheduling Problem, Not a Single Cadence
  12. The Economics of the Crawl
  13. What the Crawling Era Really Contributed
  14. Notes

Chapter 7 — LinkedIn and the Self-Updating Professional Database

  1. The Database Maintenance Problem LinkedIn Reframed
  2. From Network to B2B Business: A Compact Timeline
  3. From Online Résumé to Professional Graph
  4. Three Graphs, Three Jobs
  5. Why Members Maintain the Record
  6. Self-Reported Is Not the Same as Verified—and Verification Is Not Complete
  7. Job Changes Turn a Profile Edit Into a Signal
  8. The Network Changes “Who Should We Contact?”
  9. When People Data Becomes Company Data
  10. Sales Navigator: Turning the Graph Into a Sales Product
  11. Enterprise Mapping: Page, Account and Legal Entity
  12. The Graph Is Visible, but It Is Not Open
  13. What LinkedIn Changed—and What It Did Not
  14. Notes

Chapter 8 — Crowdsourced Contact Data

  1. From Jigsaw to the Hybrid Data Factory: A Compact Timeline
  2. Why Contact Data Is a Special Kind of Fact
  3. Jigsaw’s Give-to-Get Market
  4. Four Products Hidden Inside One Contact Database
  5. From Jigsaw to Data.com
  6. Crowdsourcing Is Not One Mechanism
  7. A Contact Conflict, Properly Handled
  8. How a Work Email Becomes a Data Claim
  9. A Telephone Number Is Not a Person Either
  10. The Network Effect—and Its Adverse Selection
  11. Rights, Privacy and the Represented Person
  12. What the Retirement of Data.com Does—and Does Not Prove
  13. The Modern Contributory Network
  14. Designing a Responsible Data Cooperative
  15. What Crowdsourced Contact Data Can Answer
  16. Notes

Chapter 9 — The API-First Data Company

  1. From Database Login to Software Primitive
  2. Enrichment Is Still Not Discovery
  3. An API Response Is an Operational Promise
  4. From Request to Corrected State
  5. API-First Does Not Mean API-Only
  6. Clearbit: The Data Primitive Moves Into the System of Record
  7. People Data Labs: A Family of Data Primitives
  8. Coresignal: APIs Over a Public-Web Data Factory
  9. FullContact: From Contact Enrichment to Identity Resolution
  10. Four API-First Models Compared
  11. The Interface Changes the Buyer
  12. What API-First Requires Operationally
  13. Failure Modes of API-Delivered Data
  14. An Acceptance Test for a Data API
  15. What the API Era Solved—and What It Did Not
  16. Notes

Chapter 10 — Social, Community and Public Activity Signals

  1. A Controlled Vocabulary for Public Traces
  2. Public Activity Is Not One Data Class
  3. The Anatomy of an Activity Record
  4. Different Platforms Produce Different Evidence
  5. LinkedIn: The Professional Context Layer
  6. X: The Real-Time Public Conversation
  7. GitHub: Specific Technical Evidence with Organizational Ambiguity
  8. Reddit: Rich Problem Language, Weak Enterprise Identity
  9. Product and Employer Reviews: Reported Experience, Not Neutral Sensors
  10. Slack and Discord: The Boundary Between Community Data and Private Communication
  11. Facebook and Instagram: Operational Evidence Beyond Enterprise Personas
  12. Conferences, Webinars and Event Data
  13. The Person-to-Company Inference Ladder
  14. Bots, Campaigns and Performative Activity
  15. Access Is Part of Data Quality
  16. A Worked Example: From Traces to a Bounded Account Hypothesis
  17. Measuring a Public-Activity Data Product
  18. What Public Activity Adds to the Four Questions
  19. Notes

Chapter 11 — Vertical Company Intelligence

  1. Horizontal, Vertical and Specialized-Horizontal Data
  2. What Makes a Dataset Vertical
  3. Why Vertical Data Can Command a Premium
  4. PitchBook: A Specialized-Horizontal Model of Private Capital
  5. Crunchbase: Specialized-Horizontal Distribution Becomes a Data Input
  6. CB Insights: Technology Taxonomy Plus Editorial Judgment
  7. DataFox: Specialized Company Data Finds a Workflow Owner
  8. Mattermark: A Specialized-Horizontal Caution Against Single-Cause Stories
  9. Industry Databases: When the Schema Becomes the Product
  10. Worked Example: An Investment-Adviser Intelligence Database
  11. The Role of Specialist Analysts in the GenAI Era
  12. Business Models for Vertical Intelligence
  13. The Vertical Wedge and the Expansion Trap
  14. Failure Modes Specific to Vertical Data
  15. What Vertical Intelligence Adds to the Four Questions
  16. Notes

Chapter 12 — The Account Becomes the Unit of Strategy

  1. Why the Lead Became the Unit in the First Place
  2. The Account Is a Model, Not a Natural Fact
  3. From Key Accounts to Account-Based Marketing
  4. Buying Committees and Buying Groups
  5. Selecting Target Accounts
  6. Tiering Is Capacity Allocation
  7. Account Advertising: Matching Before Persuasion
  8. Sales–Marketing Coordination Is the Product
  9. Worked Example: Meridian Bioanalytics
  10. Measuring Account Progress
  11. What ABM Solved—and What It Did Not
  12. What the Account Model Adds to the Four Questions
  13. Notes

Chapter 13 — The Rise of Intent Data

  1. The Word “Intent” Hides Different Products
  2. The Anatomy of an Intent Record
  3. Three Behavioral Source Classes—and a Separate Event Layer
  4. First-Party Intent: Closest to the Seller
  5. Second-Party Intent: Context from a Publisher or Marketplace
  6. Third-Party Aggregated Intent: Breadth Through Patterns
  7. Platforms That Combine Source Classes
  8. The Identity Problem: From Traffic to Company
  9. Topic Ambiguity
  10. Sample Size, Baselines and the Large-Account Bias
  11. Research Is Not Purchase
  12. Public Events Can Become Candidate Signals
  13. Worked Example: Northstar Components
  14. Comparing the Major Models
  15. From Signal to Proportionate Action
  16. Measuring Whether Intent Data Works
  17. Governance and Trust
  18. What Intent Data Adds to the Four Questions
  19. Notes

Chapter 14 — Predictive Lead Scoring and the Big-Data Promise

  1. From Hand-Built Points to Learned Scores
  2. Five Scores Commonly Collapsed into One
  3. The Training Set Is a History of Organizational Behavior
  4. Selection Bias and the Self-Fulfilling Score
  5. The Candidate-Universe Problem
  6. Label Design: Predict What, by When?
  7. Leakage: When the Future Enters the Past
  8. Scores Drift Because the Commercial System Changes
  9. Explainability and the Seller’s Right to Ignore
  10. No Action, No Value
  11. The Predictive-Marketing Vendor Wave
  12. What the Consolidation Means
  13. How to Evaluate a Predictive GTM Model
  14. Prediction Versus Discovery, Causation and Decision
  15. What Predictive Scoring Adds to the Four Questions
  16. Notes

Chapter 15 — From Data Vendor to GTM Orchestrator

  1. The Stack Had Data but No Shared Decision
  2. Four Operating Objects
  3. Apollo: The Database Moves Downstream
  4. Clay: The Programmable Enrichment Table
  5. Common Room: The Signal-Centric Account
  6. Product-Led Signals: Usage Is Evidence, Not a Buyer
  7. Reverse ETL: From the Warehouse Back to Work
  8. GTM Engineering: Revenue Logic Becomes Production Logic
  9. Automated Outbound: When Inference Can Send an Email
  10. Worked Example: Orchestrating One Account Across the Stack
  11. Orchestration Does Not Eliminate Data Strategy
  12. Why the Decision Surface Captures Value
  13. Measuring an Orchestrator
  14. What Orchestration Adds to the Four Questions
  15. Notes

Chapter 16 — From Advertisements to Outcomes

  1. The Units of Value
  2. The Advertiser Pays to Be Found
  3. The Buyer Pays for a Decision Report
  4. List Rental: Pay for Permissioned Use, Not Ownership
  5. Subscription and Per-Seat SaaS: Pay for an Information Environment
  6. Per Record and Per API Call
  7. Credits and Actions: A Currency for Heterogeneous Costs
  8. Cooperative Access: Contribute to the Pool
  9. OEM Rights and Cloud Delivery
  10. Workflow Bundles: Pay for the Job, Not the Field
  11. Verified Accounts and Outcome Pricing
  12. Data SaaS Does Not Have Zero Marginal Cost
  13. Refresh Is an Economic Commitment
  14. Worked Example: Six Ways to Buy the Same Market
  15. Choosing the Economic Unit
  16. What the Business Model Adds to the Four Questions
  17. Notes

Chapter 17 — What Creates a Data Moat?

  1. A Moat Must Survive the Replication Test
  2. Exclusive Access and Permitted Use
  3. User-Generated and Contributory Networks
  4. The Identifier Is More Valuable Than It Looks
  5. The Entity-Resolution History
  6. Vertical Ontology: Knowing What the Market Means
  7. Proprietary Relationship and Event Histories
  8. Refresh Infrastructure and Data Memory
  9. Evidence and Provenance as a Trust Asset
  10. Corrections and Evaluation Sets
  11. Workflow Integration and Decision History
  12. Brand, Auditability and Regulatory Position
  13. Why Public-Web Crawling Is Rarely Enough
  14. False Moats
  15. Moats Can Erode
  16. Worked Example: A Moat for a New Market
  17. A Builder’s Moat Scorecard
  18. What a Data Moat Adds to the Four Questions
  19. Notes

Chapter 18 — Distribution Often Beats the Better Dataset

  1. Data Does Not Distribute Itself
  2. The Distribution Stack
  3. Enterprise Sales Sells Confidence
  4. Product-Led Distribution Sells the First Useful Moment
  5. Developers Can Be a Channel
  6. Content Is a Preview of the Database
  7. Communities and Templates Distribute Workflows
  8. Marketplaces Put the Product Near Existing Trust
  9. OEM Embedding Lets Someone Else Own the Interface
  10. Cloud Marketplaces Distribute Procurement and Delivery
  11. Agencies and Consultants Carry the Last Mile
  12. Six Companies, Six Distribution Logics
  13. Channel–Product Fit
  14. Distribution Changes the Product
  15. Channels Can Conflict
  16. The Distribution–Data Flywheel
  17. Measure Distribution as a Funnel
  18. Worked Example: Distributing a Contractor Market
  19. A Distribution Scorecard
  20. What Distribution Adds to the Four Questions
  21. Notes

Chapter 19 — Acquisition Is Not a Single Definition of Success

  1. The Acquisition Scoreboard
  2. The Headline Price Is Not the Payout
  3. Announcements State Intentions
  4. Why GTM-Data Companies Are Acquired
  5. Jigsaw and Salesforce: The Product Can End After the Capability Spreads
  6. Eloqua and Oracle: Brand Survival With Long Product Continuity
  7. DataFox and Oracle: Capability Survival Is Harder to Trace
  8. ExactTarget, Pardot and Salesforce: A Nested Acquisition
  9. Marketo and Adobe: Preservation as a Strategic Choice
  10. LinkedIn and Microsoft: Independence Can Be Part of Integration
  11. Clearbit and HubSpot: Native Data, Standalone Product and Portfolio Pruning
  12. Lattice Engines and Dun & Bradstreet: From Acquired Platform to Portfolio Layer
  13. Mattermark and FullContact: Investor Outcome, Product Outcome and Data Outcome Diverged
  14. Claap and lemlist: A Strong Thesis Is Not Yet a Long-Term Outcome
  15. A Comparative View
  16. Product Survival Has Layers
  17. Data Absorption Requires Its Own Audit
  18. The Customer’s Acquisition Checklist
  19. The Founder’s Acquisition Checklist
  20. The Acquirer’s Integration Thesis and Evidence Plan
  21. Worked Example: One Deal, Five Verdicts
  22. What Acquisition Adds to the Four Questions
  23. Notes

Chapter 20 — Why GTM-Data Companies Fail

  1. Failure Is Usually a System, Not an Incident
  2. The Data Factory Has Unit Economics
  3. The Decay Tax
  4. The Data Is Available Elsewhere
  5. Broad but Shallow
  6. The Product Cannot Prove Incremental Value
  7. The Buyer Likes It but Does Not Own a Budget
  8. Distribution Costs More Than the Dataset Can Carry
  9. Platform Dependency Is a Concentrated Liability
  10. The Services Trap
  11. The Signal Is Interesting but Not Actionable
  12. A Feature Can Be Absorbed by the Suite
  13. Timing, Capital and the Founder
  14. Historical Endpoints Do Not Reveal One Cause
  15. The Failure Dashboard
  16. Recovery Is a Strategic Narrowing
  17. Worked Example: GridSite Intelligence
  18. What Failure Adds to the Four Questions
  19. Notes

Chapter 21 — Privacy, Ownership and Trust

  1. “Public” Is an Access Condition
  2. A Source Has a Rights Envelope
  3. Facts, Expression and Databases Are Not the Same Object
  4. Scraping Is a Bundle of Questions
  5. Personal Data Does Not Stop at the Office Door
  6. GDPR: Lawful Basis Is Only the Beginning
  7. The KASPR Case: Visibility Did Not Establish Fair Processing
  8. California: Data Brokers and Deletion at Scale
  9. Permissible Purpose Is a Product Boundary
  10. Collecting a Contact Is Not Permission to Use Every Channel
  11. Inference Creates New Responsibility
  12. Customer Uploads Change the Vendor’s Role
  13. Trust Requires Product Controls
  14. A Policy Engine Must Sit Before Activation
  15. Worked Example: European Hotel Ownership and Management Intelligence
  16. The Trust Review
  17. What Trust Adds to the Four Questions
  18. Notes

Chapter 22 — From Filters to Meaning

  1. What Filters Do Well
  2. Three Ways Text Becomes Searchable
  3. Hybrid Search Preserves Different Kinds of Evidence
  4. The Query Must Become a Set of Tests
  5. Terminology Expansion Is a Hypothesis Generator
  6. Classification From Prose Creates Query-Specific Fields
  7. Relationship Extraction Requires Direction
  8. Event Extraction Must Separate Publication From Occurrence
  9. Multilingual Search Changes the Reach of the Market
  10. Query Rewriting Can Help—and Invent
  11. Reranking Is Not Qualification
  12. Semantic Search Does Not Produce Coverage by Itself
  13. Identity Remains Outside Similarity
  14. A Production Semantic Stack
  15. Semantic Judgments Need Stable Claims
  16. Unknown Is a Necessary Output
  17. Contradiction Search Must Be Deliberate
  18. Semantic Quality Must Fit a Cost and Latency Budget
  19. Evaluate the Component That Can Fail
  20. Worked Example: One Company, Four Interpretations
  21. What Meaning Adds to the Four Questions
  22. Notes

Chapter 23 — Constructing a Market That Does Not Exist as a Table

  1. Three Products That Are Often Called a Database
  2. Enrichment Cannot Discover a Missing Row
  3. Membership Must Be Defined Before Search
  4. Build a Source Map Before Building a Crawler
  5. Candidate Generation Is a Coverage Program
  6. The Candidate Ledger Preserves How a Company Entered
  7. Resolve the Organization Before Publishing the Account
  8. Evidence Packets Turn Mentions Into Claims
  9. Qualification Is Criterion by Criterion
  10. Store Base Facts Separately From Membership
  11. Publish a Market as a Versioned Product
  12. Maintenance Makes the Market a Product
  13. Measuring a Market When the Universe Is Unknown
  14. Worked Example One: European Cold-Chain Oncology Distributors
  15. Worked Example Two: Southeast Asian Direct-to-Chip Channel Partners
  16. Before and After Market Construction
  17. What Market Construction Adds to the Four Questions
  18. Notes

Chapter 24 — Why the Company Graph Must Include Time and Evidence

  1. Why a Flat Company Row Breaks
  2. A Graph Is a Data Model, Not Necessarily a Graph Database
  3. Keep Identity, Claim and Evidence Separate
  4. A Claim Needs Qualifiers
  5. Business Time and Knowledge Time Are Different
  6. Publication Date Is Not Event Date
  7. Open-Ended Does Not Mean Permanent
  8. “Current” Is a Computation
  9. Contradictions Are Data
  10. Do Not Complete the Graph by Imagination
  11. Events Change State; They Are Not State
  12. Confidence Is Not One Number
  13. Merges and Splits Must Preserve History
  14. Add, Update, Merge, Retract and Delete Are Different
  15. Evidence Must Be Immutable Enough to Audit and Mutable Enough to Govern
  16. The Current Company Record Is a Materialized View
  17. Worked Example: Reconstructing One Account Through Time
  18. Graph Quality Is Not Graph Size
  19. What the Temporal Evidence Graph Adds to the Four Questions
  20. Notes

Chapter 25 — The Agentic GTM Workflow

  1. An Agent Is Not a Workflow
  2. Begin With the Business Decision
  3. The Workflow Needs Durable Objects
  4. The End-to-End Agentic GTM Flow
  5. Allocate Work by Failure Mode
  6. Use an Autonomy Ladder
  7. The Workflow Is a State Machine
  8. Side Effects Must Be Idempotent
  9. Typed Interfaces Make Handoffs Testable
  10. Tool Access Is Decision Authority
  11. Search and Reasoning Need Budgets
  12. Exceptions Are a First-Class Product Surface
  13. Observability Must Follow the Business Object
  14. Feedback Must Be Typed Before It Trains Anything
  15. Worked Example: Product Approval to One Governed Action
  16. Measuring the Agentic Workflow
  17. What the Agentic Workflow Adds to the Four Questions
  18. Notes

Chapter 26 — AI-Native Intent and the New Signal Frontier

  1. “Intent” Now Covers Too Much
  2. The Signal Chain Has Five Stages
  3. The Signal Record Needs More Than a Type and Date
  4. Custom Semantics Expand the Event and Candidate-Signal Frontier
  5. Signal Discovery and Account Monitoring Are Different Problems
  6. Absence and Negative Change Require Special Care
  7. Champion Movement Is a Relationship Signal
  8. Competitor Connections Need Direction and Strength
  9. Partnerships and Channel Changes Create New Markets
  10. Capacity Expansion Is a State Machine
  11. Hiring Shows Resource Allocation, Not Completed Capability
  12. Regulatory Events Can Create Need Without Expressing Demand
  13. Physical-World Proxies Expand the Frontier—and the Risk
  14. Composite Signals Can Be Stronger—And More Opaque
  15. Novelty, Materiality and Fit Must Be Separate
  16. Signals Need Expiry, Suppression and Saturation
  17. Outcome Learning Must Not Rewrite Observation
  18. Measure the Entire Signal Chain
  19. Worked Example: A Signal Stack for Warehouse Automation
  20. What AI-Native Signals Add to the Four Questions
  21. Notes

Chapter 27 — What Becomes Valuable When Intelligence Becomes Cheap?

  1. Cheap Is Not Free
  2. Separate Commodity Capabilities From Compounding Assets
  3. The Replication Clock Reveals the Asset
  4. Value Migrates From the Answer to the Assurance
  5. Model Independence Becomes a Strategic Asset
  6. Pricing Moves Toward Maintained Outcomes
  7. Exclusive and Permissioned Observations Appreciate
  8. Trusted Identity Becomes More Valuable, Not Less
  9. Longitudinal State Outlives the Page
  10. Evidence Converts Interpretation Into a Trust Asset
  11. Ontology Becomes the Product’s Judgment
  12. Evaluation Sets Become Scarcer Than Prompts
  13. Customer Corrections Can Create a Network—or a Bias Machine
  14. Workflow Integration Creates Decision Memory
  15. Rights and Governance Appreciate With Scale
  16. Distribution Determines Whether the Asset Is Used
  17. Reliability and Delivery Become Part of the Data
  18. The Assets Form a Compounding Loop
  19. Assets That Depreciate
  20. Own the Decision Layer; Buy the Commodity Layer
  21. Worked Comparison: Two Battery-Recycling Data Products
  22. A Buyer’s Test
  23. What Appreciating Intelligence Adds to the Four Questions
  24. Notes

Chapter 28 — Choosing a Data Wedge

  1. Start With the Decision, Not the Dataset
  2. A Wedge Has Five Boundaries
  3. Seven Common Wedge Shapes
  4. The Wedge Is an Intersection
  5. Apply Fatal Gates Before Scoring
  6. Test Founder–Source–Market Fit
  7. Score the Surviving Wedges
  8. Put Evidence Beside the Score
  9. Missing Data Must Be Proven, Not Assumed
  10. Evaluate the Source Before the Model
  11. Estimate Maintenance Before Coverage
  12. Choose the Unit Before the Market
  13. Distribution Is Part of Wedge Selection
  14. Match the Business Model to the Data Work
  15. A Wedge Should Expand Along the Same Data Object
  16. Avoid the Services Disguise
  17. Worked Scorecard: Four Candidate Wedges
  18. The Wedge Selection Sprint
  19. The Final Wedge Test
  20. What Wedge Choice Adds to the Four Questions
  21. Notes

Chapter 29 — Building the Minimum Credible Data Product

  1. The Credibility Boundary
  2. Begin With a Product Promise
  3. The Minimum Credible Stack
  4. Let the Interface Expose the Data Model
  5. Bound the Universe and Version the Definition
  6. Choose the Canonical Object and Give It a Stable Identity
  7. Publish a Controlled Schema, Not a Bag of Attributes
  8. Attach Evidence to Claims, Not Merely Rows
  9. Make Time a First-Class Field
  10. Preserve the Observation-to-Action Boundary
  11. Separate Unknown, Inferred, Disputed and False
  12. Divide Model Work From Deterministic State
  13. Treat Rights, Privacy and Security as Release Inputs
  14. Put Human Review Where the Error Is Expensive
  15. Corrections Must Survive the Next Run
  16. Delivery Is Part of the Product
  17. Design the Refresh Before Launch
  18. State a Service Level the Team Can Operate
  19. Make Customer Matching an Explicit Stage
  20. What Can Wait—and What Cannot
  21. Worked Release: From Demo to Data Product
  22. The Release Gates
  23. Credibility Is an Economic Choice
  24. What a Credible Product Adds to the Four Questions
  25. Notes

Chapter 30 — Proving That New Data Is Valuable

  1. Define the Decision Being Tested
  2. Lock the Query Before Running the Systems
  3. Use a Portfolio of Difficult Queries
  4. Prevent Benchmark Leakage
  5. Compare Workflows, Not Brand Names
  6. Pool the Candidates Before Judging Them
  7. Blind the Adjudication
  8. Seed Known Positives and Hard Negatives
  9. Measure Discovery and Qualification Separately
  10. Evaluate the Threshold, Ranking and Review Queue
  11. Measure Identity as Its Own System
  12. Report Sampling Uncertainty and Reviewer Disagreement
  13. Score Relationships and Events at Claim Level
  14. Measure Evidence Sufficiency, Not Explanation Quality
  15. Put Time and Stability Into the Benchmark
  16. Audit Sources and Shared Blind Spots
  17. Price the Accepted Increment, Not the Generated Candidate
  18. Test Workflow Value After Data Quality
  19. Precommit Success Thresholds
  20. Convert Metrics Into a Deployment Decision
  21. Worked Pilot: Five Market Questions
  22. Diagnose Failure Instead of Averaging It Away
  23. The Evidence Package for a Buyer
  24. What Proof Adds to the Four Questions
  25. Notes

Chapter 31 — Selling a Data Product

  1. Choose Which Kind of Buyer You Are Selling To
  2. Sell the Missing Decision, Not “Better Data”
  3. Find the Owner Whose Metric Changes
  4. Map the Buying Committee Early
  5. Multithread Without Bypassing the Champion
  6. Conduct Discovery Around Real Artifacts
  7. Qualify the Opportunity Before Offering a Pilot
  8. Write the Pilot as a Decision Document
  9. Avoid Pilot Purgatory
  10. Show a Methodology Package Before It Is Requested
  11. Choose a Delivery Model That Fits the Buyer
  12. Decide Whether to Sell Direct, OEM or Both
  13. Price the Maintained Decision
  14. Build the Economic Case Against the Real Alternative
  15. Treat Rights Review as a Product Conversation
  16. Make Security and Supply-Chain Diligence Easy
  17. Negotiate Service Levels Around the Data, Not Only the API
  18. Plan the First Ninety Days Before Signature
  19. Design Renewal at the Start
  20. Handle the Common Objections Directly
  21. Worked Sale: A GTM Platform
  22. What Selling Adds to the Four Questions
  23. Notes

Chapter 32 — The Next Data Companies

  1. The Feasibility Frontier Has Moved
  2. The Next Company Is Not Another Universal Database
  3. The Three-Layer Test
  4. A Taxonomy of the Next Data Companies
  5. From Buyer Prompts to Maintained Objects
  6. Local and Jurisdictional Company Databases
  7. Industry-Specific Intelligence Systems
  8. Enterprise Maps and Relationship Graphs
  9. Living-Data Companies
  10. Market Compilers
  11. Evidence and Assurance Rails
  12. Organize Around Responsibility, Not a Model
  13. Identity Is the First Strategic Choice
  14. Relationship Becomes the New Firmographic
  15. Change Becomes a Maintained State Machine
  16. Distribution Moves From Export to Participation
  17. Company Form Shapes Distribution
  18. Founder–Market Fit Becomes Founder–Data Fit
  19. Narrow Does Not Necessarily Mean Small
  20. An Opportunity Atlas
  21. Counterarguments—and What They Get Right
  22. Failure Modes of the Next Data Company
  23. A Test for the Next Data Company
  24. How to Read the Four Build Playbooks
  25. What the Next Data Companies Add to the Four Questions
  26. Notes

Chapter 33 — Local AI-Built Company Databases

  1. Decide What “Every Business” Means
  2. Maintain a Jurisdictional Population Ledger
  3. California Is the State Spine, Not the Finished Product
  4. Federal Sources Add Slices, Not Completeness
  5. Add State and Local Operating Layers
  6. Build an Object Graph Before Building the “Golden Row”
  7. Match the Operating Web, Not Merely a Domain String
  8. Treat Locations and Opening Hours as Time-Varying Claims
  9. Owners, Officers and Contacts Must Remain Separate
  10. Divide the Work Between Deterministic Systems, Crawlers and Models
  11. Build Rights and Privacy Into the Source Architecture
  12. Refresh Each Claim at the Speed It Decays
  13. Worked Example: A California Refrigeration Contractor
  14. Measure Coverage by Layer
  15. Choose a Wedge Inside the Jurisdiction
  16. Expand From California to a Jurisdictional Network
  17. Who Buys the Finished Product?
  18. What Local Databases Add to the Four Questions
  19. Notes

Chapter 34 — Industry-Specific Intelligence Systems

  1. An Industry Is Not a Filter
  2. Begin With the Industry Decision
  3. Own a Domain Ontology That Changes Decisions
  4. Build a Source Spine, Not a Source Pile
  5. Preserve Every Identity the Domain Needs
  6. Authorization Is Not Operating or Commercial Capability
  7. Model Capability and Relationship Claims With Qualifiers
  8. Write Specialist Evidence Policies
  9. Operate a Claim Lifecycle, Not a Scraping Project
  10. Refresh According to How Each Fact Decays
  11. Specialist Review Is Part of the Product
  12. From the Chapter 23 Cohort to a Vertical Product
  13. Retail Requires a Different Vertical System
  14. Financial Services Requires Yet Another Model
  15. Decide What to Build and What to Buy
  16. Match Distribution and Business Model to the Workflow
  17. Defensibility Comes From Compounding Decisions
  18. Expand Along the Domain Graph
  19. Failure Modes of an Industry-Specific System
  20. A Vertical-System Readiness Test
  21. What Industry-Specific Intelligence Adds to the Four Questions
  22. Notes

Chapter 35 — Enterprise Maps and Relationship Graphs

  1. An Account Is a View, Not an Entity
  2. The Five Maps Inside an Enterprise
  3. Begin With Controlled Nodes and Edges
  4. Legal Hierarchy Provides a Spine, Not the Whole Map
  5. One Enterprise Has Several Hierarchies
  6. Completeness Is Relative to a Question
  7. Write a Graph Coverage Contract
  8. Brands, Domains and Sites Need Their Own Identities
  9. Buying Centers Are Claims, Not Org Charts
  10. Relationship Sources Have Different Meanings
  11. Government Awards Show the Value of Typed Edges
  12. Relationships Need Direction, Scope and Time
  13. Model Change as an Operation
  14. Absence Is Not a Negative Edge
  15. Discovery and Resolution Are Different Problems
  16. Build the Graph From Evidence Packets
  17. The Minimum Useful Enterprise Map
  18. Product Surfaces for the Same Graph
  19. Move From Read-Only Map to Controlled Action
  20. Business Models and Buyers
  21. Protect the Boundary Between Public Graph and Customer Memory
  22. Deploy the Map in Four Stages
  23. Evaluate the Map at the Edge Level
  24. Failure Modes
  25. Worked Example: Mapping a Global Manufacturing Account
  26. What Enterprise Maps Add to the Four Questions
  27. Notes

Chapter 36 — Living Data: CRM Health, Intent, and the Maintained Answerable Market

  1. A CRM Is a Belief System
  2. Cleaning Is a State Transition, Not a Batch Project
  3. Maintain Identity Before Modeling Events or Deriving Candidate Signals
  4. Preserve the Controlled Boundary
  5. Give Events a Validity Window and Candidate Signals a Commercial Half-Life
  6. Maintain Source-Specific Change
  7. A Living Market Has Three Loops
  8. Maintain Suppression and Negative Decisions
  9. The Maintained Answerable Market
  10. Refresh the Definition as Well as the Data
  11. Decide What May Act Automatically
  12. Measure Living Data as a System
  13. Evaluate the Incremental Decision, Not the Alert
  14. Make the System Replayable
  15. Corrections Maintain the Living System
  16. The Business Models for Living Data
  17. Failure Modes of the Living Market
  18. Worked Market: From Question to Continuous State
  19. The Four Questions, Reassembled
  20. The Final Thesis
  21. Notes

Appendix A — Timeline, 1841–2026

  1. Era I — Directories, Credit and Addressable Lists, 1841–1979
  2. Era II — Desktop CRM, Online Company Databases and Marketing Automation, 1980–2003
  3. Era III — Crowdsourced Contacts, Web-Scale Data and the API Turn, 2004–2013
  4. Era IV — Professional Graphs, ABM, Intent, Predictive Scoring and Data APIs, 2014–2021
  5. Era V — Generative AI, GTM Orchestration and Agentic Data Work, 2022–2026
  6. Closing Synthesis — What Changed, and What Did Not
  7. Notes

Appendix B — Vendor Case Matrix

  1. How to Read the Matrix
  2. Complete Comparison Matrix
  3. I. Directories, Identity and Contact Data
  4. II. Private-Company and Market Intelligence
  5. III. Intent, Reviews and Account-Based Marketing
  6. IV. Predictive Scoring and Customer Intelligence
  7. V. Prospecting and GTM Orchestration
  8. VI. Marketing Automation and Workflow Distribution
  9. VII. Adjacent Data Lineages
  10. VIII. Local and Vertical Market Data
  11. What the Cases Show
  12. Notes

Appendix C — The B2B Data Source Atlas

  1. How to Read the Atlas
  2. 1. Corporate Registries, Filings and Legal Identity
  3. 2. Licensing, Regulation and Enforcement
  4. 3. Procurement, Awards and Public Funding
  5. 4. Company-Owned Web Sources
  6. 5. Jobs and Workforce Change
  7. 6. Partner, Channel and Marketplace Sources
  8. 7. News, Media and Events
  9. 8. Professional, Social and Community Sources
  10. 9. Developer and Open-Source Sources
  11. 10. People and Contact Sources
  12. 11. First-Party and Customer-Private Data
  13. 12. Behavioral and Intent Sources
  14. 13. Maps, Permits and the Physical World
  15. 14. Shipping, Customs and Supply Chains
  16. 15. Commercial Provider Feeds
  17. Select Sources From the Claim Backward
  18. Source Competence Matrix
  19. Refresh by Decay, Not by Row
  20. Worked Source Map: Hospital Microgrid Integrators
  21. A Source Intake Card
  22. Final Rule
  23. Notes

Appendix D — Data-Quality Benchmark Template

  1. How to Use This Appendix
  2. 1. Benchmark Decision Sheet
  3. 2. Market-Definition Sheet
  4. 3. System and Effort Boundary
  5. 4. Query Portfolio
  6. 5. Candidate Ledger
  7. 6. Identity Adjudication
  8. 7. Blinded Qualification Rubric
  9. 8. Core Metric Dictionary
  10. 9. Economics and Workflow Metrics
  11. 10. Precommitted Thresholds
  12. 11. Result Table
  13. 12. Error Taxonomy
  14. 13. Source and Shared-Blind-Spot Audit
  15. 14. Repeat-Run and Maintenance Test
  16. 15. Worked Fictional Result
  17. 16. Production Decision Memo
  18. 17. Anti-Gaming Review
  19. 18. Final Benchmark Checklist
  20. Notes

Appendix E — Privacy and Source-Rights Checklist

  1. The Six Questions That Must Not Collapse Into “Can We Scrape It?”
  2. 1. Product and Purpose Intake
  3. 2. The Source Register
  4. 3. Facts, Expression and Database Rights
  5. 4. Field and Inference Register
  6. 5. Notice, Rights, Accuracy and Suppression
  7. 6. Retention, Security and Vendor Controls
  8. 7. AI, Semantic Extraction and Agents
  9. 8. Customer Entitlement and Downstream Use
  10. 9. Outreach and Communication Channels
  11. 10. Jurisdiction and Data-Broker Register
  12. 11. Go, Conditional Go and Stop
  13. 12. Founder Release Checklist
  14. 13. Customer Diligence Checklist
  15. 14. Investor and Acquirer Diligence Checklist
  16. 15. One-Page Approval Record
  17. Final Principle
  18. Notes

The Leanpub 60 Day 100% Happiness Guarantee

Within 60 days of purchase you can get a 100% refund on any Leanpub purchase, in two clicks.

See full terms...

Earn $8 on a $10 Purchase, and $16 on a $20 Purchase

We pay 80% royalties on purchases of $7.99 or more, and 80% royalties minus a 50 cent flat fee on purchases between $0.99 and $7.98. You earn $8 on a $10 sale, and $16 on a $20 sale. So, if we sell 5000 non-refunded copies of your book for $20, you'll earn $80,000.

(Yes, some authors have already earned much more than that on Leanpub.)

In fact, authors have earned over $15 million writing, publishing and selling on Leanpub.

Learn more about writing on Leanpub

Free Updates. DRM Free.

If you buy a Leanpub book, you get free updates for as long as the author updates the book! Many authors use Leanpub to publish their books in-progress, while they are writing them. All readers get free updates, regardless of when they bought the book or how much they paid (including free).

Most Leanpub books are available in PDF (for computers) and EPUB (for phones, tablets and Kindle). The formats that a book includes are shown at the top right corner of this page.

Finally, Leanpub books don't have any DRM copy-protection nonsense, so you can easily read them on any supported device.

Learn more about Leanpub's ebook formats and where to read them

Write and Publish on Leanpub

You can use Leanpub to easily write, publish and sell in-progress and completed ebooks and online courses!

Leanpub is a powerful platform for serious authors, combining a simple, elegant writing and publishing workflow with a store focused on selling in-progress ebooks.

Leanpub is a magical typewriter for authors: just write in plain text, and to publish your ebook, just click a button. (Or, if you are producing your ebook your own way, you can even upload your own PDF and/or EPUB files and then publish with one click!) It really is that easy.

Learn more about writing on Leanpub