A book in progress

The Data Collection Handbook

I have spent years building systems that collect web data at scale, and most of what I know arrived the expensive way: through production incidents, wrong datasets, and lessons no book had written down. This handbook is my attempt to write them down, in the open, for the whole industry: the engineers and analysts who build collectors, and the founders and buyers who pay for the results.

Three rules govern every chapter. First, plain language: no jargon without a one-line definition. Second, evidence: each chapter runs one analytical experiment, with data and runnable code published alongside it, so you can replicate every result instead of taking my word. Third, the written analysis carries the full conclusion, so you get the finding even if you never run the code.

One scope note. This book covers collecting publicly visible data at fair load. It does not teach collecting from behind logins, defeating CAPTCHAs, or evading any specific vendor's detection product. Where anti-bot systems appear, the treatment is how detection works and why coherent, respectful traffic is the durable strategy. Those are principled boundaries, not oversights.

This page is the book's index. Each chapter link below leads to that chapter's plan today, and to the full chapter as it publishes. The structure went through independent review from two directions, an engineer's read and a buyer's read, before the first chapter was drafted.

Part I. Foundations

Part II. Getting the Data

Part III. Trusting the Data

Part IV. Running It at Scale

Part V. The Other Side of the Table

Get new chapters by email as they publish.