The Data Collection Handbook · Part I. Foundations

Chapter 1. The Data Collection Landscape

In progress: this is the chapter plan. The full chapter, its charts, and its runnable experiment publish on this page.

What this chapter answers

Every data project starts with the same three questions: what is out there, how do people get it, and should we collect it ourselves or buy it? This chapter maps the industry so you can answer all three without guessing.

What you will learn

In this chapter

The experiment

The question: At what monthly volume of collected records does building in-house become cheaper than buying from a vendor?

A transparent cost model: engineer time, proxy and compute cost, and a maintenance tax that grows with every source added, against parameterized vendor price tiers. The model sweeps volume from ten thousand to one hundred million records per month and finds the breakeven point, with a sensitivity analysis on the most uncertain inputs. The data and code ship with the chapter, sized to run offline on a laptop with one command.

My working hypothesis is that the breakeven sits much higher than most teams assume, and that the maintenance tax, not the initial build, is what drives it. The published chapter will report whatever the model actually shows.

← Handbook index Chapter 2. Legality, Ethics, and Fair Use →

Get new chapters by email as they publish.