In David’s Words
What David Says About Data Collection
Lines from his talks. Seven featured at a time, reshuffled every month.
He is writing these lessons down in The Data Collection Handbook, a free open book with a replicable experiment in every chapter.
Your competitors publish their strategy every day. It’s called their website.
Good data collection is boring on purpose. The excitement belongs in the decisions, not in whether the data arrived.
Companies don’t have a data problem. They have a collection problem. The answers are public, sitting on pages they never collect.
The best datasets aren’t downloaded. They are built, page by page, with someone accountable for every row.
In data work, trust is the product. The rows are just the packaging.
In travel, a price is true for minutes. By the time a quarterly report prints it, it is history.
A web page looks free. Collecting millions of them, correctly, every day, is one of the hardest engineering problems I know.
July 2026 selection
The Full List
All quotes by David Martin Riveros, from his talks on data collection, web data, and AI.
Every business decision is a bet. Fresh market data doesn’t remove the risk. It just means you saw the table before you sat down.
A price on a website is an opinion. The same price, tracked across every competitor for a full season, is a strategy.
You can’t out-analyze bad collection. If the data coming in is incomplete, every dashboard downstream is fiction.
Complete beats fast. A dataset with holes doesn’t tell you less. It lies to you.
The gap between what a company knows and what it could know is usually one uncollected dataset wide.
Good data collection is boring on purpose. The excitement belongs in the decisions, not in whether the data arrived.
Data pipelines fail like supply chains. Everyone audits the warehouse. Nobody checks whether the trucks came back full.
If your market moves daily and you collect monthly, that’s not analytics. That’s archaeology.
In data work, trust is the product. The rows are just the packaging.
The most expensive data is the data you didn’t collect when it was still there. The web forgets fast.
The web is the largest market report ever written. It rewrites itself every second, and most companies still read it once a quarter.
Your competitors publish their strategy every day. It’s called their website.
A web page looks free. Collecting millions of them, correctly, every day, is one of the hardest engineering problems I know.
AI didn’t reduce the need for data collection. It made every company hungry for exactly the data that is hardest to get.
AI agents will shop, book, and compare on our behalf. What they recommend will depend on the data they were fed.
Machines read the web differently than people do. A crawler catches the price change at 3 a.m. that no analyst was awake to see.
In travel, a price is true for minutes. By the time a quarterly report prints it, it is history.
Companies don’t have a data problem. They have a collection problem. The answers are public, sitting on pages they never collect.
The best datasets aren’t downloaded. They are built, page by page, with someone accountable for every row.
Turning data into advantage is not about volume. It is about collecting the right pages on the day they change.