What a Web Page Costs

Every serious data-collection operation keeps a number the public never sees: what one successful page actually costs. Ours moves around, by site, by country, by how hard the target defends itself, and we have watched it for years the way a manufacturer watches the price of steel. Until this month, that number had one strange property. It was always discovered, never posted. Nobody sold access to the web at a listed price; you paid in infrastructure, retries, and patience, and found out afterwards what a page had cost you.

That changed, at least formally, on July 1. Cloudflare — which proxies nearly one website in four (W3Techs) — announced that sites on its network will classify automated visitors by declared purpose (Search, Agent, or Training), and that from September 15 its defaults flip: for domains newly onboarding, Training and Agent crawlers get blocked on ad-carrying pages unless the site decides otherwise. Its year-old Pay Per Crawl program, where a publisher sets one flat price per fetched page with a floor of $0.01, is evolving into what it calls Pay Per Use: paying publishers when content shapes an AI answer, not merely when it is fetched, piloted with Ceramic.ai and You.com.

Most coverage treated this as a story about AI companies versus publishers. I read it as something narrower and more interesting: the first posted price for machine access to the web, from the one company positioned to enforce one. And the first thing any operator does with a posted price is compare it with the real one.

The toll versus the fence

Here is the comparison nobody in the coverage seems to have run. The commercial market for getting pages out of defended websites is mature and publicly priced. The large providers sell "unlocking" as a service — you hand them a URL, they return the page, you pay only on success — at list prices of roughly $1 to $3 per thousand pages. That is $0.001 to $0.003 per page, for the hard case: sites that actively resist.

Cloudflare's official market opens at $0.01 per page — the floor, before any publisher decides their content is worth more. So the posted, permissioned price starts at three to ten times the going rate for simply defeating the defenses, as listed openly by companies with sales teams and terms of service.

The posted price opens above the unofficial marketColumn chart. Commercial unlocking APIs list at 1 to 3 dollars per thousand successful pages. Cloudflare Pay Per Crawl's price floor is 10 dollars per thousand — three to ten times higher.The posted price opens above the unofficial marketWhat 1,000 successful page fetches cost, public list prices, July 2026$0$5$10USD per 1,000 pages$1Unlocking APIs · low list$3Unlocking APIs · high list$10Pay Per Crawl · price floor3–10× the going rateSources: Cloudflare Pay Per Crawl documentation ($0.01/page floor); Bright Data Web Unlocker public list pricing, July 2026
View data table
Cost of 1,000 successful page fetches, public list prices
OfferPrice per 1,000 pages
Unlocking APIs — low list$1
Unlocking APIs — high list$3
Cloudflare Pay Per Crawl — price floor$10

A market where the legal product costs several multiples of the unofficial one does not clear on price. If pay-per-crawl were only selling access, it would have almost no buyers — and a year in, Cloudflare's own repositioning toward paying per answer reads as an admission that per-fetch pricing alone was not finding them. But I don't think access is what this market is really selling.

What the toll actually buys

What a paid, permissioned fetch produces that no unofficial fetch can produce is a receipt. When institutions buy web data — and I have sat through those diligence calls on the vendor side — the questions that decide the deal are about provenance: when was this observed, from where, under what right. I wrote in an earlier essay that defensible provenance is the scarce asset in this market, and it is getting scarcer as generated content floods the open web. A licensing regime with per-page billing events creates provenance as a by-product. The page was fetched with permission, at a recorded time, under recorded terms, and there is a paper trail proving it.

Seen that way, the price gap stops looking like a market failure. The spread between $0.001 and $0.01 is the compliance premium — what it costs to hold data whose right to exist survives a diligence call. Whether that premium is worth ten times the base price depends entirely on who your buyer is. For training corpora at web scale, probably not, which is why the labs haven't rushed in. For data feeding regulated or institutional decisions, the receipt can be worth more than the page.

The part that will not work as advertised

I am less convinced by the purpose labels. Search, Agent, and Training are declarations, and at the wire level the three fetches are identical; the packet does not carry a purpose. What a network can verify is identity — cryptographically, who is crawling — and identity is where the serious work is heading. Intent is a promise, and pricing by promise invites reclassification games rather than honesty. Nothing prevents tomorrow's crawler from calling itself a search engine that merely remembers what it read.

The default-block also lands unevenly in a way the announcement does not dwell on. It binds the crawlers that identify themselves — the declared, polite ones — because a default can only block what it can name. Cloudflare's own data shows even the polite ones have been careless: more than half of AI-crawler traffic re-fetches pages that have not changed since the last visit. But the traffic that never declares itself is untouched by any of this. Thales's 2026 bot report put automated traffic at 53% of the web last year, 40 points of it classified as malicious, and that 40% did not ask anyone's permission and will not start paying now. A posted price disciplines only the traffic that shows its name.

Automated traffic is now the majority of the webColumn chart. The automated share of all web traffic rose from 42.3 percent in 2021 to 47.4 in 2022, 49.6 in 2023, 51 in 2024 and 53 percent in 2025, crossing the 50 percent majority line in 2024.Automated traffic is now the majority of the webShare of all web traffic that is automated, Imperva / Thales Bad Bot Reports0%20%40%60%Share of traffic (%)50% — majority line42.3%202147.4%202249.6%202351%202453%2025+10.7 pp in four yearsSource: Imperva Bad Bot Reports 2022–2025; Thales 2026 Bad Bot Report (data years 2021–2025)
View data table
Automated share of all web traffic
YearAutomated share
202142.3%
202247.4%
202349.6%
202451%
202553%

Why the free-crawl bargain died now

The web's original deal with crawlers was barter, not money: read my pages, send me readers. Search engines held that bargain up for twenty-five years. The reason tolls are appearing in 2026 is that the new crawlers consume like search engines but refer like nothing at all. Cloudflare's Radar unit publishes the asymmetry as a number, the crawl-to-refer ratio, and this spring's readings made it concrete: for every visitor Anthropic's crawler sends back to a website, it reads on the order of eleven thousand pages. OpenAI's ratio was measured in the high hundreds to one. Google's — the old bargain — about five to one. When one side of a barter stops delivering its half, the other side starts asking for cash. That, more than any lawsuit, is what changed.

The crawl-for-referral bargain has collapsedThree figures. Pages crawled for every one visitor referred back to websites in late May 2026: Anthropic about 11,122 to 1, OpenAI about 857 to 1, Google about 5 to 1.The crawl-for-referral bargain has collapsedPages crawled for every visitor referred back to websites, late May 202611,122 : 1Anthropic857 : 1OpenAI5 : 1Googlethe old bargain, still intactSource: Cloudflare Radar crawl-to-refer measurements, week of May 25–June 1, 2026
View data table
Pages crawled per visitor referred, late May 2026
OperatorCrawl-to-referral ratio
Anthropic11,122 : 1
OpenAI857 : 1
Google5 : 1

Where this goes, I think, is a web with two published prices: one for identity-verified, permissioned collection with receipts, and the effective price of everything else. The first will fall as the market finds real clearing prices below today's optimistic floors. The second will rise, because defenses keep improving and the arms race is not free. Somewhere in the middle they meet, and the interesting question for anyone who buys, sells, or invests in web data is which side of that line their supply chain sits on — and whether anyone can prove it.

The practical version of that question fits in one diligence line. Ask your data vendor what a page costs them, and how they know. An operator with a real answer understands their own supply chain. One without an answer is telling you the provenance story ends at their front door.