The First Public Map of What AI Actually Crawled
The report my billing system produces fastest is the one the EU just turned into law. For any month since we started, I can pull our crawled domains ranked by pages collected, because customers pay per page and disputes get settled from that ledger. It exists for invoicing, not for virtue. Under the AI Act, every company training a general-purpose AI model on scraped web data now has to publish roughly that report. Not to a regulator, under seal. Publicly, on its own website, updated every six months.
The template was written for whoever runs the crawl
The legal basis is Article 53(1)(d) of the AI Act, which requires providers of general-purpose AI models to publish a "sufficiently detailed summary" of their training content, in a mandatory template the European Commission's AI Office released on July 24, 2025. Most of the template reads like ordinary disclosure paperwork until you reach the section on data crawled and scraped from online sources. That section asks for the names of the crawlers used. The period of collection. A description of what was scraped, the types of websites, their geography, their languages. And then the line with teeth: a list of the top 10% of all internet domains scraped, measured by the size of the content collected. Smaller companies get a concession, the top 5% or 1,000 domains, whichever is lower.
I have read a fair amount of AI regulation written by people who have never operated anything, and this is not that. You cannot produce that domain list from memory, or from a lawyer's interview with the engineering team. You can only produce it if the crawl infrastructure kept per-domain accounting while it ran. Quietly, the template is a test of whether a company's data pipeline has bookkeeping at all.
View data table
| Item | What must be published |
|---|---|
| Crawlers | Names of all crawlers used to collect the training data |
| Period | The period of collection for each crawl |
| Content description | Types of websites, their geography, their languages |
| Domain list | Top 10% of all domains scraped by size of content collected (SMEs: top 5% or 1,000 domains, whichever is lower) |
| Cadence | Summary refreshed every six months; pre-August 2025 models file by August 2, 2027 |
A year of very quiet compliance
The obligation went live on August 2, 2025 for every model placed on the EU market from that date. In January 2026, three researchers from the AI Accountability Lab at Trinity College Dublin and Mozilla went looking for the results: search engines, model repositories, provider legal pages, an exhaustive sweep they later published at ACM FAccT 2026. They found four summaries. Hugging Face had filed one for SmolLM3, a 3-billion-parameter open model. The Swiss National AI Initiative had filed one for its Apertus family. The Polish open-source collective SpeakLeash had filed one for Bielik. Bria AI, an image company that trains on licensed data, had filed one for Bria 3.2. From the six frontier providers the team checked first, Anthropic, Google, Meta, Mistral, OpenAI and xAI, they found nothing. There was also a file in Microsoft's Phi-4 repository that resembled the template without saying so; graded on its content, it scored D for transparency and F for usefulness.
In March the same authors wrote it up for TechPolicy.Press under a headline that needed no qualifier: big AI developers were skirting the mandate, and the small teams filing careful summaries had proven the task possible on a fraction of the budget.
Then the calendar did what the mandate alone could not. The lab's public tracker now lists fifteen graded summaries, and the missing names have appeared on it: GPT-5.5, Gemini 3 Pro, Grok 4.5, Meta's two Muse image models, Microsoft's MAI-Image-2. Look at the grades, though, and the map has two resolutions. Apertus still holds the only A. The B grades belong to Bria, Hugging Face, SpeakLeash and Domyn. The frontier labs cluster at C, enough structure to claim the exercise, not enough content to be useful, with usefulness graded D+ for OpenAI's GPT-5.5, Meta's Muse pair and Microsoft's MAI-Image-2. And four models on the tracker still show no summary at all: OpenAI's GPT-OSS, GPT-5.6 Sol and Sora 2, and Anthropic's Claude Fable 5. The first public map of what AI actually crawled exists now, and it is sharpest exactly where the stakes are smallest.
View data table
| Grade band | Models | Count |
|---|---|---|
| A | Apertus (Swiss AI Initiative) | 1 |
| B / B+ | Bria 3.2, SmolLM3-3B, Bielik v3 11B Instruct, Domyn Large | 4 |
| C / C+ | GPT-5.5 (OpenAI), Gemini 3 Pro (Google), Grok 4.5 (xAI), Muse Image (Meta), Muse Spark (Meta), MAI-Image-2 (Microsoft), Adobe Firefly, Inkling (Thinking Machines), FastwebMIIA (Fastweb) | 9 |
| D | Phi-4 (Microsoft) | 1 |
| No summary | GPT-OSS (OpenAI), GPT-5.6 Sol (OpenAI), Sora 2 (OpenAI), Claude Fable 5 (Anthropic) | 4 |
The one deadline the omnibus did not move
On August 2, 2026, five days after this essay publishes, the Commission's AI Office gains its enforcement powers over general-purpose AI providers: the power to demand documentation, run model evaluations, order corrective measures, restrict a model from the EU market, and fine up to 3% of global annual turnover or 15 million euros, whichever is higher. A missing or hollow training-content summary is squarely inside that power.
Context makes the date sharper. When Brussels agreed its digital omnibus package in May, it pushed the high-risk AI deadlines back by sixteen months, to December 2027, and the rules for AI embedded in regulated products to 2028. The GPAI enforcement date it left exactly where it was. Whatever the omnibus says about Europe's appetite for delay, that appetite did not extend to this obligation.
And the map is scheduled to get denser. Models placed on the market before August 2025 owe their summaries by August 2, 2027. That covers the first generation of frontier models, the ones trained when nobody was writing anything down. Their providers either produce a domain ledger for that era or explain to a regulator why they cannot.
View data table
| Date | Summaries found | From frontier labs |
|---|---|---|
| January 2026 | 4 | 0 |
| July 2026 | 15 graded | 6 entries (OpenAI, Google, xAI, Meta ×2, Microsoft), all in the C band |
What I would do with the map
I run crawlers for a living and sit on the vendor side of data diligence calls, so here is the map read as an operator's document rather than a compliance story.
If you invest in or buy from AI companies, read the summary before the pitch deck. The C grades are a tell, and not mainly a legal one. To publish your top 10% of domains you need per-domain accounting inside the crawl pipeline, and a provider that ships a vague summary is telling you one of two things: it chose vagueness with a regulator watching, or its pipeline cannot produce the number. Both answers are worth having before you believe anything else the data section of the deck says.
If you publish content, the lists are leverage arriving on a schedule. Once the legacy models file in 2027, a publisher will be able to look across the industry's disclosures and see whose corpora its work fed. A licensing conversation that starts from a published fact is a different conversation from one that starts with a subpoena.
And if you collect data, the lesson is one this site keeps arriving at from different directions. Two weeks ago I argued that what buyers of web data are really buying is provenance. The EU has now made a slice of provenance public infrastructure: comparable, falsifiable, refreshed every six months. The operators who kept honest per-domain logs, the boring accounting nobody ever asked to see, turn out to have been building the deliverable all along.
For thirty years, what the crawlers actually took was the most consequential dataset nobody could see. The first public version of it is live, graded, and mostly C's. August 2 is when a C starts costing money.