The First Public Map of What AI Actually Crawled

The report my billing system produces fastest is the one the EU just turned into law. For any month since we started, I can pull our crawled domains ranked by pages collected, because customers pay per page and disputes get settled from that ledger. It exists for invoicing, not for virtue. Under the AI Act, every company training a general-purpose AI model on scraped web data now has to publish roughly that report. Not to a regulator, under seal. Publicly, on its own website, updated every six months.

The template was written for whoever runs the crawl

The legal basis is Article 53(1)(d) of the AI Act, which requires providers of general-purpose AI models to publish a "sufficiently detailed summary" of their training content, in a mandatory template the European Commission's AI Office released on July 24, 2025. Most of the template reads like ordinary disclosure paperwork until you reach the section on data crawled and scraped from online sources. That section asks for the names of the crawlers used. The period of collection. A description of what was scraped, the types of websites, their geography, their languages. And then the line with teeth: a list of the top 10% of all internet domains scraped, measured by the size of the content collected. Smaller companies get a concession, the top 5% or 1,000 domains, whichever is lower.

I have read a fair amount of AI regulation written by people who have never operated anything, and this is not that. You cannot produce that domain list from memory, or from a lawyer's interview with the engineering team. You can only produce it if the crawl infrastructure kept per-domain accounting while it ran. Quietly, the template is a test of whether a company's data pipeline has bookkeeping at all.

What goes public for every scraped corpusThe EU template's four disclosure items for data crawled from online sources. One: the names of the crawlers used. Two: the period of collection. Three: a description of the content, covering types of websites, geography and languages. Four, highlighted: a list of the top 10 percent of all domains scraped, by size of content collected, with smaller companies disclosing the top 5 percent or 1,000 domains, whichever is lower. Summaries must be refreshed every six months, and models on the market before August 2025 must file by August 2, 2027.What goes public for every scraped corpusThe template's disclosure items for data crawled and scraped from online sources1Names of the crawlersevery crawler that fed the training set, identified by name2Period of collectionwhen each crawl ran3Description of the contenttypes of websites, their geography, their languages4Top 10% of all domains scraped, by content volumethe crawl ledger itself; SMEs disclose the top 5% or 1,000 domains, whichever is lowerRefreshed every six months. Models on the EU market before August 2, 2025 must file by August 2, 2027.Source: European Commission AI Office, Template for the public summary of training content, July 24, 2025
View data table
Disclosure items required for web-scraped training data, EU AI Act Article 53(1)(d) template
ItemWhat must be published
CrawlersNames of all crawlers used to collect the training data
PeriodThe period of collection for each crawl
Content descriptionTypes of websites, their geography, their languages
Domain listTop 10% of all domains scraped by size of content collected (SMEs: top 5% or 1,000 domains, whichever is lower)
CadenceSummary refreshed every six months; pre-August 2025 models file by August 2, 2027

A year of very quiet compliance

The obligation went live on August 2, 2025 for every model placed on the EU market from that date. In January 2026, three researchers from the AI Accountability Lab at Trinity College Dublin and Mozilla went looking for the results: search engines, model repositories, provider legal pages, an exhaustive sweep they later published at ACM FAccT 2026. They found four summaries. Hugging Face had filed one for SmolLM3, a 3-billion-parameter open model. The Swiss National AI Initiative had filed one for its Apertus family. The Polish open-source collective SpeakLeash had filed one for Bielik. Bria AI, an image company that trains on licensed data, had filed one for Bria 3.2. From the six frontier providers the team checked first, Anthropic, Google, Meta, Mistral, OpenAI and xAI, they found nothing. There was also a file in Microsoft's Phi-4 repository that resembled the template without saying so; graded on its content, it scored D for transparency and F for usefulness.

In March the same authors wrote it up for TechPolicy.Press under a headline that needed no qualifier: big AI developers were skirting the mandate, and the small teams filing careful summaries had proven the task possible on a fraction of the budget.

Then the calendar did what the mandate alone could not. The lab's public tracker now lists fifteen graded summaries, and the missing names have appeared on it: GPT-5.5, Gemini 3 Pro, Grok 4.5, Meta's two Muse image models, Microsoft's MAI-Image-2. Look at the grades, though, and the map has two resolutions. Apertus still holds the only A. The B grades belong to Bria, Hugging Face, SpeakLeash and Domyn. The frontier labs cluster at C, enough structure to claim the exercise, not enough content to be useful, with usefulness graded D+ for OpenAI's GPT-5.5, Meta's Muse pair and Microsoft's MAI-Image-2. And four models on the tracker still show no summary at all: OpenAI's GPT-OSS, GPT-5.6 Sol and Sora 2, and Anthropic's Claude Fable 5. The first public map of what AI actually crawled exists now, and it is sharpest exactly where the stakes are smallest.

The map has two resolutionsGrade ladder of the 19 models tracked by the AI Accountability Lab as of July 27, 2026. Grade A: one model, Apertus by the Swiss AI Initiative. Grade B band: four models, Bria 3.2, SmolLM3-3B, Bielik v3 11B and Domyn Large. Grade C band: nine models, including GPT-5.5, Gemini 3 Pro, Grok 4.5, Meta's Muse Image and Muse Spark, MAI-Image-2, Adobe Firefly, Inkling and FastwebMIIA. This is where the frontier labs all land. Grade D: one model, Microsoft's Phi-4. No summary at all: four models, GPT-OSS, GPT-5.6 Sol, Sora 2 and Claude Fable 5.The map has two resolutionsTransparency grades of published AI training-content summaries, July 2026A1Apertus (Swiss AI Initiative)B4Bria 3.2 · SmolLM3-3B · Bielik v3 11B · Domyn LargeC9the frontier labs all land hereGPT-5.5 · Gemini 3 Pro · Grok 4.5 · Muse Image · Muse SparkMAI-Image-2 · Adobe Firefly · Inkling · FastwebMIIAD1Phi-4 (Microsoft)No summary4GPT-OSS · GPT-5.6 Sol · Sora 2 · Claude Fable 5Grade bands include plus grades. Models, not providers: Meta and Microsoft each appear twice, OpenAI four times.Source: AI Accountability Lab, Trinity College Dublin, GPAI training transparency tracker (aial.ie), accessed July 27, 2026
View data table
Transparency grades of tracked training-content summaries, AI Accountability Lab, July 27, 2026
Grade bandModelsCount
AApertus (Swiss AI Initiative)1
B / B+Bria 3.2, SmolLM3-3B, Bielik v3 11B Instruct, Domyn Large4
C / C+GPT-5.5 (OpenAI), Gemini 3 Pro (Google), Grok 4.5 (xAI), Muse Image (Meta), Muse Spark (Meta), MAI-Image-2 (Microsoft), Adobe Firefly, Inkling (Thinking Machines), FastwebMIIA (Fastweb)9
DPhi-4 (Microsoft)1
No summaryGPT-OSS (OpenAI), GPT-5.6 Sol (OpenAI), Sora 2 (OpenAI), Claude Fable 5 (Anthropic)4

The one deadline the omnibus did not move

On August 2, 2026, five days after this essay publishes, the Commission's AI Office gains its enforcement powers over general-purpose AI providers: the power to demand documentation, run model evaluations, order corrective measures, restrict a model from the EU market, and fine up to 3% of global annual turnover or 15 million euros, whichever is higher. A missing or hollow training-content summary is squarely inside that power.

Context makes the date sharper. When Brussels agreed its digital omnibus package in May, it pushed the high-risk AI deadlines back by sixteen months, to December 2027, and the rules for AI embedded in regulated products to 2028. The GPAI enforcement date it left exactly where it was. Whatever the omnibus says about Europe's appetite for delay, that appetite did not extend to this obligation.

And the map is scheduled to get denser. Models placed on the market before August 2025 owe their summaries by August 2, 2027. That covers the first generation of frontier models, the ones trained when nobody was writing anything down. Their providers either produce a domain ledger for that era or explain to a regulator why they cannot.

Compliance arrived with the deadline, not the obligationColumn chart. In January 2026, an exhaustive search by the AI Accountability Lab found 4 published training-content summaries, none from the frontier labs. By July 2026 the lab's tracker lists 15 graded summaries, including six frontier-lab entries, all graded C. Enforcement begins August 2, 2026, with fines up to 3 percent of global turnover.Compliance arrived with the deadline, not the obligationPublished training-content summaries found by the AI Accountability Lab051015Summaries found4none from the frontier labsJanuary 202615July 2026six frontier-lab entries,all graded CAug 2, 2026: fines up to3% of global turnoverSources: Blankvoort, Pandit and Gahntz, ACM FAccT 2026; AI Accountability Lab tracker (aial.ie), July 27, 2026
View data table
Published training-content summaries found, January vs July 2026
DateSummaries foundFrom frontier labs
January 202640
July 202615 graded6 entries (OpenAI, Google, xAI, Meta ×2, Microsoft), all in the C band

What I would do with the map

I run crawlers for a living and sit on the vendor side of data diligence calls, so here is the map read as an operator's document rather than a compliance story.

If you invest in or buy from AI companies, read the summary before the pitch deck. The C grades are a tell, and not mainly a legal one. To publish your top 10% of domains you need per-domain accounting inside the crawl pipeline, and a provider that ships a vague summary is telling you one of two things: it chose vagueness with a regulator watching, or its pipeline cannot produce the number. Both answers are worth having before you believe anything else the data section of the deck says.

If you publish content, the lists are leverage arriving on a schedule. Once the legacy models file in 2027, a publisher will be able to look across the industry's disclosures and see whose corpora its work fed. A licensing conversation that starts from a published fact is a different conversation from one that starts with a subpoena.

And if you collect data, the lesson is one this site keeps arriving at from different directions. Two weeks ago I argued that what buyers of web data are really buying is provenance. The EU has now made a slice of provenance public infrastructure: comparable, falsifiable, refreshed every six months. The operators who kept honest per-domain logs, the boring accounting nobody ever asked to see, turn out to have been building the deliverable all along.

For thirty years, what the crawlers actually took was the most consequential dataset nobody could see. The first public version of it is live, graded, and mostly C's. August 2 is when a C starts costing money.