Is Robots.txt Legally Binding? Two Courts Answered in the Same Week
The first file any serious crawler requests from a new website is robots.txt. Ours does, on every target, before anything else, and for most of my career the file had a strange status: universally read, universally understood, and legally weightless. It dates to 1994, when the early web agreed on a convention for telling robots where not to go, and it survived twenty-eight years on politeness alone before the IETF even wrote it down as a formal standard, RFC 9309, in 2022. Nobody knew what it was worth in court, because almost nobody had asked a court.
Then, in five days last December, two courts answered, and they disagreed about almost everything except the file's name.
The American answer: a sign, not a fence
On December 15, 2025, Judge Sidney Stein of the Southern District of New York dismissed Ziff Davis's anti-circumvention claim against OpenAI. Ziff Davis, which publishes IGN, Mashable and PCMag, had argued that its robots.txt directives telling GPTBot to stay out were a "technological measure that effectively controls access" to its copyrighted work under the DMCA, so ignoring them was circumvention, the same legal category as breaking encryption.
The court said no, in an image that will follow this file around for years: robots.txt is a "keep off the grass" sign. It asks; it does not prevent. Circumvention means affirmatively disabling a control, something like breaking into a locked system, and a crawler that reads a request and proceeds anyway has disabled nothing. Three days later the judge denied Ziff Davis's request to amend the claim. The rest of the case survives, including contributory infringement and removal of copyright-management information, so OpenAI did not walk away clean. But in the US, as of December, ignoring robots.txt is not hacking. It may still be infringement; it is not circumvention.
The German answer: the only sign that counts
Five days earlier, on December 10, the Higher Regional Court of Hamburg decided Kneschke v. LAION, the appeal in Europe's first real text-and-data-mining case, and reached the mirror image. Robert Kneschke is a photographer whose images ended up in LAION's training dataset. EU law since 2019 has given him a genuine legal lever the American publisher lacked: under Article 4(3) of the DSM Directive, commercial text and data mining is allowed unless the rightsholder reserves their rights, and for online content that reservation must be machine-readable.
Kneschke had a reservation. The stock-photo site hosting his images carried plain-language terms forbidding automated scraping. The court held that this was not enough: for the 2021 crawl at issue, a reservation written for humans did not meet the machine-readability requirement, and the judges pointed to what would have counted, robots.txt entries, the TDM Reservation Protocol, metadata tags. The court also confirmed the uncomfortable part for rightsholders, that the TDM exception covers training generative AI at all, and it allowed a further appeal to Germany's Federal Court of Justice, so this is not the final word.
Put the two rulings side by side and the asymmetry is the story. The American publisher posted the machine-readable sign and was told a sign is not a fence. The German photographer wrote his objection in human language and was told only the machine-readable sign would have counted. One file, one week, two legal systems reading it in opposite directions. In the United States robots.txt turned out to be too weak to trigger the anti-hacking statute. In the European Union it turned out to be, more or less exactly, the legal instrument the statute demands.
View data table
| Jurisdiction | Legal reading | Basis |
|---|---|---|
| United States | A request, not a barrier; ignoring it is not DMCA circumvention | Ziff Davis v. OpenAI, S.D.N.Y., Dec. 15, 2025 |
| European Union | A binding, machine-readable rights reservation AI trainers must honor; fines from Aug. 2, 2026 | Kneschke v. LAION, OLG Hamburg, Dec. 10, 2025; EU AI Act Arts. 53, 101 |
August 2 is when the EU reading gets teeth
The German reading would matter less if it stayed a private-law dispute between a photographer and a research collective. It has not. The EU AI Act's Article 53 requires every provider of a general-purpose AI model offered in the EU to maintain a copyright policy that identifies and complies with those Article 4(3) reservations, robots.txt included. That obligation has formally applied since August 2, 2025. What changes on August 2, 2026, less than two weeks out, is enforcement: the European Commission's AI Office gains the power to demand documentation, order corrective measures, restrict a model from the EU market, and fine providers up to 3% of global annual turnover or 15 million euros, whichever is higher.
There is one honest complication. Machine-readable is a standard without a spec: the Hamburg court named examples, not a definition, and the Commission is running a stakeholder process to publish a list of agreed opt-out protocols, expected late 2026. Until that list exists, providers and publishers are both aiming at a target that is still being drawn. But the direction is set. In the EU, a crawler operator's treatment of robots.txt is no longer etiquette. It is a documented compliance obligation with a regulator attached.
View data table
| Date | Event |
|---|---|
| April 2019 | EU DSM Directive Art. 4(3): commercial TDM allowed unless rights are reserved in machine-readable form |
| August 2, 2025 | EU AI Act Art. 53 applies: GPAI providers must identify and honor machine-readable reservations |
| December 10, 2025 | OLG Hamburg (Kneschke v. LAION): a plain-language opt-out is not machine-readable |
| December 15, 2025 | S.D.N.Y. (Ziff Davis v. OpenAI): robots.txt is a request, not a barrier; ignoring it is not DMCA circumvention |
| August 2, 2026 | European Commission enforcement powers begin: fines up to 3% of global turnover or €15M |
Meanwhile, the signs keep going up
Whatever the courts say the file means, publishers are posting it at a rate that would have been unthinkable three years ago. The News Homepages archive, which checks the robots.txt of 1,155 news publishers worldwide twice a day, counted 629 of them, 54.5%, blocking at least one of the three major AI-adjacent crawlers as of July 17, 2026. OpenAI's GPTBot and Common Crawl's CCBot are each blocked by 49.8% of the sample; Google's AI-training crawler, Google-Extended, by 44.8%. Half the professional news web has now put up the sign. The December rulings decided, jurisdiction by jurisdiction, whether the sign is a request or a right.
View data table
| Crawler | Publishers blocking | Share |
|---|---|---|
| Any of the three below | 629 | 54.5% |
| GPTBot (OpenAI) | 575 | 49.8% |
| CCBot (Common Crawl) | 575 | 49.8% |
| Google-Extended | 518 | 44.8% |
What I tell people in diligence now
I run crawlers for a living, and I sit on the vendor side of data-diligence calls, so here is the operational translation. Robots.txt used to be one question on a compliance checklist. It is now a jurisdictional fork in the supply chain of every dataset that touched the open web. Data collected for EU-bound AI training against a machine-readable reservation carries a defect that a regulator, not just a plaintiff, can price from next month. Data collected in the US past a robots.txt disallow is not "circumvention", after December, but the surviving claims in the same lawsuit are a reminder that it is not a safe harbor either, and contract and copyright theories are still very much alive.
So the diligence question has changed shape. It used to be "do you respect robots.txt", a yes/no that told you almost nothing. The version worth asking now is: can your vendor show you, per source and per date, what the machine-readable signals said when the data was collected, and what the collection system did about them. That is a logging problem, and operators either built the log or they did not. A "keep off the grass" sign binds nobody in Manhattan and binds a foundation-model provider in Brussels. The file in the middle did not change at all. The record of how you read it is becoming the product.