
Anyone who has spent the last few years watching large language models grow ever hungrier for training data has probably wondered, at least in passing, whether the trillions of words scraped from the open web actually comply with European data protection law. On 8 July 2026, the European Data Protection Board answered that question with unusual clarity. The EDPB adopted new guidelines, formally titled Guidelines 03/2026, stating in plain terms that generative AI web scraping falls squarely within the scope of the GDPR whenever the material collected includes identifiable personal data. For an industry that has often treated “publicly available” as a synonym for “fair game,” this is a meaningful correction.
The guidelines were released alongside two related documents: a revised framework on anonymisation, updated to reflect a September 2025 ruling from the Court of Justice of the EU, and the final version of the Board’s long-awaited blockchain guidance. But it is the web scraping text that will land hardest on the AI industry, because it addresses a practice that sits at the very foundation of how modern foundation models are built. Rather than issuing a blanket prohibition, the EDPB has chosen to walk through the entire lifecycle of a scraping operation and identify where GDPR obligations attach, from the initial crawl to the eventual outputs of a trained model. That lifecycle approach is deliberate: regulators have grown tired of enforcement actions that focus only on the moment of collection while ignoring what happens to the data afterward.
The first thing any organisation engaged in generative AI web scraping needs to establish, according to the guidelines, is its role under the regulation. Companies scraping data to train their own models, those fine-tuning a third-party model on scraped content, and those simply purchasing a scraped dataset from a vendor may all end up classified differently as controllers, processors, or joint controllers, and each status carries distinct obligations. The EDPB is explicit that this determination cannot be waved away with generic disclaimers; it requires a genuine, documented assessment of who decides the purposes and means of processing. This matters enormously for the smaller companies now building products on top of foundation models, many of which assumed that responsibility for the underlying training data stopped with the model provider.
On the question of lawful basis, the guidelines confirm what most privacy lawyers had already suspected: legitimate interest, not consent, will be the basis most organisations lean on to justify large-scale generative AI web scraping, since asking billions of website authors for individual permission is plainly unworkable. But the Board is not handing out a free pass. It expects a detailed, case-by-case legitimate interest assessment that weighs the necessity of the scraping against the rights and reasonable expectations of the people whose words, photos, and comments end up in the dataset. Vague references to “innovation” or “the public interest in AI development” are unlikely to satisfy that balancing test on their own. Where scraped content incidentally captures special category data, such as information revealing health conditions, political opinions, or sexual orientation, the guidelines require both an Article 6 lawful basis and a separate Article 9(2) exception, drawing on the CJEU’s earlier ruling in the GC and Others case about incidental collection of sensitive data during otherwise legitimate processing.
Transparency is the third pillar, and here the EDPB acknowledges a genuine practical tension. Individually notifying every person whose data might appear somewhere in a scraped web corpus is, in most cases, simply not feasible. The Board’s compromise is to require organisations to publish clear, accessible privacy information describing their scraping practices, combined with “additional appropriate measures and safeguards” that go beyond a buried clause in a terms-of-service document. Data minimisation runs through the entire text as well: the guidelines call for filtering mechanisms applied before collection even begins, exclusion lists for categories of sensitive websites, and continuous controls throughout the data lifecycle, including anonymisation or the use of synthetic data as substitutes for raw personal information wherever that is technically achievable. For any residual sensitive data that slips through collection filters, the Board wants to see deletion mechanisms and, notably, filters applied at the model output stage as a final backstop.
It’s worth pausing on why this guidance lands now. The Board’s own recent record shows a steady drumbeat of AI-adjacent enforcement and rulemaking: the anonymisation update issued the same week responds directly to a 2025 CJEU judgment tightening the definition of when data truly ceases to be personal, built around a three-part test of whether a record can be isolated, linked to other records, or used to infer information about an individual. Meanwhile national authorities have kept up their own pressure independently of Brussels; France’s CNIL fined the health-data broker IQVIA five million euros earlier this year over pseudonymisation failures, and the EDPB itself spent the same week ordering Belgium’s data protection authority to properly examine a long-running noyb complaint about deceptive cookie banners rather than settling it quietly. Taken together, these actions describe a regulatory environment where AI companies can no longer assume that data practices developed for a pre-generative-AI internet will survive scrutiny unchanged.
For now, the web scraping guidelines remain in draft form, with the EDPB inviting public comment through 30 October 2026, which gives AI developers, publishers, and civil society groups a real window to shape the final text before it hardens into the kind of reference document national regulators cite in enforcement decisions. Organisations that scrape data for AI training, whether for chatbots, image generators, or narrower fine-tuning projects, would be well advised to use that window productively: mapping their own scraping pipelines against the lifecycle the EDPB describes, documenting legitimate interest assessments now rather than after a complaint arrives, and testing whether their anonymisation techniques would actually satisfy the new three-part test. Generative AI web scraping was always going to attract regulatory attention eventually; the only real surprise is how specific and operational the EDPB’s first comprehensive answer has turned out to be.