Update: October 2026
Web scraping is legal when the data isn’t personal, the site’s terms allow it, and you don’t infringe any copyright or database right. When the data includes personal information about people in the EU, the GDPR applies even if the information is public, and you need a lawful basis and the right safeguards.
The picture changed in July 2026, when the European Data Protection Board (EDPB) adopted draft Guidelines 03/2026 on web scraping in the context of generative AI. This guide explains the tests a scraper must pass, what the draft guidelines say, and how contract law and court rulings add separate risks.
• The GDPR applies to web scraping whenever personal data is collected, stored or used, and publicly available data does not fall outside it. A scraper needs a lawful basis. The EDPB’s draft guidelines say legitimate interests are often used and that consent would most likely not apply, because scrapers cannot obtain it from everyone at scale.
• The EDPB’s draft Guidelines 03/2026 (adopted 7 July 2026, open for comment until 30 October 2026) recommend collecting only what is needed, excluding sites that block scraping through robots.txt, ai.txt or CAPTCHAs, and publishing privacy information about the scraping.
• Scrapers must still tell people their data is being collected. Article 14 requires privacy information, the disproportionate-effort exception is narrow, and infringements of data subject rights can be fined up to €20 million or 4% of worldwide annual turnover.
No single law decides whether scraping is legal. Each of the following can apply independently, and passing one does not clear the others.
Data protection law applies if the data relates to an identifiable person. Contract law applies if the site’s terms forbid scraping and you agreed to them. Copyright and database rights apply if you copy protected content. Scraping that avoids personal data and respects a site’s terms and technical limits, for example, collecting product prices, sits at the low-risk end. Scraping names, photos or profiles at scale sits at the high-risk end.
The GDPR defines processing to include collection, storage, organisation and retrieval. The EDPB’s draft guidelines confirm that web scraping is covered whenever it includes those operations on personal data. Data someone publishes openly remains personal data, and the person who scrapes it becomes a controller with full GDPR duties.
The draft guidelines describe scraping as a large-scale activity that often takes place without the people concerned knowing. The EDPB says that creates particular risks for their rights. They apply to private entities scraping data to train generative AI models, which is the setting the document covers. The reasoning on lawful basis, transparency and minimisation also gives a clear picture of how regulators approach any large-scale scraping of personal data.
The same logic applies to a sales team scraping contact details from professional profiles and to a researcher building a dataset from public posts. If the data can identify a person, the GDPR applies.
Consent is unworkable for most scraping. A scraper has no direct relationship with the people whose data it collects and cannot ask millions of them for permission. The EDPB draft guidelines say consent would most probably not be an applicable legal basis, and that a missing robots.txt file does not amount to consent. They note that private entities often use legitimate interests under Article 6(1)(f) to scrape data for generative AI training.
Legitimate interests require three conditions to be met together. The controller or a third party must pursue a legitimate interest. Processing personal data must be necessary for that interest. The interests or fundamental rights of the people concerned must not override it; this is the balancing test.
The draft guidelines treat the balancing test as the hardest part. Among the factors, they say people’s reasonable expectations matter. If a site uses robots.txt, ai.txt files, authentication or CAPTCHAs to keep scrapers out, and people know it does, they are less likely to expect a scraper to process their data. Safeguards can tip the balance. The EDPB’s examples include excluding certain data categories and sources by default, limiting collection to freely accessible data, deleting or anonymising data quickly, and making it easy for people to exercise their rights.
The draft guidelines are open for comment until 30 October 2026, and the final text may change. They already show where the EDPB expects organisations to end up.
The draft guidelines tie data minimisation to the scraper’s design. Before collecting, a controller should consider whether synthetic data could replace personal data, define precise collection criteria, map the data it expects to collect, and apply filters that exclude categories of data it does not need. After collecting, the draft points to filters that catch data identifiable by its format, such as telephone or social security numbers, and to replacing real data with synthetic data or anonymising or pseudonymising it where feasible.
The draft guidelines recommend excluding websites that clearly oppose scraping through technical measures such as authentication, robots.txt or ai.txt files, or CAPTCHAs. They also recommend excluding personal data on sites accessible only after logging in, such as social network content that requires an account.
In practice, a scraper that ignores these signals weakens its legitimate interests argument because the balancing test considers what people could reasonably expect.
Processing special categories of personal data, such as health information, is prohibited unless an Article 9(2) derogation applies, in addition to an Article 6 basis. The EDPB says the Court of Justice’s ruling in GC and Others (C-136/17) can be relevant to incidental and residual collection of such data during scraping. That applies where the controller has put technical and organisational measures in place to prevent collecting and spreading it, and the EDPB asks for a case-by-case assessment.
When a controller does not collect personal data directly from the person, Article 14 still requires it to provide privacy information. For scraping, individual notices are often impossible. Article 14(5)(b) allows an exception where informing people proves impossible or would take disproportionate effort.
The draft guidelines say controllers should not rely on that exception routinely. They should weigh the effort of informing people against the impact of not informing them, and assess this for the dataset as a whole, taking into account factors such as the number of people and the age of the data. Where an exception applies, the controller must still make the information public. The EDPB says the notice should include the categories of data, the purposes, the legal basis and a precise indication of the sources. Other safeguards it names include carrying out a data protection impact assessment and making it public, minimising the data and the storage period, and applying strong security.
A scraper can comply with the GDPR and still break a contract. Two cases show how courts have handled it.
In Ryanair v PR Aviation (C-30/14, 15 January 2015), a flight-comparison site collected Ryanair’s fares and schedules automatically, and Ryanair’s terms banned automated extraction. The Court of Justice held that when a database falls outside the Database Directive’s protection, its owner can, subject to national law, set contractual terms for its use. In plain terms, a site’s terms can restrict scraping of data that no intellectual property right protects.
In the US case hiQ v LinkedIn, the Ninth Circuit found that hiQ had raised at least serious questions that scraping public LinkedIn profile data was lawful under the Computer Fraud and Abuse Act. That is the part most often quoted. In November 2022, the district court ruled that hiQ’s scraping and its use of fake accounts breached LinkedIn’s user agreement, subject to hiQ’s equitable defences. The case ended in December 2022 with a consent judgment: $500,000 against hiQ, a permanent injunction, and an order to delete its LinkedIn scraping code and all LinkedIn member profile data it held. The parties stipulated to those terms, so they don’t set a precedent.
Courts can enforce a site’s terms as a contract even when the data is public.
Copying text, images or a substantial part of a protected database can infringe copyright or the sui generis database right under Directive 96/9/EC, whether or not personal data is involved. Rights holders can also reserve their works from text and data mining under Directive 2019/790, and the EDPB draft includes a favourable balancing test that excludes content where rights holders have objected. These rights sit outside the GDPR, so a scraping project needs a separate review.
The clearest enforcement example is Clearview AI, a US company that scraped facial photos from the internet and turned them into biometric codes. The Dutch Data Protection Authority fined it €30.5 million in September 2024. It found that Clearview should never have built the database, gave people too little information about its use of their data, and did not cooperate with access requests. It ordered the company to stop, with penalties of up to €5.1 million on top of the fine. The authority also said that using Clearview’s services is illegal. Clearview did not object, so it could not appeal.
Article 83 sets the maximum for infringements of the data protection principles, lawful basis rules and data subject rights at €20 million or 4% of worldwide annual turnover, whichever is higher.
Before any project that touches personal data, write down the purpose and the legitimate interest, and test whether you need personal data at all. Map the sources and exclude sites that block scraping, require a login or hold sensitive content. Check each source’s terms and any copyright or database rights. Decide how you will inform people and whether the Article 14(5)(b) exception applies. Run a data protection impact assessment where the scale or the data types make a high risk likely, and keep the record. For AI training in particular, the EDPB draft adds measures to prevent the model from memorising and regurgitating personal data.
GDPRLocal advises on lawful bases, DPIAs and AI governance for projects that collect data at scale. See our GDPR consultant services and AI governance services, or read our guide to the legal and privacy challenges of data scraping.
Whether scraping is legal depends on four things: what data you collect, which lawful basis you use, what the source’s terms and technical signals say, and whom you tell. The GDPR allows scraping based on a documented legitimate interest, minimal collection, and public information about what you do.
The EDPB’s draft guidelines are open for comment until 30 October 2026, and the final text will set the benchmark for regulators across the EU. An organisation that can show documented interest, a narrow collection scope, and respect for sites that block scrapers will be well placed when the guidelines are finalised.
No single rule makes scraping legal or illegal. Non-personal public data is generally low risk if you respect the site’s terms and copyright. Personal data needs a lawful basis under the GDPR, usually legitimate interests, plus safeguards.
Information that identifies a person remains personal data when published openly. The person who scrapes it is a controller and must meet the GDPR’s duties, including lawful basis, minimisation and transparency.
Consent is rarely practical, because scrapers have no direct relationship with the people concerned. The EDPB’s draft guidelines describe legitimate interests as the basis most often used, and they require a documented balancing test.
Disclaimer: This blog post is intended solely for informational purposes. It does not offer legal advice or opinions. This article is not a guide for resolving legal issues or managing litigation on your own. It should not be considered a replacement for professional legal counsel and does not provide legal advice for any specific situation or employer.