News: 1691527293

  ARM Give a man a fire and he's warm for a day, but set fire to him and he's warm for the rest of his life (Terry Pratchett, Jingo)

How to identify OpenAI's crawler bot to stop it slurping websites for training data

(2023/08/08)


OpenAI, the maker of machine learning models trained on public web data, has published the specifications for its web crawler so that publishers and site owners can opt out of having their content scraped.

The newly released [1]technical document describes how to identify OpenAI's web crawler GPTBot through its user agent token and string, which get emitted by the company's software in the HTTP request header sent to ask a server for a web page.

Web publishers can thus add an entry into their web server's [2]robots.txt file to tell the crawler how it should behave, assuming GPTBot was designed to heed the [3]Robots Exclusion Protocol – not all bots do so. For example, the following set of robots.txt key/value pairs would instruct GPTBot to stay out of the root directory and everything else on the site. User-agent: GPTBot

Disallow: /

However, OpenAI insists that allowing its bot to collect site data can improve the quality of AI models the biz builds and scraping can be done without gathering sensitive information – for which OpenAI and Microsoft [4]were recently sued .

"Web pages crawled with the GPTBot user agent may potentially be used to improve future models and are filtered to remove sources that require paywall access, are known to gather personally identifiable information (PII), or have text that violates our policies," the ML super-lab's documentation reads.

Allowing GPTBot to access your site can help AI models become more accurate and improve their general capabilities and safety

"Allowing GPTBot to access your site can help AI models become more accurate and improve their general capabilities and safety."

And who wouldn't want to save OpenAI the time and expense of making its models more capable and less risky?

[5]

Even so, OpenAI's acknowledgement that it trains its large language models on the public internet has coincided with efforts by organizations to limit automated access to information via the web. AI software makers enjoy grabbing all kinds of info from sites to train their models to bank millions if not billions of dollars in revenues. Some businesses are putting their foot down, and closing off access if they're not going to get a cut of that income.

[6]

[7]

Reddit, for example, recently [8]changed its API terms to better enable the company to monetize the content created free-of-charge by its users. And Twitter recently [9]sued four unidentified entities to prevent site data from being scraped for AI training.

Unleash the legal eagles!

OpenAI did not immediately respond to a request to explain why it published details about GPTBot. But it may not be a coincidence that there have been several recent lawsuits filed against the Microsoft-championed biz for allegedly using publicly accessible data without consent, or in contravention of stated licensing terms.

Beyond the privacy lawsuit noted above, OpenAI, Microsoft, and the latter's GitHub subsidiary [10]were sued in November for allegedly ingesting license-encumbered source code to train OpenAI's Codex model, and then reproducing that code through GitHub's Copilot source-suggestion service. Several book authors last month [11]filed a similar lawsuit alleging OpenAI trained ChatGPT on their work without permission.

Google, DeepMind, and parent Alphabet [12]have also been sued over similar claims.

[13]

Given the [14]legal uncertainty arising from scraping public data and using that info to train AI models, it's perhaps unsurprising that Google – an OpenAI rival – last month proposed [15]rethinking how the Robots Exclusion Protocol works.

[16]OpenAI pulls AI text detector due to it being a bit crap

[17]ChatGPT's odds of getting code questions correct are worse than a coin flip

[18]AI on AI action: Googler uses GPT-4 chatbot to defeat image classifier's guardian

[19]How to make today's top-end AI chatbots rebel against their creators and plot our doom

Israel Krush, CEO and co-founder of Hyro, which makes an AI assistant for the healthcare industry, told The Register there are two main issues with the way web crawling works.

"Firstly, the default setup involves publishers having to actively opt out if they don't want their websites to be crawled and used for fine-tuning," he said. "This process is quite different from how search engines operate, where crawling serves as a reference to direct users to the publishers' sites.

"With OpenAI and AI assistants, the content becomes a direct part of the product, which could sometimes lead to inaccuracies. The fact that publishers have to opt out raises a big concern."

Krush said integrating this content into someone else's product and potentially changing it raises another potential issue.

Microsoft Azure OpenAI lets enterprises feed corporate secrets to ChatGPT [20]READ MORE

"The second problem is with OpenAI's statement about excluding websites 'known for using Personally Identifiable Information (PII),'" he said. "This statement is a bit puzzling."

"Take news publishers, for instance; they naturally include some identifiable information. Even websites that aren't specifically thought of as holding PII might still have some. Any content involving PII needs to be properly redacted."

[21]

Krush argued that compliance concerns and responsible model use require stronger safeguards, noting that his own firm only scrapes data with explicit permission and handles personal information appropriately.

"Instead of just focusing on scraping websites already flagged for PII, OpenAI should assume there's potential for PII across all sites, particularly with publishers," he said. "They should take proactive steps to make sure the scraped info aligns with compliance rules." ®

Get our [22]Tech Resources



[1] https://platform.openai.com/docs/gptbot

[2] https://developers.google.com/search/docs/crawling-indexing/robots/create-robots-txt

[3] https://www.rfc-editor.org/rfc/rfc9309.html

[4] https://www.theregister.com/2023/06/28/microsoft_openai_sued_privacy/

[5] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=2&c=2ZNK7Awr1ElK92yHRnUakeQAAAJI&t=ct%3Dns%26unitnum%3D2%26raptor%3Dcondor%26pos%3Dtop%26test%3D0

[6] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44ZNK7Awr1ElK92yHRnUakeQAAAJI&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0

[7] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33ZNK7Awr1ElK92yHRnUakeQAAAJI&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0

[8] https://www.reddit.com/r/redditdev/comments/14nbw6g/updated_rate_limits_going_into_effect_over_the/

[9] https://www.theregister.com/2023/07/14/twitter_scraping_lawsuit/

[10] https://www.theregister.com/2023/06/09/github_copilot_lawsuit/

[11] https://www.theregister.com/2023/07/10/in_brief_ai/

[12] https://www.theregister.com/2023/07/12/google_alphabet_deepmind_bard_complaint/

[13] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44ZNK7Awr1ElK92yHRnUakeQAAAJI&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0

[14] https://blog.ericgoldman.org/archives/2022/12/hello-youve-been-referred-here-because-youre-wrong-about-web-scraping-laws-guest-blog-post-part-2-of-2.htm

[15] https://blog.google/technology/ai/ai-web-publisher-controls-sign-up/

[16] https://www.theregister.com/2023/07/26/openai_ai_classifier/

[17] https://www.theregister.com/2023/08/07/chatgpt_stack_overflow_ai/

[18] https://www.theregister.com/2023/08/01/google_boffin_breaks_ai_model/

[19] https://www.theregister.com/2023/07/27/llm_automated_attacks/

[20] https://www.theregister.com/2023/06/22/microsoft_azure_ai_data/

[21] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33ZNK7Awr1ElK92yHRnUakeQAAAJI&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0

[22] https://whitepapers.theregister.com/



Or alternatively...

Jamie Jones

Redirect such requests to pages and pages of nonsense. If you don't have any, just slurp just about any facebook or twitter feed.

Re: Or alternatively...

stiine

But I don't have 10PB of random words that I can send them....

The risk with Robots.txt

Len

Can someone enlighten me?

Robots.txt has been used for/against search engine bots for ages. It has always come with the warning that if you don't want people/bots/crawlers to know that directory B exists you should not explicitly allow crawling of directory A while explicitly blocking crawling of directory B. You're asking not to look somewhere and only scrupulous bots would honour that. The solution is usually to sort it out at page level so a page you don't want to be crawled has a meta tag blocking it.

How would one do this with the ChatGPT bot? Will it also look at the page meta tags? I don't trust some makers of AI bots to not use a block attempt to explicitly go and harvest data from directories that I have disallowed. It might just give them an edge over OpenAI.

How could I let ChatGPT freely crawl my FAQ or About Us page but not my Content page?

Freedom begins when you tell Mrs. Grundy to go fly a kite.