Google reCAPTCHA service under the microscope: Questions raised over privacy promises, cookie use
- Reference: 1604304912
- News link: https://www.theregister.co.uk/2020/11/02/google_ad_recaptcha/
- Source link:
The v2 update in 2014 added an iframe or HTML Inline Frame, which is a way of embedding one web page in another. Then there was the v3 update in 2018, which added machine learning to the mix, to reduce the need for interaction with bot detection challenges.
reCAPTCHA makes it possible for the internet giant to challenge netizens to prove they are real people, by completing picture puzzles and the like, while providing plumbing to potentially funnel information about folks into its advertising business. Google insists it doesn't use reCAPTCHA data for personalized adverts, and says as much in the reCAPTCHA terms of service.
Yet the Silicon Valley corp's fine-print and other disclosures stop short of saying reCAPTCHA is completely quarantined from all ad-related data collection. And privacy researchers now argue that the company needs to clarify that point.
Zach Edwards, co-founder of web analytics biz Victory Medium, found that Google's reCAPTCHA's JavaScript code makes it possible for the mega-corp to conduct "triangle syncing," a way for two distinct web domains to associate the cookies they set for a given individual. In such an event, if a person visits a website implementing tracking scripts tied to either those two advertising domains, both companies would receive network requests linked to the visitor and either could display an ad targeting that particular individual.
Two different domains generally shouldn't have access to the same set of cookie data, based on the distinction between first-party and third-party resources in the web browser security model. But triangle syncing dissolves that separation.
Triangle of ad success?
"Triangle syncs expand an advertising universe and make it possible to target someone across more domains," Edwards told The Register .
It's a common practice in advertising, he said, so that two separate companies with two separate domains can share data, such as the identifiers associated with a particular individual. And it's also done within a single company like Google that operates more than one domain and wants to track internet users across the different domains.
"So reCAPTCHA's gstatic.com domain doing a triangle sync to google.com basically ensures that a user can be found/tracked if either of those domains is embedded into a website," Edwards said.
Cloudflare dumps Google's reCAPTCHA, moves to hCaptcha as free ride ends (and something about privacy) [1]READ MORE
According to Google, the company doesn't use reCAPTCHA for triangle syncing and reCAPTCHA loads static resources from two places on gstatic.com , with no cookies written or read. No triangle request or sync is done as part of this process, we were told. And the gstatic.com domain is supposedly "cookieless," in that it has been designed to be unable to collect cookie data.
Yet, reCAPTCHA JavaScript [2]code hosted at Google's gstatic.com domain includes multiple references to cookies. And visiting a web page embedded with a reCAPTCHA widget does set a google.com [3]"NID" preference cookie , even if you try to block third-party cookies.
Edwards says what's going on isn't typical triangle syncing. He says if you embed a reCAPTCHA on a site like ncrts.com , for example, the gstatic.com requests then redirect to a new request to google.com and then google.com sets its cookie. "It's a triangle sync not in a traditional cookie match sync on both sides, but in a request + cookie match," he said.
He also points out that [4]Google's privacy policy identifies the gstatic.com domain specifically as one of many domains used to set cookies for its advertising products.
Google maintains gstatic.com doesn't read or write cookies, but it appears the domain invites google.com to set them.
T&Cs
Edwards argues Google isn't being straightforward about how it handles cookies, noting that in a Safari browser test he conducted, the Google domain sets session keys, a form of temporary browser data storage linked to a server, instead of cookies.
Google's reCAPTCHA terms of service state that the service sends device and application data to the company. It specifies how it handles that data thus: "The information collected in connection with your use of the service will be used for improving reCAPTCHA and for general security purposes. It will not be used for personalized advertising by Google."
The Register specifically asked Google whether reCAPTCHA data might be used for some aspect of the ad business other than personalized advertising. It might, for example, be helpful to fight ad fraud.
Google's spokesperson cited the policy spelled out above – the data improves reCAPTCHA and may be used for general security purposes, whatever that means.
Via Twitter, Ashkan Soltani, a privacy researcher and former Federal Trade Commission technologist, said what Google is doing looks a lot like what the company did in 2011 and 2012 to bypass Safari's third-party cookie blocking.
Here’s a quick clip showing how Google’s ReCaptcha sets a 3rd party cookie even when [5]@Mozilla [6]@Firefox is set to block “cross-site tracking cookies” [7]#privacy [8]#cookiewars [9]pic.twitter.com/vWgibJ20ty — ashkan soltani (@ashk4n) [10]October 30, 2020
In 2012, America's consumer watchdog the FTC [11]fined Google $22.5m for misrepresenting to Safari users that it would not place tracking cookies.
Solanti also suggested Facebook's [12]2019 settlement with the FTC may be relevant. In that case, Facebook was penalized for collecting data for one purpose (security) and also using it for another (ads).
In an email to The Register , Soltani said he had tested Edwards's claims and confirmed that reCAPTCHA sets google.com cookies even when the user's browser has been configured to block third-party cookies.
He subsequently posted the video depicting the network requests from visiting the hubspot.com/abuse-complaints page, which calls a google.com -hosted reCAPTCHA script that runs gstatic.com -hosted code for invoking a reCAPTCHA puzzle.
Discussing what was going on, Soltani said the main issue is whether those who rely reCAPTCHA for security are exposing users to profiling by Google for the purpose of advertising.
Google's privacy disclosures may be adequate to cover reCAPTCHA's role if it were found to play a role in the company's ad business. Google does disclose that it sets advertising cookies via its gstatic.com domain.
Data CAPTCHA
Edwards however argues that Google hasn't been sufficiently clear that reCAPTCHA uses this domain.
"It's problematic for publishers who care about user privacy," he said, because if you implement reCAPTCHA on your website and don't disclose that you set google.com cookies, that runs the risk of violating some aspects of the "right to know" requirement under the California Consumer Privacy Act.
Edwards contends that websites in Europe will need to rethink how they use reCAPTCHA for bot defense.
"In my opinion, organizations in Europe that use reCAPTCHA for spam protection now need to move reCAPTCHA behind their consent walls," he said.
"It's a huge stretch to call syncing cookies to google.com mandatory in any way, and it doesn't seem possible to deploy reCAPTCHA in any way anymore that doesn't do that sync."
Google already recommends that in reCAPTCHA's terms of service, which state, "For users in the European Union, you and your API Client(s) must comply with the [13]EU User Consent Policy ." ®
Get our [14]Tech Resources
[1] https://www.theregister.com/2020/04/09/cloudflare_dumps_recaptcha/
[2] https://www.gstatic.com/recaptcha/releases/T9w1ROdplctW2nVKvNJYXH8o/recaptcha__en.js
[3] https://policies.google.com/technologies/types?hl=en-US
[4] https://policies.google.com/technologies/types?hl=en-US
[5] https://twitter.com/mozilla?ref_src=twsrc%5Etfw
[6] https://twitter.com/firefox?ref_src=twsrc%5Etfw
[7] https://twitter.com/hashtag/privacy?src=hash&ref_src=twsrc%5Etfw
[8] https://twitter.com/hashtag/cookiewars?src=hash&ref_src=twsrc%5Etfw
[9] https://t.co/vWgibJ20ty
[10] https://twitter.com/ashk4n/status/1322260097716793344?ref_src=twsrc%5Etfw
[11] https://www.ftc.gov/news-events/press-releases/2012/08/google-will-pay-225-million-settle-ftc-charges-it-misrepresented
[12] https://www.ftc.gov/news-events/blogs/business-blog/2019/07/ftcs-5-billion-facebook-settlement-record-breaking-history
[13] https://www.google.com/about/company/user-consent-policy.html
[14] https://whitepapers.theregister.com/
Google may have been caught not respecting the policies they pubish but those are not necessarily the policies they operate under.
They should be broken up to make way for someone else to slurp our data!
In a way this is all encouraging for me. Pretty much everywhere I go that uses reCRAPCHA I get asked to identify the buses/hydrants/bikes/hills/bridges every single time. Having to do this every single time is annoying, but is balanced by the fact that it tells me that the Google monster is having trouble tracking me with my choice of VPN, browser and add-ons.
Re: In a way this is all encouraging for me
Definitely this, if you have your browser relatively locked down, recaptcha will ask every single time for image verification on sites you visit repeatedly. This should indicate that they can't fingerprint you so have to repeatedly verify you're not just a random bot.
Of course, the conspiracy version is that they can still track people running noscript / ublock etc just fine. They're only pretending they can't to make you feel safe....
Re: In a way this is all encouraging for me
Unfortunately I believe they managed to grab a full set on info on me when I renewed my subscription to Tin Foil Beanie Monthly! I had to take mine off to get measurements while doing the renewal, and they must have read everything they needed while it was off!
CAPTCHA is a PITA
and the only reason that google.com isn't blocked at my firewall.
Refuzniks of the world unite - avoid Google, Facebook, Amazon etc.
hCaptcha
In my quest to reduce my (and that of my users) exposure to Google I have started migrating to [1]hCaptcha . I can recommend it, it works.
Still looking for a European alternative so no data gets send to the other side of the planet at all but this is a start.
[1] https://www.hcaptcha.com/
Helping Google
Privacy aside, the reason I hate them is I'm effectively helping google build it's AI by finding fire hydrants, traffic lights, etc. They have effectively built the world's largest free human labour image recognition machine - and that's what gets my goat most.
Re: Helping Google
Along with the assumption that the entire world speaks or translates to American English. What the hell is a cross walk?
I know it is obvious but they have so much "intelligence" with language settings and tracking then have the courtesy to make the stupid puzzle appropriate to the region the request is coming from.
Google are lying
I can prove Google are lying, when I need to log in to a site using recaptcha, but also when I use two different accounts to do so.
One is a personal account, I have a dedicated email address for it used only there, not used for anything but this one site. As a general standard, I delete cookies from all sites I visit after I leave (I have about 6 sites that I allow primary domain cookies, but block all third party) . I have blocked all Google domain and advertising sites for years and if a site uses resources from a Google domain like fonts or scripts and it doesn't work with the blocks in place, I go else where. I do not use any Google services, except my Android phone does have an account that's not used anywhere else and that doesn't work too well as it likes to continually nag me, as though there is a virus running, "App permission management is running" and I have blocked as much as I can there too.
The second account is a work colleague's address. Although he occasionally does add blocking, he does little else and remains logged in to many sites/services.
When I use the target site, I have to allow Google.com and gstatic.com (used to be recaptcha.net and gstatic,com, I wonder why that changed?). I login, order some parts I need, log out, delete cookies and re-block the two domains and the main site I just used. I do this whether I use my account or my colleagues, there is NO data on my machine to show if I have visited before and that I passed the recaptcha.
When I login as me, I need to select the images to prove I'm not a bot, when I use my colleague's account, all I need to do is check the box. This is from the same Linux PC, both accounts, I don't use my colleagues machine to login.
How do they know it's a human when I use my colleague's email, but not when I use my mainly Google protected email, I have to be tested? Where are they getting the data from to decide my colleague with lots of Google data is human, but my low profile account needs a bot test?
The simplest explanation is that as on other occasions, Google are lying, they are using their trove of personal data and making the experience of non Googled people worse. I'd give this as evidence in a legal hearing.
As I understand it, reCAPTCHA only issues an actual puzzle to users it has reasons to doubt. In most cases, I only have to click a checkbox marked I am not a robot; and in a few rare cases, like if I'm in incognito mode, I have to click the parts of the image containing a car or something. Which clearly means that Google is fingerprinting me in order to guess whether I'm a human or not, and is storing information over time to facilitate the fingerprinting. It's not possible to get out of that. And ultimately, data that they have is data that they can use for ads; and we're supposed to "trust" them that they don't. Privacy policies are very nice and all, but they've been caught not respecting their own policies, haven't they? Though I guess whatever data they could get through reCAPTCHA would be rather insignificant compared to the firehose the vast majority of users sends to them anyway...