Forget security – Google's reCAPTCHA v2 is exploiting users for profit | Web puzzles don't protect against bots, but humans have spent 819 million unpaid hours solving them

ForgottenFlux@lemmy.world · 7 months ago

Forget security – Google's reCAPTCHA v2 is exploiting users for profit | Web puzzles don't protect against bots, but humans have spent 819 million unpaid hours solving them

aaaaace@lemmy.blahaj.zone · 7 months ago

Try the headphone option.

brbposting@sh.itjust.works · 7 months ago

Finally heard a clear audio CAPTCHA for the first time in my life this past month. It was glorious. There was slight garbling before and after the characters were read, but that’s it.

Besides that singular experience, all audio CAPTCHAs have been utterly 100% impossible to interpret. Blaring white noise followed by a small squeak of “threeve” or “eleventeen”.

TheObviousSolution@lemm.ee · 7 months ago

Sometimes I think writers just try to find things to be edgy about. The straws this grasps at it are incredible. Might as well complain from the billions of unpaid man hours people provide by providing common courtesy for free.

Blackmist@feddit.uk · 7 months ago

I thought the whole point of reCaptcha was to provide a reliable set of data to train bots. Entering a fuzzy scanned word, identifying bikes and traffic lights, etc.

The fact that they’ve now got that, and the bots are trained is hardly a surprise.

Without captchas the problem of spambots would still be a million times worse.

polonius-rex@kbin.run · 7 months ago

Google should bear the cost of detecting bots, rather than shifting it to users

how?

radivojevic@discuss.online · 7 months ago

Yeah. Written by someone who doesn’t really understand the internet.

siph@lemmy.world · 7 months ago

Considering the article states that reCAPTCHA v2 and v3 can be broken/bypassed by bots 70-100% of the time, they are obviously not the solution.

radivojevic@discuss.online · 7 months ago

“Google should bear the cost”

Google should shut it down and make sites roll their own verification. Give everyone a month to implement a new solution on millions of websites.

Chozo@fedia.io · 7 months ago

Then what is?

siph@lemmy.world · 7 months ago

Maybe a billion dollar company has the budget to come up with something?

Looking at the numbers in this post, reCAPTCHA exists to make Google money, not to keep bots out.

I’d rather have no reCAPTCHA than the current state.

OsrsNeedsF2P@lemmy.ml · edit-2 7 months ago

Hi it’s me. I work for a billion dollar company with a budget. We have no ethical ideas on how to stop bots. Thanks for coming to my tech talk.

siph@lemmy.world · 7 months ago

Yeah, that’s about the way I’d expect it to go.

“Traffic resulting from reCAPTCHA consumed 134 petabytes of bandwidth, which translates into about 7.5 million kWhs of energy, corresponding to 7.5 million pounds of CO2. In addition, Google has potentially profited $888 billion from cookies [created by reCAPTCHA sessions] and $8.75–32.3 billion per each sale of their total labeled data set.”

There might be a tiny chance they’re not interested in changing things.

Anti_Iridium@lemmy.world · 7 months ago

Something something free market?

conciselyverbose@sh.itjust.works · 7 months ago

At what cost?

100% success rate isn’t even moderately useful if it costs $5 per pass. The discussion is completely pointless without a concrete, documented analysis of the actual hardware and energy costs involved.

polonius-rex@kbin.run · 7 months ago

how do you get the metric of 70-100% of the time?

the best bots doing it 70-100% of the time is very different to the kind of bot your average spammer will have access to

siph@lemmy.world · 7 months ago

Did you read the article or the TL:DR in the post body?

The paper, released in November 2023, notes that even back in 2016 researchers were able to defeat reCAPTCHA v2 image challenges 70 percent of the time. The reCAPTCHA v2 checkbox challenge is even more vulnerable – the researchers claim it can be defeated 100 percent of the time.

reCAPTCHA v3 has fared no better. In 2019, researchers devised a reinforcement learning attack that breaks reCAPTCHAv3’s behavior-based challenges 97 percent of the time.

So yeah, while these are research numbers, it wouldn’t be surprising if many larger bots have access to ways around that - especially since those numbers are from 2016 and 2019 respectively. Surely it is even easier nowadays.

polonius-rex@kbin.run · 7 months ago

researchers were able to defeat reCAPTCHA v2 image challenges 70 percent of the time

that doesn’t answer the question?

researchers devised a reinforcement learning attack that breaks reCAPTCHAv3’s behavior-based challenges 97 percent of the time

i’d argue “bespoke system, deployed in a very limited context, built by researchers at the top of their field” is kind of out of reach for most people? and any bot network scaled up automatically becomes easier to detect the further you scale it

the cost of just paying humans to break these already at or below pennies per challenge

IphtashuFitz@lemmy.world · 7 months ago

Don’t know why you’re being downvoted… My employer sees a lot of bot activity on our sites, which are hosted in AWS and protected by Akamai. It’s Akamai that informs us when a bot visits our site, and Akamai that lets us block it. Google never sees this traffic.

Midnight Wolf@lemmy.world · 7 months ago

I thought this was old news 20 years ago?

Etterra@lemmy.world · 7 months ago

We already knew that, but it’s nice re to have data.

sarmale@lemmy.zip · 7 months ago

I thought it was detecting bots based on how you are moving your mouse, etc to solve it, but if they can be solved by AI do they want their AI trained by other AI?

radivojevic@discuss.online · 7 months ago

This is bullshit. Author is literally insane.

HiramFromTheChi@lemmy.world · 7 months ago

There’s nothing that can express my disdain for Google’s reCaptcha.

😒 We’re training its AI models 😒 It’s free labor for Google 😒 Sometimes it wants the corner of an object, sometimes it doesn’t 😒 Wildly inconsistent 😒 Always blurry and hard to see 😒 Seemingly endless 😒 It’s the robot asking us humans if we’re the robots

شاهد على إبادة@lemm.ee · 7 months ago

They were using us to label the data.

Benaaasaaas@lemmy.world · 7 months ago

That’s why you always make sure that labeling is “garbage in” and label whatever

lud@lemm.ee · edit-2 7 months ago

Alright, I don’t use google.com

Edit: this was in reply to someone. I guess my app fucked up the reply.

Rin@lemm.ee · 7 months ago

But you might still be using their captcha

Mubelotix@jlai.lu · 7 months ago

I bypassed 35000 google recaptcha v2 using bots. Don’t ever rely on this for security

Caboose12000@lemmy.world · 7 months ago

Where can I learn this power?

Mubelotix@jlai.lu · 7 months ago

I just spent 3$ worth of bitcoin on NoCaptchaAI. I used their web extension on a server which had a browser opened and controlled by a custom webextension I made so that a solved challenge would be returned to a swarm of clients upon request

Gregor@gregtech.eu · 6 months ago

Your extension is archived, I’d rather not use it.

Mubelotix@jlai.lu · 6 months ago

It’s a custom extension solving my very specific problem on a specific internal website. It was never meant for you to use it, it’s just there to serve as inspiration to others

snooggums@midwest.social · 7 months ago

The conclusion can be extended that the true purpose of reCAPTCHA v2 is a free image-labeling labor and tracking cookie farm for advertising and data profit masquerading as a security service,” the paper declares.

I thought this was known since it came out. It seemed even more obvious when the images leaned in heavily to traffic related pictures like stoplights.

TypicalHog@lemm.ee · 7 months ago

I always thought they are just getting the training data for AI using these.

serenissi@lemmy.world · 7 months ago

The objective of reCAPTCHA (or any captcha) isn’t to detect bots. It is more of stopping automated requests and rate limiting. The captcha is ‘defeated’ if the time complexity to solve it, whether human or bot, is less than what expected. Now humans are very slow, hence they can’t beat them anyway.

smb@lemmy.ml · 7 months ago

[…] reCAPTCHA […] isn’t to detect bots. It is more of stopping automated requests […]

which is bots. bots do automated requests and every automated request doer can also be called a bot (i.e. web crawlers are called bots too and -if kind- also respect robots.txt which has “bots” in its name for this very reason and bots is the shortcut for robots) use of different words does not change reality behind it, but may add a fact of someone trying something on the other.

serenissi@lemmy.world · 7 months ago

There isn’t a good way to classify human users with scripts without adding too much friction to normal use. Also bots are sometimes welcome amd useful, it’s a problem when someone tries to mine data in large volume or effectively DoS the server.

Forget bots, there exist centers in India and other countries where you can employ humans to do ‘automated things’ (youtube like count, watch hour for example) at the same expense of bots. There are similar CAPTCHA services too. Good luck with those :)

Only rate limiting is the effective option.

smb@lemmy.ml · 7 months ago

Only rate limiting is the effective option.

i doubt that. you could maybe ratelimit per IP and the abusers will change their IP whenever needed. if you ratelimit the whole service over all users in the world, then your service dies as quickly into uselessness as effective your ratelimiter is. if you ratelimit actions of logged in users, then your ratelimiting is limited by your ability to identify fake or duplicate accounts, where captchas are not helpful at all.

at the same expense of bots. they might be cheap, but i doubt that anyway, bots don’t need sleep.

i was answering about that wording (that captchas were “not” about bots but about “stopping automated requests”) and that automated requests “are” bots instead.

call centers are neither bots nor automated requests (the opposite IS their advantage) and thus have no relation to what i was specifically saying in reply to that post that suggested automated requests and bots would be different things in this context.

i wasn’t talking about effectiveness of captchas either or if bots should be banned or not, only about bots beeing automated requests (and vice versa) from the perspective of the platform stopping bots. and that trying to use different words for things, (claiming like “X isn’t X, it is really U!”* or automated requests aren’t bots) does not change the reality of the thing itself.

*) unrelated to any (a-)social media platform

serenissi@lemmy.world · 7 months ago

stopping automated requests

yeah my bad. I meant too many automated requests. Both humans and bot generate spams and the issue is high influx of it. Legitimate users also use bots and by no means it’s harmful. That way you do not encounter captcha everytime you visit any google page, nor a couple of scraping scripts gets a problem. Recaptcha (or hcaptcha, say) triggers when there is high volume of request coming from same ip. Instead of blocking everyone out to protect their servers, they might allow slower requests so legitimate users face mininimal hindrance.

Most google services nowadays require accounts with stronger (like cell phone) verification so automated spam isn’t a big deal.

smb@lemmy.ml · 7 months ago

since bots are better at solving captchas and humanoid services exist that solve them, the only ones negatively affected by captchas are regular legitimate users. the bad guys use bots or services and are done. regular users have to endure while no security is added, and for the influx i guess it is much more like with the better lock on the front door: if your lock is a bit better than that of your neigbhour, theirs might be force-opened more likely than yours. it might help you, but its not a real but only relative and also very subjective feeling of 'security".

beeing slower than the wolves also isn’t as bad as long as you are not the slowest in your group (some people say)… so doing a bit more than others always is a good choice (just better don’t put that bar too low like using crowdsnakeoil for anything)

serenissi@lemmy.world · 7 months ago

the bad guys use bots or services and are done. regular users have to endure while no security is added

put in other words, common users can’t easily become ‘bad guy’ ie cost of attack is higher hence lower number of script kiddies and automated attacks. You want to reduce number. These protections are nothing for bitnet owners or other high profile bad actors.

ps: recaptcha (or captcha in general) isn’t a security feature. At most it can be a safety feature.

tb_@lemmy.world · 7 months ago

I thought captcha’s worked in a way where they provided some known good examples, some known bad examples, and a few examples which aren’t certain yet. Then the model is trained depending on whether the user selects the uncertain examples.

Also it’s very evident what’s being trained. First it was obscured words for OCR, then Google Maps screenshots for detecting things, now you see them with clearly machine-generated images.

nickwitha_k (he/him)@lemmy.sdf.org · 7 months ago

There are much better ways of rate limiting that don’t steal labor from people.

serenissi@lemmy.world · 7 months ago

hCaptcha, Microsoft CAPTCHA all do the same. Can you give example of some that can’t easily be overcome just by better compute hardware?

nickwitha_k (he/him)@lemmy.sdf.org · 7 months ago

The problem is the unethical use of software that does not do what it claims and instead uses end users for free labor. The solution is not to use it. For rate limiting a proxy/load-balancer like HAProxy will accomplish the task easily. Ex:

Forget security – Google's reCAPTCHA v2 is exploiting users for profit | Web puzzles don't protect against bots, but humans have spent 819 million unpaid hours solving them

Forget security – Google's reCAPTCHA v2 is exploiting users for profit | Web puzzles don't protect against bots, but humans have spent 819 million unpaid hours solving them

Google's reCAPTCHA v2 just labor exploitation, boffins say