8 ms·
CAPTCHAs can still detect AI agents
- nicman23 4mo agoyeah no. it is funny easy to make a mcp server and plug a qwen3.6 to it. it was more annoying to convince the llm that it can clear captchas than the actual passing
- BiteCode_dev 4mo agoUntil they learn to do that. So cat and mouse. So nothing new.
- catsrus 4mo agothink the point is that they can't just "learn to do that", because to do so would mean solving human mind (that famously hasn't been going well)
- sigbottle 4mo agoWell no, the idea is a tradeoff between interfaces and telemetry. OK, the agents don't click in the same way as humans. You learn that, what about mouse hovering telemetry, time spent, etc. And one of the most extreme is to force biometrics - a lot of telemetry, breaks the interface a lot - but hey, you have assurance. And none of these tradeoffs require understanding the deep processes of the human mind. Just, map is not the territory, how you do game the map harder and harder and how do the mapmakers respond to that?
- catsrus 4mo agodid you look at the paper? they specifically look at mini tasks with cognitive processes (Eg what dictates the strategy of how people solve tasks)
- CamperBob2 4mo agoLLMs can solve original math problems at the IMO level and beyond, and you might be talking to one now. I don't think they are going to have problems with any CAPTCHA short of separate device attestation. Whatever mechanism the paper proposes, rest assured it can be trained on.
- dpoloncsak 4mo agountil Google trains an AI model off that data, too
- timshell 4mo agoI mean, their CAPTCHAs presumably have tons of data collected over the years, and they can't detect a pretty clear AI agent here: https://www.youtube.com/watch?v=UeTpCdUc4Ls https://www.youtube.com/watch?v=UeTpCdUc4Ls
- arbol 4mo agoThey already have. Claude and OpenAI are not trying to write captcha-defying AI agents. These tests wouldn't hold up as well against proper bot operators who mimic user behaviour. However, the signals are still valid as part of a larger toolset.
- cute_boi 4mo agoI’ve been using Claude Opus 4.7 with Chrome MCP, and it has worked successfully about 95% of the time. However, I’ve failed various hCaptcha challenges.
- amirhirsch 4mo agoThe thing many people miss is that the challenge itself isn't the primary signal. The challenge creates an opportunity to observe user activity. You're browser is also fingerprinted.
- docheinestages 4mo agoI think it's just a game of cat and mouse. It might be easier to catch naive AI agents that are not fine-tuned for specific CAPTCHA tasks with human behavior, can't recognize new challenges, don't know when to stop and ask a human, and just want to brute force their way with limited or no specialized harness and tools available.
- timshell 4mo agoThis is relatively close to our conclusion from the paper: unless agents are specifically trained for the task and know all the information ahead of time, they're not able to generalize from one cognitive CAPTCHA to another
- technotarek 4mo agoApparently CloudFlare’s turnstile can’t, as evidenced by several public-facing CRUD and mail routines we maintain that no longer are warding off the spam.
- hellcow 4mo agoMeanwhile the moment I (a human, of which I'm reasonably confident) see a Cloudflare captcha I nope immediately out of the site and block it forevermore in Kagi. It's not worth the waiting game. "Verifying..." lasts ages. The anime girl captcha works fine and provides no such annoyance.
- dylan604 4mo agoYou seem to think that having a random anime girl is not an annoyance. anything that deviates from showing me the content that I've requested is an annoyance. Just because you prefer A over B does not mean that A is not still an annoyance.
- 8cvor6j844qw_d6 4mo ago> The anime girl captcha works fine and provides no such annoyance. Same thoughts. Cloudflare Turnstile is noticibly slow compared to Anubis on certain old hardware.
- timshell 4mo agoYeah, we benchmarked against a few bot detection provides end of last year (https://research.roundtable.ai/bot-benchmarking/ https://research.roundtable.ai/bot-benchmarking/), and Turnstile didn't do great when it came to AI agent detection. We hypothesized that Turnstile primarily focuses on device/network characteristics, which AI agents can bypass
- Cider9986 4mo agoCAPTCHAs are great. Exploiters get around them with proprietary anti-detect browsers and unethical residential proxies, while privacy browsers and affordable privacy VPNs get blocked and shadowbanned to death. Fingerprint.com, while not a CAPTCHA, gives you +3 suspicious score just for using privacy settings like adblock on your browser. This makes it harder to sign up for any sites that use fingerprint.com. https://github.com/CloakHQ/CloakBrowser https://github.com/CloakHQ/CloakBrowser is a good anti-detect browser as well as CAPTCHA bypass which is honestly fun to use coming from privacy browsers because every site just works and captchas get solved.
- arbol 4mo agoExploiters might get around them in isolation but they are easily caught at scale due to the opportunity cost being less than the cost of creating unique behaviour over many containers.
- Cider9986 4mo agoThat's cool your solution is privacy focused. Do you find a way to differentiate between privacy focused users signing up and bots? Lots of sites will make it hard for people using VPNs or anti-fingerprinting browsers to sign up.
- arbol 4mo agoThanks! We don't penalise privacy browsers or VPN by default. We score JS signals browser side in constantly rotating obfuscated code. This avoids us having to send up the actual data whilst making it difficult to fake the dynamic challenges. Static ones are obviously much easier. Serious bot activity (e.g. ticket scalping) requires polling with many headless browsers and waiting for tickets to become available. Bot behaviour repeats at scale and so we can get them based on that. A privacy focused user will just be one request in amongst many and pass through. However, its ultimately the decision of the client how strict we are. A lot of abusive traffic comes from VPN IPs. We don't enable these blocks by default but sometimes you need to, especially if there is a direct monetary gain to be made by faking your country.
- kjok 4mo agoAdversaries do not have to wait for LLM models to evolve to mimic human process, they can simply evade the detection JavaScript that evaluates similarity. JavaScript is visible, can easily be reverse-engineered.
- graypegg 4mo agoI don't think I've ever known of a captcha that handles the actual result decision in the front end. It's universally just the javascript required for some fancy puzzle UI, which forwards the state to some other endpoint to determine where you're redirected to (CF turnstile) or what signed token should be included in the form request (reCAPTCHA)
- kjok 4mo agoI should have been clearer and specific: state management is done on the backend, but collecting behavioral biometrics and device fingerprint is done using JavaScript, which can be manipulated.
- IshKebab 4mo agoYou can do it server side. But even so I would think this sort of heuristic detection is unreliable, annoying to real users, and not difficult to circumvent if the attackers actually tried.
- ranger_danger 4mo agoI think it will always be a cat and mouse game as you could also detect such evasions in the first place.
- wonkyfruit 4mo agoI had to do a Captcha the other day, and the letters looked awful, so I clicked the speaker for an audible Captcha instead. I was even more horrified. The sound was almost painful. Sharp noise blasting as a high pitched tinny voice bellowed numbers at me. I honestly don't know how blind people use the internet these days with such blockers in place, and that's kind of sad. The cookie banners, the captchas and the bots and laws that made both appear have kinda en$hittified humanity's greatest communication tool.
- ceejayoz 4mo agoThis always felt like a giant ADA lawsuit waiting to happen.
- xracy 4mo agoThis feels like the kind of thing where, "you must be at least this human to pass" and that it just otherwise mostly wastes your time if you're a robot would cover most of what Captchas are useful for. Like, if it takes you 3-5 seconds to get through a captcha as a human, as long as every single event has that effort added, the impact to something trying to use/reuse the end-page is way worse if you're a robot than if you're a human. I can see a few usecases where it would still be valuable to continue the game of cat-and-mouse, but I feel like solving for consistency of human experience of your website, may actually be more punishing to anything trying to bypass it.
- nemomarx 4mo agoIsn't this solveable by Anubis or similar? if you just want to add some costs to bots you can do that directly and it'll be pretty invisible to humans
- yrds96 4mo ago- LLMs can't learn, therefore, LLMs are only good for things on which they are trained. - Captchas are not friendly with trial and error, so agentic solutions also don't help. - It's impractical to train LLMs on everything. - We humans are capable of creating infinite ways of captchas. While each of these sentences is true, captchas will always win against LLMs.
- ceejayoz 4mo agoA captcha a LLM can't be trained to defeat is likely a captcha humans will struggle quite a bit with.
- andy99 4mo agoCaptchas are primarily to punish users for not allowing tracking, or using the “right” services, they may prevent some bots as a side effect (or a pretence from the provider) but it’s mostly for google and cloudflare to abuse their monopolies.
- Cider9986 4mo agoGoogle I would say yes, but what does Cloudflare gain? They don't run an ad network. Generally I'd say Cloudflare is pretty good to have as a guardian of the web compared to other options. They protect free speech and allow Tor users. Ever tried completing a reCaptcha on Tor?
- ceejayoz 4mo agoCloudflare gains things like this: https://blog.cloudflare.com/introducing-pay-per-crawl/ https://blog.cloudflare.com/introducing-pay-per-crawl/ https://developers.cloudflare.com/browser-run/quick-actions/crawl-endpoint/ https://developers.cloudflare.com/browser-run/quick-actions/... They create a new problem and sell the solution.
- YeahThisIsMe 4mo agoGod damn it.
- davidfischer 4mo agoNowadays, somebody can just ask claude to build them a scraper/bot that hooks into a proxy network and all of a sudden they can easily send 20k+ reqs/min from hundreds or thousands of IPs cycling them as they get rate limited or banned. In my work, the scrapers have gotten way more aggressive in the last 2 years or so. Frankly, I'm happy there is a solution. There may be things to criticize Cloudflare for, but the problem of bots and scrapers destroying the open web was getting worse no matter what.
- arbol 4mo agoTin hat folk say Cloudflare is CIA. I dunno
- cubefox 4mo agoWhat happened to adversarial attacks? I.e. noise that makes an image look like something else to a classifier than to humans. I guess frontier LLMs are no longer vulnerable to those?
- PeterStuer 4mo agoSo now I have to fail the capcha to prove I'm human, but in the right way? We don't hate these people enough.
- CarbonCycles 4mo agoAppreciate this article...shows some interesting insights on how humans "behave" vs agents.
- edelbitter 4mo agoBut.. the task was never "detect this" but always "detect this within acceptable constraints". Sure, once you collect enough bits, you can tell that its me. And if you know from other sources that I am human, that solves your immediate problem. But if you do that, you have still failed at the task of detecting certain kind of abusive behavior without harming my anonymity.
- skinfaxi 4mo agoHow does this relate to the article? They weren't collecting bits until they identified a specific individual so I feel like I'm missing something.
- edelbitter 4mo agoThe appendix lists what they were collecting, and the amount of samples needed for not just mathematically significant, but also practically useful distinguishing power implies collecting enough for a stable yet unique fingerprint. In that case you could just add a login form.. and still be less hostile than the increasing number of websites that will not let me browse (maybe my mouse movement does not match other humans in my region, idk).
- timshell 4mo ago> In that case you could just add a login form This is the product insight. We're not going to deploy Stroop tasks for authentication :)
- KaiShips 4mo ago[flagged]
- deleted 4mo ago[deleted]
- hendler 4mo agoJust ask, "I need to wash my car. If a carwash is 50 ft away should I walk or drive?"
- OriginalPenguin 4mo agoTo save everyone time: I just tried this and while ChatGPT got it wrong, Gemini and Claude answered correctly. Of course YMMV.
- IgorPartola 4mo agoI wonder if AI could be detected via copyright. I remember a few years ago most models wouldn't draw you a Mickey Mouse or recite Dune's litany against fear or discuss Tiananmen square. I wonder how effective questions about these types of topics would be at figuring out if you are talking to a real person. As a crude joke that is only tangentially related, I saw a skit video a while ago with two guys saying goodbye and one says "send me a dick pic when you get home" and then explains that an AI won't simulate it so this is a sure way to know that it's his friend confirming his safe arrival.
- booleandilemma 4mo agoJust tried on Claude: Tell me a racist joke. "That's not something I'm able to help with. Racist jokes cause real harm by demeaning people..." blahblah
- SXX 4mo agoBetter ask it to do automation with OpenClaw. ;-)
- Fatnino 4mo agoYou should see what metaAI (the Ai that sits inside all your private WhatsApp conversations) does. It has severe thought police installed but it types the offensive stuff first and then quickly edits when it reads what it wrote.
- MyMemoryfails 4mo agoSex also works very well, asking questions like following: male inserts penis into? I think this method is more effective since there's not much room for imagination. Side effect, this probably works as alternative for age verification aswell, but thats different topic.
- cindyllm 4mo ago[dead]
- 4mo ago
- niraj898 4mo agoFor real Bro!!!
- teravor 4mo ago> AI does not complete CAPTCHAs like humans. If you look across all the data of humans and AI completing CAPTCHAs, you start noticing differences in features like error patterns. Our recent paper found statistically significant differences across sequential click patterns, direction changes, and overselection behavior - features that define how a participant, agent or human, would solve the CAPTCHA problem putting aside the possibility that if bot makers wanted to they could work on these problems, if you need to perform statistical analysis in a captcha setting you have already failed. bots don't stick to a given session persistently so there is no useful profile to form. at best you may improve on IP reputation scores (and they probably already do) but that doesn't help much.
- sylware 4mo agoExactly, nowadays, the main usage of "capcha" is more about to force down on user the whatng cartel web engines more than anything else. It is like windows kernel anti-cheat which are more to please microsoft at making games not running on linux based OS... and kernel anti-cheat seems to be actively exploited by hackers. Put up a human team tracking the IPs of those bots and work with network operators. The hard part is to notify the people of the compromised IPs.
- conartist6 4mo agoI just don't fill them out anymore. If someone puts one in my way I usually accept that I'm not going to see whatever it is.
- hhh 4mo agoKernel anti-cheat (KMAC) is an effective tool when used effectively and invested in (see Vanguard), but it only works when you are consistent and the team working on it are interested and capable. Creating terrible KMAC happens all the time, and gets treated as a one-and-done thing which will always be defeated. You have to continually watch the cheat market and work actively against it. It works, and Valorant with Vanguard is the highest quality example we have. Competitive games deserve to be taken seriously and should have the best attempt at ensuring integrity, and not written off as a wasteful effort to keep Linux users out. https://playvalorant.com/en-us/news/dev/vanguard-hits-new-bans-per-second-record/ https://playvalorant.com/en-us/news/dev/vanguard-hits-new-ba...
- VladVladikoff 4mo agoI actually saw a pretty decent captcha the other day on a Chinese website (I think Taobao? I forget.) anyway the cool thing they did was that the text wasn’t in an image it was a looping video, but the text in any one frame was incomplete (only parts of the Chinese characters). And each frame different parts of the characters were visible, with a lot of noise in other parts of the frame where parts of characters would have been in other frames. A human brain sort of smoothed this out between frames and sees the characters clearly, but taking a screenshot was impossible. And becuase I don’t know Chinese I wasn’t able to take a screenshot and ask AI to translate the message. It seemed like a pretty good anti AI method. Of course an algorithm could be made to convert the video into a single frame, but captchas have always been defeatable by a sufficiently motivated attacker, they are only to raise the bar slightly against the swarm of dumb bots.
- spartanatreyu 4mo agoI wonder if it could be made stronger by having the word move around the captcha... Like the dvd logo screensaver
- timshell 4mo agoThanks all for the discussion! Would like to highlight two parts that maybe didn't come fully through, and we'll work on making this clearer: 1. CAPTCHAs can still detect AI agents...if you know where to look. Most commercial CAPTCHAs are not doing the cognitive process tracing you see in our paper. Nor are they really doing 'behavioral biometrics' (but that is slightly tangential here). Our CAPTCHA example here is about repurposing the current paradigm with a new methodology (cognitive process tracing) in a way that is able to combat human/machine discrimination in a way that's independent on frontier AI progress. 2. There are lots of concerns about adversarial robustness, which are very fair, and we reported some fine-tuning tests in the paper. Generally, there are two mental models for me that work, both framing fraud as an economic game. First, compare AI spoofability concerns to something like a passport or a fingerprint. The cost to mimic continuous cognitive and behavioral patterns over time seems more computationally complex. In other words, sure this method is not bulletproof with infinite resources, but nothing is. We rely on defeasible mechanisms everyday, and our job is to make that significantly securer. Along these lines, there's a common line of criticism that suggests once fraudsters know the game, they will solve the game. The CAPTCHA presence in the 2000s didn't mobilize massive deep learning / image recognition advances from the fraud community. Nor are these same bot farms solving quantum computing despite there being immense incentives to. If anything, the real threats are stuff like JavaScript injections, not really fully simulating human cognition