10 ms·
Cloudflare's new marketplace lets websites charge AI bots for scraping
- zebomon 2y agoHere's a look at my AI Audit on Bingeclock for anyone who's curious. Interesting drop in the last 48 hours given that it coincided with Cloudflare's announcement. https://www.bingeclock.com/blog/img/ai-audit-cloudflare-092324.png https://www.bingeclock.com/blog/img/ai-audit-cloudflare-0923... The payment program sounds intriguing, I suppose. I can't imagine it will do much to move the needle for websites that will become unviable due to traffic drain. Without a doubt, AI scrapers will (quite rationally from their POV) avoid anything but nominal payments until they're forced to do otherwise.
- AtNightWeCode 2y agoMaybe they could solve some of the core issues instead. It is like CF lost the source code and just pushing new more or less useless features all the time. Even though I think this is a fair change.
- dageshi 2y agoAhhh I love it. The era of silo's has well and truly arrived, I hope websites milk every dollar they can from the AI startups, they can afford it!
- kylehotchkiss 2y agoIs anybody else seeing an absolutely massive amount of Amazonbot crawls on their site? What are they up to? And why so aggressively?
- n_ary 2y agoMost likely aspiring AI startups gathering as much data as they can before regulation jaws snap shut around them cutting off the blood stream. In this AI race(hype), data is finally the ultimate gold. Also at the rate the information is polluted by GenAI junk all over, any remnants of real data is holy grail.
- kylehotchkiss 2y agoSo any unknown or upcoming AIs would just show as Amazon?
- nitwit005 2y agoThey have documentation on verifying if it is indeed their bot: https://developer.amazon.com/amazonbot https://developer.amazon.com/amazonbot
- osigurdson 2y agoNext step: generate reams of content using generative AI and get paid by Cloudflare when this is scanned by generative AI.
- micromacrofoot 2y agoabsent of legal changes this mostly rewards companies that figure out how to scrape without being detected, this problem has existed before AI
- rahimnathwani 2y agoWhile it’s a bold idea, Cloudflare is not sharing a fully fleshed-out idea of what its marketplace will look like.
- synack 2y agoAre they gonna let me block the scrapers that run on Cloudflare Workers?
- renewiltord 2y agoJust use some residential proxy network and slam your target. They can't detect you.
- brikym 2y agoCloudflare has probably noticed those proxy networks are quite expensive.
- renewiltord 2y agoSometimes get hit by the captcha but captcha solvers are cheap (0.3 cents a captcha).
- CaptainFever 2y agoRelated article: "An AI can beat CAPTCHA tests 100 per cent of the time" https://www.shiningscience.com/2024/09/an-ai-can-beat-captcha-tests-100-per.html https://www.shiningscience.com/2024/09/an-ai-can-beat-captch...
- j45 2y agoNeat licensing idea - look forward to seeing some case studies.
- datavirtue 2y agoWasn't the web designed to be scraped?
- marcus_holmes 2y ago> If you don’t compensate creators one way or another, then they stop creating, and that’s the bit which has to get solved I'm not sure this is true. Maybe they stop creating commercial stuff for sale, and go do something else for money, but generally creative people don't stop creating just because they can't get paid for it.
- CatWChainsaw 2y agoI guess Web3 will exist after all. In a microtransaction-per-webpage-utilized sense. No way websites don't start charging real people when there's money to be made.
- boristsr 2y agoI'm pretty interested in how companies are exploring how to properly monetize or compensate for scraped content to help keep a strong ecosystem of quality content. Id love to see more efforts like this.
- hedora 2y agoCompanies have been trying to find novel ways to bypass fair use / public domain laws for a long time. Each time they do, we see more consolidation of the media, and lower pay for the people that produce the content. I don’t see why this particular effort will turn out differently.
- bippihippi1 2y agoI wonder if there's a way to test this hypothesis. Does content being freely reproducible with minor modification increase the demand for content creators since new content is more valuable than the existing that can be copied. I'd guess that since AI can fair-useify a work faster than any human, that fair-use reviewers, compilers/collagers, re-imaginers, etc content creators will be devalued. However, AIs are as yet unable to create work as innovative as humans. Therefore new work should be more valuable since now there is demand from people and AIs for their work. I'm assuming that AI companies pay for the work that they use in some way. Hopefully the aggregation sites continue to compete for content creators.
- chrisweekly 2y ago> "I'm assuming that AI companies pay for the work that they use in some way." That mistaken assumption is at the heart of the problem under discussion.
- kordlessagain 2y agoThere's a HTTP code for charging for access: https://developer.mozilla.org/en-US/docs/Web/HTTP/Status/402 https://developer.mozilla.org/en-US/docs/Web/HTTP/Status/402 Then there's a Lightning Network protocol for it: https://docs.lightning.engineering/the-lightning-network/l402 https://docs.lightning.engineering/the-lightning-network/l40... With the Cloudflare stuff, it just seems like an excuse to sell Cloudflare services (and continue to force everyone to use it) as opposed to just figuring out a standard way of using what is already built to provide access for some type of micropayment.
- Mistletoe 2y agoHow will Scraping Chad deal with this? https://www.reddit.com/r/webscraping/comments/w1ve97/virgin_api_consumer_vs_chad_thirdparty_scraper/ https://www.reddit.com/r/webscraping/comments/w1ve97/virgin_...
- sunshadow 2y agoThere is no difference between this and a well known bot prevention mechanism, from the scraper perspective.
- neilv 2y ago> A demo of AI Audit shared with TechCrunch showed how website owners can use the tool to see how AI models are scraping their sites. Cloudflare’s tool is able to see where each scraper that visits your site comes from, and offers selective windows to see how many times scrapers from OpenAI, Meta, Amazon, and other AI model providers are visiting your site. And if I didn't authorize the freeloading copyright-laundering service companies to pound my server and take my content, then I need a really good lawyer, with big teeth and claws.
- yard2010 2y agoThis is such a nice opportunity for 4chan weirdos to teach the AI some new slurs.
- BSDobelix 2y agoI would say let's get rid of copyright and software patents altogether ;)
- blibble 2y agothey're already gone but only if you're well funded (OpenAI)
- CaptainFever 2y agoRemember that open source AI exists.
- mdaniel 2y agoI've always heard it as "the golden rule:" those who have the gold make the rules
- giancarlostoro 2y agoI really love Cloudflare. They're always up to something interesting and different. I hope we see more companies rise up similar to Cloudflare. I almost want to say Cloudflare is everything we hoped Google would be, but Google became another corporate cog machine that innovates and then scraps things up in one swoop. I don't recall the last I heard of Cloudflare spinning something up just to wind it back down? I don't think its impossible for them to make a bad choice, but I think they really think their projects through typically. My biggest problem with AI will be once it starts getting legislated, it will just be limited in how it can function / be built, we are going to lock in existing LLMs like ChatGPT in the lead and stop anyone from competing since they wont be able to train on the same data. My other biggest problem is "AI" or really LLMs which is what everyones hyped about, is lack of offline first capabilities.
- clvx 2y agoSomeone somewhere outside of your country's legal entities can still do all the things your country doesn't like and there's little to stop them. Governments might limit legal or commercial usage but it doesn't mean it won't exist.
- giancarlostoro 2y agoIts much harder to pull off when you're hitting an international market, are you really going to ignore an entire country? Maybe if it was a small country with few citizens, but if the EU or US passes a law, you're going to miss out on an entire market.
- nindalf 2y ago> last I heard of Cloudflare spinning something up just to wind it back down Cloudflare bet big on NFTs (https://blog.cloudflare.com/cloudflare-stream-now-supports-nfts/ https://blog.cloudflare.com/cloudflare-stream-now-supports-n...), Web3 (https://blog.cloudflare.com/get-started-web3/ https://blog.cloudflare.com/get-started-web3/), Proof of stake (https://blog.cloudflare.com/next-gen-web3-network/ https://blog.cloudflare.com/next-gen-web3-network/). In fact they "bet on blockchain" way back in 2017 (https://blog.cloudflare.com/betting-on-blockchain/ https://blog.cloudflare.com/betting-on-blockchain/) but it's telling that they haven't published anything in the last couple of years (since Nov 2022). Since then the only crypto related content on blog.cloudflare.com is real cryptography - like data encryption. I'm not criticising. I'm just saying they're part of an industry that thought web3 was the Next Big Thing between 2017-2022 and then pivoted when ChatGPT released in Nov 2022. Now AI is the Next Big Thing. I wouldn't be surprised if a lot of the blockchain stuff got sunset over the next few years. Can't run those in perpetuity, especially if there aren't any takers.
- FlyingSnake 2y agoMore details here at the Cloudflare blog: https://blog.cloudflare.com/cloudflare-ai-audit-control-ai-content-crawlers/ https://blog.cloudflare.com/cloudflare-ai-audit-control-ai-c...
- neilv 2y agoCloudflare found a new variation on their traditional service of protecting from abusers. This time, Cloudflare has formed a "marketplace" for the abuse from which they're protecting you, partnering with the abusers. And requiring you to use Cloudflare's service, or the abusers will just keep abusing you, without even a token payment. I'd need to ask the lawyer how close this is to technically being a protection racket, or other no-no.
- theamk 2y agodoesn't seem this way? > Website owners can block all web scrapers using AI Audit, or let certain web scrapers through if they have deals or find their scraping beneficial. You don't have to make any deals, or participate in the marketplace, "block all" is right there. And if you are not using Cloudflare, you are going to be abused. This is a sad fact, but I have no idea why you are blaming Cloudflare and not AI companies.
- flir 2y agoI dunno. If Cloudflare's protection doesn't work (and lets face it, it doesn't), why are you paying for it?
- TZubiri 2y agoAssociating a cost with a detrimental action is a well established defense against sybil attacks.
- mrits 2y ago[flagged]
- brigadier132 2y ago[flagged]
- neilv 2y ago> This kind of cynicism is boring. IMHO, this kind of thinking is only cynicism iff you're only looking for your angle to profit, and someone is peeing on your parade, every time they boorishly mention irrelevant, imaginary concerns like "ethics", "legality", or "Geneva Convention".
- Workaccount2 2y agoProps to cloudaflare for referring to it as "scanning your data", which is probably the most technically accurate way to describe what AI training bots are doing.
- creatonez 2y agoThis seems like a gimmick. Isn't preventing crawling a sisyphean task? The only real difference this will make is further entrenching big players who have already crawled a ton of data. And if this feature comes at the cost of false positives and overbearing captchas, it will start to affect users.
- hipadev23 2y agoCompanies have been trying and failing to prevent large scale crawling for 25 years. It’s a constant arms race and the scrapers always win. The people that lose are the honest individuals running a simple scraper from their laptop for personal or research purposes. Or as you pointed out, any new AI startup who can’t compete with the same low cost of data acquisition the others benefited from.
- jeroenhd 2y agoThe people that lose are the ones left with bandwidth charges and overloaded servers. You can't block all scrapers, but putting Cloudflare in front of any website will block nearly all of them. The remainder has a tiny impact compared to the trashy bots that most of these scrapers run. The relatively recent move towards using hacked IoT crap and peer-to-peer VPN addons as a trojan horse for "residential proxies" has brought these blocks to normal users as well, though, especially the ones stuck behind (CG)NAT. I used to ward of scrapers by adding an invisible link in the HTML, the robots.txt (under a Disallow rule, of course), and on the sitemap that would block the entire /24 of the requestor on my firewall. Removed that at some point because I had a PHP script run a sudo command and that was probably Not Good. Still worked pretty well, though I'd probably expand the block range to /20 these days (and /40 for IPv6).
- digging 2y ago> The people that lose ... are also everyone who makes (literally) any effort in the direction of digital privacy, whose internet experience is degraded and frustrating due to increasingly bad captchas or just outright refusal of service.
- andyp-kw 2y ago
- kelsey98765431 2y agolol good luck
- johnsutor 2y agoOr, you know, just create your own API for your platform and charge people per request to that.
- kijin 2y agoAI scrapers are parasites. I don't care whether you're OpenAI, Amazon, Meta, or some unknown startup. As soon as you generate a noticeable load on any of the servers I keep my eyes on, you'll get a blank 403 from all of the servers, permanently. I might allow a few select bots once there is clear evidence that they help bring revenue-generating visitors, like a major search engine does. Until then, if you want training data for your LLM, you're going to buy it with your own money, not my AWS bill.
- h8hawk 2y ago> AI scrapers are parasites. I've been making crawlers for a living! Thanks for informing me that I'm a parasite.
- kccqzy 2y agoThe AI scrapers are failing to discover something old-style search engines have been doing for decades: respecting a host and not giving them too much load. I'd say you did a good job banning those that generate noticeable load.
- billyhoffman 2y agoCommon Crawl is shown in their screen shot of "Providers" along side OpenAI and Antropic. The challenge is that Common Crawl is used for a lot of things that are not AI training. For example, it's a major source of content for the Wayback machine. In fact, that's the entire point of the Common Crawl project. Instead of dozens of companies writing and running their (poorly) designed crawlers and hitting everyone's site, Common Crawl runs once and exposes the data in industry standard formats like WARC for other consumers. Their crawler is quite well behaved (exponential backoff, obeys Crawl-Delay, will use SiteMaps.xml to know when to revisit, follows Robots.txt, etc.). There are significant knock-on effects if CloudFlare starts (literally) gatekeeping content. This feels like a step down the path to a world where the majority of websites use sophisticated security products that gatekeep access to those who pay and those who don't, and that applied whether they are bots or people.
- shadowgovt 2y ago> This feels like a step down the path to a world where the majority of websites use sophisticated security products that gatekeep access to those who pay and those who don't ... and that future has been a long time coming. People who want an alternative to advertising-supported online content? This is what that alternative looks like. Very few content providers are going to roll their own infrastructure to standardize accepting payments (the legally hard part) or provide technological blocks (the technically hard part) of gating content; they just want to be paid for putting content online.
- Terr_ 2y ago> People who want an alternative to advertising-supported online content? This is what that alternative looks like. Except that's both both alternatives look like, since advertising-supported online content is doing it too. Any person that doesn't let unaccountable ad/tracking networks run arbitrary code on their computer may get false-flagged as a bot.
- AlienRobot 2y agoI think this is a temporary problem. In a few years many AI companies will run out of VC money, others will be only after "low-background" content made before AI spam. Maybe one day nature will heal.
- xyzzy_plugh 2y agoAh yes, the ol' monopoly invents an illusionary marketplace ploy. Cloudflare is obviously right here. AI has changed things so an open web is no longer possible. /s What absolute garbage.
- zkid18 2y agoWhat's wrong with AI agents accessing website content? We seem to have been happy with Google doing that for ages in exchange for displaying the website in search results.
- spiderfarmer 2y agoAnd AI agents scrape your content in exchange for what exactly?
- zkid18 2y agoSorry, I distinguish here an AI agent that basically automate the visual lookup and scraping to feed into LLMs by big tech. I don't see any problem with the first one tbh.
- red_admiral 2y agoThe website owner chooses. They can say "nope" in robots.txt. Not everyone respects this, but Google does. Google can choose not to show that site as a result, if they want to. This adds a third option besides yes and no, which is "here's my price". Also, because cloudflare is involved, bots that just ignore a "nope" might find their lives a bit harder.
- lolinder 2y agoRobots.txt is for crawlers. It's explicitly not meant to say one-off requests from user agents can't access the site, because that would break the open web.
- Spivak 2y agoYep, there's really two parts to this. * Some company's crawler they're planning to use for AI training data. * User agents that make web requests on behalf of a person. Blocking the second one because the user's preferred browser is ChatGPT isn't really in keeping with the hacker spirit. The client shouldn't matter, I would hope that the web is made to be consumed by more than just Chrome.
- johnisgood 2y agoHow are they going to pay? How much? Can it be enforced?
- NoMoreNicksLeft 2y agoGreat. The HR software my company uses can charge me when my own bot "scrapes" my paystub pdf.
- flaburgan 2y agoI was recently speaking with people from OpenFoodFacts and OpenStreetMap, and I guess Wikipedia as the same issue. They are under constantly DDoS by bots which are scraping everything, even if the full dataset can be downloaded for free with a single HTTP request. They said this useless traffic was a huge cost for them. This is not about copyright, just about bots being stupid and people behind them not caring at all. We for sure need a solution to this. To maintain a system online nowadays means not only they get your data but you pay for that!
- epc 2y agoI’ve just taken to blocking entire swaths of cloud services IP networks. I don’t care what the intentions are, my personal sites don’t get the infinite bandwidth to put up with a thousands of poorly written spiders.
- MathMonkeyMan 2y agoI use a VPN when bittorrent is running, and I've found that several websites outright block me "for security reasons." They like to show me my IP address, too, like a great secret has been revealed and the SWAT team is on their way.
- neilv 2y agoIs there a public list of those address blocks, which you'd recommend?
- epc 2y agoNot that I know of, but each service seems to publish a list (some in text, some JSON). I’ll reply later with the URLs of the ones I have.
- epc 2y agoThis is what I have, see another reply for shared IP lists: https://ip-ranges.amazonaws.com/ip-ranges.json https://www.digitalocean.com/geo/google.csv https://www.gstatic.com/ipranges/cloud.json
- 015a 2y agoOne minor, tedious thing that I've become so tired of lately is showcased very plainly in the screenshot in this article: That the Cloudflare admin dashboard has now prominently placed "AI Audit (ALPHA)" as a top-level navigation menu item at the very top of the list of a Cloudflare Account's products. Everyone is doing this, for AI products or whatever came before them, and it genuinely pushes me away from paying for Cloudflare, as I get the distinct sense that they aren't building the things or fixing the problems that I feel are important to me. I would greatly appreciate the ability to customize the items and ordering of those items in this sidebar.
- siliconc0w 2y agoAny recommendations for simple WAF tool that will stop the majority of the abuse without having to use Cloudflare? I use Cloudflare just to keep that noise away from my logs but I'm not super keen to be dependent on them.
- meiraleal 2y agoWow, a big tech thinking about creators not about how to extract all they can but to give back. That became so uncommon nowadays. Cloudflare deserves their exponential growth. Kudos for them.
- sdflhasjd 2y agoHow long does the world-wide-web have left? It's always felt like it would be around forever, but at some point it will fade into obscurity like IRC has done. The golden age, I feel, has been gone a while, but "AI" seems like the beginning of the end.
- ivanjermakov 2y ago"AI" is the beginning of the end the same way as spam, malware and bot content were perceived in the past. To every action there is a reaction and "AI" won't be an exception.
- sharpshadow 2y agoIt is indeed a huge waste to scrape the same whole site for changes and new content. If Cloudflare is capable to maintain an overview about changes and updates it could save a lot of resources. The site could tell cloudflare directly what changed and cloudflare could tell the AI. The AI buys the changes and cloudflare pays the site keeps a margin.
- jsheard 2y agoThe sitemap.xml spec already has fields for indicating the last time a page was changed and how often it's expected to change in the future, so that search engines can optimize their updates accordingly, but AI scrapers tend to disregard that and just download the same unchanged page 10,000 times for the hell of it.
- Aachen 2y ago> sitemap.xml spec already has fields for indicating the last time a page was changed I did not know that bit! I'm considering adding this to my site now, because it sounds like it would save a lot of resources for everyone. Do (m)any crawlers use this information in your experience?
- jsheard 2y agohttps://developers.google.com/search/docs/crawling-indexing/sitemaps/build-sitemap#xml https://developers.google.com/search/docs/crawling-indexing/... Google ignores the priority and change-frequency fields, but they do use the last-modified field to skip pulling pages which haven't changed since their crawler last visited. Not sure exactly which signals Bing uses but they definitely use last-modified as well.
- delanyoyoko 2y agoI guess with marketplace like this, if webmasters are happy and the AI agents are also happy, then we'll be seeing quite a few services to come up with similar solution. Then end goal will be, from search engine optimization to something like LLM optimization or prompt engine optimization.
- dangoodmanUT 2y agothe blog makes it seem like the bot buys access but if they are only tracking the bot via the user agent then can't i piggyback on that user agent? no ai scraper is going to include an auth header when accessing your website...
- zackmorris 2y agoBoy I'm sick of clicking "Verify you are human" on everything from GitLab to banking apps running Cloudflare. Sick enough that I hope someone prominent at the EFF or similar takes Cloudflare to court over it. One company shouldn't be allowed to police access to the internet. And certainly shouldn't be allowed to start gatekeeping what is viewable by discriminating against the person or software doing the viewing. I worry that Cloudflare will keep escalating this unless they're sent a strong signal that it's not supported by the tech community. If you work there, it might be time to consider getting a different job. If you own stock, maybe divest. If you're connected, perhaps your associates can buy from competitors. That's probably the only way to get the board and CEO replaced these days.
- laserbeam 2y agoSomething I never considered, I wonder how clicking to be a human works for people with disabilities. There’s gotta be accessibility features there, and I bet bots are abusing them.
- gruez 2y agoAt least for cloudflare "captchas", you don't have no solve anything, only click a button. Therefore it's pretty accessible. My guess is that they care less about whether you're a human or not, and more about imposing resource costs on any attacker, because solving those challenges requires a full browser runtime (ie. hundreds of megs of memory + some non-trivial amount of CPU time). That's significantly more expensive than you spamming requests.post() with on a thousand threads.
- Wingman4l7 2y agoOr, the company leaves the accessibility alternative broken, and shrugs.
- Icathian 2y agoDo you also get mad at companies that make locks when people install them on their front doors?
- 2y ago