8 ms·
Cloudflare to introduce pay-per-crawl for AI bots
- yodon 1y agoCurrently in private beta
- jgrahamc 1y agohttps://techcrunch.com/2025/07/01/cloudflare-launches-a-marketplace-that-lets-websites-charge-ai-bots-for-scraping/ https://techcrunch.com/2025/07/01/cloudflare-launches-a-mark... "Several large publishers, including Conde Nast, TIME, The Associated Press, The Atlantic, ADWEEK, and Fortune, have signed on with Cloudflare to block AI crawlers by default in support of the company’s broader goal of a “permission-based approach to crawling.”"
- Toritori12 1y agoOverall I agree with the idea, but prob will be cheaper to bypass CF considering the amount of data that big techs are consuming (also Google with get it for free because Google Search?). If successful, I wonder how agents will transfer this cost to the user.
- figassis 1y agoMore publishers will start blocking google bots as well, bc google is already killing their revenue with AI results.
- jimbohn 1y ago>Google with get it for free because Google Search What if the second step is that Google pays the page it visits? By enabling a crawler fee per page, news websites could make some articles uncrawlable unless a huge fee is paid. Just thinking aloud, but I could easily see a protocol stating pricing by different kinds of "licensing" e.g. "internal usage", "redistribution" (what google news did/does?), "LLM training", etc. Cloudflare, acting as a central point for millions of websites, makes this possible.
- ethbr1 1y agoIt'd be a fitting solution if news closed the loop, crawled Google et al. to see if any of their content showed up there, then repriced future cotent higher for any search engines that reproduced content via genai.
- vbezhenar 1y agoThe question is: who has the leverage? If some small news website denies Google Bot crawling, it'll disappear from Google and essentially it'll disappear from the Internet. People do a great lengths to appease the Google Crawler. If some huge news website demands fees from Google, it might work, I guess. But I'm not sure that it would work even for BBC or CNN.
- jimbohn 1y agoI agree about the leverage and small website reasoning, definitely some game-theory related thinking is needed to get something like this right. But it does feel like this enables the "unionization" of websites against scraping giants, google is in an especially interesting position because, as you mentioned, could blackmail you into scraping in exchange for indexing.
- ipaddr 1y agoIf its a smaller news site they have already de-ranked them, and used their content for AI answers
- yantramanav 1y agoWhile this is a neat idea, how does it negate all the data theft being done by the bots so far? I recently saw a research paper behind a paywall but ChatGPT readily gave me a detailed summary of that article. I’m afraid the cat's out of the bag now.
- teruakohatu 1y agoAll the LLMs are being trained on LibGen/Anna’s Archive so it’s not in the least surprising they can tell you about papers behind paywalls. But I don’t think they are creating accounts to scrape paywalled data from the original source itself.
- suyash 1y agoNice to see someone addressing this annoying problem, I'm seeing first hand bot traffic go up as they are just gobbling up data. However instead of relying on Cloudflare, it would be better to have a open source protocol that handles permission and payment for crawlers/scraper.
- Leynos 1y agoIs this companies collecting data for model training, or is it agentic tools operating on behalf of users?
- xena 1y agoWhatever it is, I've seen the commons abuse Gitlab servers so hard they peg 64 high wattage server cores 24/7. Installing mitigations cut their power bill in half.
- Melonai 1y agoI think in the grand scheme of things barely anyone uses agents (as of now) to crawl sites quickly apart from maybe a quick Google search or two. At least that's been my observation of my non-technical field friends using LLMs. From what it looks like in the web logs it is in fact the same few AI company web crawlers constantly crawling and recrawling the same URLs over and over, presumably to get even the slightest advantage over each other, as they are definitely in the zero-sum mindset currently.
- rswail 1y agoThe protocol that Cloudflare are proposing could be implemented by anyone. There will need to be ways for crawlers to register and pay. CF is acting as the merchant of record, so they will be the ones billing, it's unclear what cut of the price they will take (if any) or if they will include it in their bundled services. This should be expanded to allow for: * micropayments and subscriptions * integration with the browser UI/UX * multiple currencies * implementation of multiple payment systems, including national instant settlement systems like UPI, NPP, FedNow etc.
- bgwalter 1y ago
- some_furry 1y agoI would totally do this. "Read my blog for free, or pay $25/page for your AI to read it for you." This is praxis. Enshittify the enshittification machine. We should also throw ads in there, via a deliberate prompt injection that the AI companies expose through an API. I totally won't misuse it ;)
- aspenmayer 1y agoI mean, I wouldn't care who referred the paying user. I would optimize my blog to serve free and paying users to the best of my ability. I would hope that I could do that but I don't know what you are trying to do with your blog. Most blogs could probably be self-hosted with some caching and static layout where possible to perhaps avoid needing to use Cloudflare. I guess you already have to be using CF to have access to these paying AI crawlers. Do you think that there is $25 of value in the creation of your blog, to say nothing of value that AI may be able to extract from it? (I'm speaking hypothetically, as I haven't looked at your profile to see if you link your blog, but I will do so now.) Edit: I have checked, and I've read your blog before. I think the answer to the question depends on who is asking but I don't know how you feel about the matter. I think asking for folks to pay for free things is a different value proposition than a pay-per-use fee, so the economics are different. You're also offering something different when you give away a blog and monetize access to a community or something similar, which is different still to accepting donations and so on. I don't know what you do for work or if you do your blog full time, but I think it's cool that you make it all the same.
- some_furry 1y agoI pay $300/year for the privilege of being able to write what I want without the pressure to monetize or surveil my readers. Some of my blog posts are linked by programming language docs for cryptography. Others have helped queer folks transition to a higher-paying tech career. It's difficult to quantify either of those things in a dollar value. I've opted to not ever do so. But if I can make the AI slop machines pay to view my furry/tech ramblings, I will do enthusiastically.
- 1y ago
- aspenmayer 1y agoRelated from TC: > Cloudflare launches a marketplace that lets websites charge AI bots for scraping https://techcrunch.com/2025/07/01/cloudflare-launches-a-marketplace-that-lets-websites-charge-ai-bots-for-scraping/ https://techcrunch.com/2025/07/01/cloudflare-launches-a-mark... https://archive.is/6UDUv https://archive.is/6UDUv
- crgwbr 1y agoAll this is going to do is drive AI companies to mask their user agent to appear as a standard browser, resulting in a worse end state than we’re in now. It’s an exercise in futility.
- some_furry 1y agoYes, but I do feel this makes "theft" arguments stronger if they're deliberately evading the paywall, if you decided to be litigious about it.
- delusional 1y agoThis isn't the kind of problem that really ought to be solved through courts. It's obvious to anyone that this is a new kind of problem that no author of the current jurisprudence envisioned. We need new legislation to stop this kind of abuse of the commons.
- sofixa 1y agoYes. It's always weird to me that people expect laws written over centuries, using precedents from even more centuries, to be able cover scenarios their authors couldn't have possibly imagined. Civil law countries seem better at keeping their laws up to date with new threats whereas a few common law ones (most notably the US) really insist on digging through what an 18th century slave owner would have thought about e.g. AI.
- soatok 1y agoI strongly agree with you, but I have no confidence in my country's current elected representatives to ever do anything good, so our hands are tied until we vote them out.
- rkrisztian2 1y agoI agree, it's only the big tech companies who do this AI crawling, and they will always have money for it. This paywall won't stop them.
- 1y ago
- bgwalter 1y ago"AI" is certainly creating problems that can be monetized. So humans get the Cloudflare captchas in order to access their own content at Stackoverflow and Mathoverflow, and "AI" crawlers get the data highway for a fee. And all of this does not stop the incumbents who have already stolen everything.
- jgrahamc 1y agoCloudflare has not used CAPTCHAs since 2023: https://blog.cloudflare.com/turnstile-ga/ https://blog.cloudflare.com/turnstile-ga/
- _rdvw 1y agoHaving to click "Verify you are a human" every single time is still awful.
- jlokier 1y agoEspecially when it asks you to do it again, and again, until you realise you're never going to be allowed to see the page. I've had that happen a few times with sites behind Cloudflare.
- mejutoco 1y agoTo save a visit: they use turnstile, a captcha replacement. The checkbox with verify you are a human. I would call that a captcha, but it is debatable if a non-puzzle check is.
- djfivyvusn 1y agoIf I say CloudFlare captcha and you know what I mean, does it really matter?
- deleted 1y ago[deleted]
- bgwalter 1y ago
- greatgib 1y agoIn theory, why not, in practice welcome to the world where neutrality of internet explode... Soon they could decide if your requests come from a specific company IP or networks, because you look suspicious... In addition, bot fighting was never supposed to be about blocking automatic users but about blocking abusers, like spammers and co. So now it means that bad actor can have a free pass if they pay (with stolen credit cards...) What I think would have been more fair is to propose rate limiting that would apply the same to everyone, so website should be reasonable in the limit they set for the normal users to not be annoyed. And then, you could pay to be able to have higher rate limit to ressources. That will compensate for the incurred cost to the infrastructure and the website owner. With that cloudflare could be in a good position to controle the rate limit, negotiate and collect payments to give it to the website owners.
- koolba 1y ago> In theory, why not, in practice welcome to the world where neutrality of internet explode... Anybody that has the sense^Wgall to clear their cookies regularly lives in this world as you get that CloudFlare gate keeping for just about every site you visit.
- nottorp 1y ago> Soon they could decide if your requests come from a specific company IP or networks, because you look suspicious... They already are. You probably can't browse half the internet without Cloudflare's approval.
- delusional 1y agoWe don't need another technical protocol. We need legislation.
- asim 1y agoThis is basically just how we want to do micro payments. I think coinbase recently introduced a library for the same using cryptocurrency and the 402 status code. In fact yea it's called x402. https://github.com/coinbase/x402 https://github.com/coinbase/x402
- imiric 1y agoThis should be the standard business model on the web, instead of the advertising middlemen that have corrupted all our media, and the adtech that exploits our data in perpetuity. All of which is also serving to spread propaganda, corrupt democratic processes, and cause the sociopolitical unrest we've seen in the last decade+. I hope that decades from now we can accept how insidious all of this is, and prosecute and regulate these companies just like we did with Big Tobacco. Brave's BAT is also a good attempt at fixing this, but x402 seems like a more generic solution. It's a shame that neither has any chance of gaining traction, partly because of the cryptocurrency stigma, and partly because of adtech's tight grip on the current web.
- hhh 1y agocrypto seems like a massive waste for what can just be a regular transaction
- trollbridge 1y agoSomething like BAT isn't that wasteful, and without crypto you'd be stuck never getting paid by bad actors in the scheme.
- gessha 1y agoBut why exactly does it have to be on an append-only ledger where transactions are processed/validated for a fee? Why can’t it be a more conventional transaction processor like VISA on top of the banking system?
- 1y ago
- baq 1y agoCloudflare toll as a service. brb, setting orders to buy $NET.
- a_c 1y agoUse the very fund to feed generated contents to LLM crawler is the only right move
- JimDabell 1y agoThis seems like it’s going about things in entirely the wrong way. What this does is say “okay, you still do all the work of crawling, you just pay more now”. There’s no attempt by Cloudflare to offer value for this extra cost. Crawling the web is not a competitive advantage for any of these AI companies, nor challenger search engines. It’s a cost and a massive distraction. They should collaborate on shared infrastructure. Instead of all the different companies hitting sites independently, there should be a single crawler they all contribute to. They set up their filters and everybody whose filters match a URL contributes proportionately. They set up their transformations (e.g. HTML to Markdown; text to embeddings), and everybody who shares a transformation contributes proportionately. This, in turn, would reduce the load on websites massively. Instead of everybody hitting the sites, just one crawler would. And instead of hoping that all the different crawlers obey robots.txt correctly, this can be enforced at a technical and contractual level. The clients just don’t get the blocked content delivered to them – and if they want to get it anyway, the cost of that is to implement and maintain their own crawler instead of using the shared resources of everybody else – something that is a lot more unattractive than just proxying through residential IPs. And if you want to add payments on, sure, I guess. But I don’t think that’s going to get many people paid at all. Who is going to set up automated payments for content that hasn’t been seen yet? You’ll just be paying for loads of junk pages generated automatically. There’s a solution here that makes it easier and cheaper to crawl for the AI companies and search engines, while reducing load on the websites and making blocking more effective. But instead, Cloudflare just went “nah, just pay up”. It’s pretty unimaginative and not the least bit compelling.
- PeterStuer 1y agoSo: 1. Encourage fencing off everything by default to maximize need for bypass 2. Offer bypass through payment 3. Profit! You wouldn't believe the number of public administrations with public information that have (mostly unwittingly) had some lazy contractor put Cloudflare in front of their entire site, blocking even their RSS feeds from M2M. Yes, you can send them mails and call and sometimes, if they even understand the problem, they will fix it after a few months just before the next cheapest contractor is hired and we start all over again. Not saying Cloudflare is just an extortion racket, but it's getting closer by the day.
- deleted 1y ago[deleted]
- 9283409232 1y agoI don't trust Cloudflare but this is not a problem they created. They solved a real problem with DDoS protection in the beginning and now AI crawlers increasing server cost is not a negligible problem. The CEO of iFixit called out Anthropic publicly for hitting their site a million times in 24 hours to scrape it. We are passed the point of good faith action from these AI companies. They are adversarial and need to be treated as such.
- PeterStuer 1y agoBut they do create this problem. They could have specifically defaulted to not blanketing RSS feeds and other M2M specific pages. Instead, we are now in a situation where even daring to look at robots.txt can flag you as a bot.
- johnklos 1y agoTrue. For all those people who want to make excuses for Cloudflare, it's an excellent reminder that they've known about this problem for years and they still haven't fixed it. Are they inept? Or do they really only care about things that bring them profit and that normalize their marginalization of non-paying groups? Which explanation makes the most sense?
- OtherShrezzing 1y agoSomeone should use this to create a new browser. A human user drops $100 into the browser, and each website offers a per-page-view rate, gradually deducted from the $100. In exchange the user doesn't have to suffer through advertisements.
- mdrzn 1y agoAnd we are back to: how much would you pay for a website before you access said website for the first time?
- sofixa 1y agoThe Web Monetisation protocol solved this by doing a prorata based on how long you spent on the page.
- mdrzn 1y agoWhat if I save the page locally on my machine? Or archive it?
- nottorp 1y agoWhat if my phone rings and I forget to close the page? Or leave it open to read it later? Plus as one of the parent comments said, I am not paying before I get an idea what I'm paying for.
- sofixa 1y ago> Plus as one of the parent comments said, I am not paying before I get an idea what I'm paying for. The logic was, you aren't paying to a website. You're paying to a broker that distributes how much you've paid to all websites you've visited that have opted in, prorated based on how much time you spent on them.
- nottorp 1y ago
- deleted 1y ago[deleted]
- FloatArtifact 1y agoWhat about if somebody uses artificial intelligence crawler to help them navigate the web as an accessibility tool? Enabling UI automation. It already throws up a lot of... uh... troublesome verifications.
- samrus 1y agoThe site owner can allow such crawlers. There is the issue of bad actors pretending to be these types of crawlers but that could already happen to a site that want to allow google search crawlers but not gemini training data crawlers for example, so theres strong support to solve that problem
- throw10920 1y agoWe already have ARIA, which is far more deterministic and should already be present on all major sites. AI should not be used, or necessary, as an accessibility tool.
- freeone3000 1y agoIf site authors would actually use aria. Not everything is a div, italic text is not for spawning emoji… it’s not good for semantic content or aria right now. It should not be necessary, but it is.
- ziml77 1y agoThere's plenty of people who don't bother with ARIA and likely never will, so it's good to have tools that can attempt to help the user understand what's on screen. Though the scraping restrictions wouldn't be a problem in this scenario because the user's browser can be the one to pull down the page and then provide it to the AI for analysis.
- kentonv 1y agoHow would an individual user use a "crawler" to navigate the web exactly? A browser that uses AI is not automatically a "crawler"... a "crawler" is something that mass harvests entire web sites to store for later processing...
- numberless 1y agoThis pretty much solves the problem of too many bots, but only in a way that works with Cloudflare and does not help the rest of the web. They don't mention any possibility of specifying a different platform to route payments through for instance.
- ethbr1 1y agoCloudflare published details of the prototype implementation, so there's no reason if it takes off then other CDNs and hosts can't implement the same 402 protocol. There's literally nothing Cloudflare-specific about this.
- pu_pe 1y ago> The true potential of pay per crawl may emerge in an agentic world. What if an agentic paywall could operate entirely programmatically? Imagine asking your favorite deep research program to help you synthesize the latest cancer research or a legal brief, or just help you find the best restaurant in Soho — and then giving that agent a budget to spend to acquire the best and most relevant content. So the vision is a paywall around the whole internet. Content aggregators would charge AI companies to provide data relevant to specific queries. Sounds like a nightmare to me.
- deleted 1y ago[deleted]
- cedws 1y agoCan someone explain the payment headers part? Why not just have a header called X-Crawl-Key or something and intercept that header to figure out who to charge for the request?
- krab 1y agoThey have some headers for authentication. The payment part is for the price negotiation. The headers tell you that Cloudflare wants to charge you for this particular content and you tell CF that you're OK with being charhed up to $AMOUNT.
- mhandley 1y agoThat sounds reasonable for access to actual content, but it produces a huge new incentive to constantly produce vast amounts of AI-generated slop served via Cloudflare. Is there a way to disincentivize this?
- yen223 1y agoI presume the onus will now be on the AI scrapers to decide whether that AI-slop site is worth paying for. How they will figure this out will be interesting to see.
- samrus 1y agoThats a more general problem. As content gets cheaper to produce with AI, how do consumers discriminate between good content and slop. We already have this problem with youtube and twitter and reddit Its interesting that the AI companies will now be on the other end of this issue
- asimpletune 1y agoIt’s a step in the right direction but I think there’s a long ways to go. Even better would be pay-for-usage. So if you want to crawl a site for research, then it should be practically free, for example. If you want to crawl a site to train a bot that will be sold then it should cost a lot. I am truly sorry to even be thinking along these lines, but the alternative mindset has been made practically illegal in the modern world. I would 100% be fine with there being a world library that strives to provide access to any and all information for free, while also aiming to find a fair way to compensate ip owners… technology has removed most of the technical limitations to making this a reality AND I think the net benefit to humanity would be vastly superior to the cartel approach we see today. For now though that door is closed so instead pay me.
- danaris 1y agoThe problem with this is that people who want to make money will always be highly motivated to either find loopholes to abuse the system, outright lie about their intentions, buy and resell the data for less (making profit on volume), or just break in. "Ah, it's free for research? Well, that's what I'm doing! I'm conducting research! Ignore the fact that once I have the data, I'm going to turn around and give it to this company that is coincidentally also owned by me to sell it!"
- joosters 1y agoYou can tell the difference between the two by checking if the Evil bit is set in the corresponding IP packet - RFC 3514 already standardised this.
- Intralexical 1y agoIf that doesn't work, you can also add rate limiting by enforcing compliance with RFC 1149.
- stego-tech 1y agoLiterally this. It’s why I advocate for regulations over technological solutions nowadays. We have all the technology we need to solve today’s ills (or support the R&D needed to solve today’s ills). The problem is that this technology isn’t being used to make life better, just more extractive of resources from those without towards those who have too much. The solution to that isn’t more technology (France already PoC’ed the Guillotine, after all), but more regulations that eliminate loopholes and punish bad actors while preserving the interests of the general public/commons. Bad actors can’t be innovated away with new technological innovations; the only response to them has always been rules and punishments.
- skenderbeu 1y agoHow long before we get pay per browse and the internet is 6ft under?
- freeone3000 1y agoHonestly preferable to the insane amounts of paywalls and advertising
- squigz 1y agoThis is a paywall.
- BenjiWiebe 1y agoI'd rather pay 5c for one article than subscribe for $10/yr to view one article. Still a paywall, but less annoying.
- nerdix 1y agoThat won't end ads. Just like paid cable subscriptions didn't end TV ads. Or how ads are slowly creeping into the various streaming platforms with "ad supported tiers".
- nosioptar 1y agoA week. I'm constantly getting cloudflare nonsense that thinks I'm a bot. (Boring firefox + ublock setup.) I wouldn't be surprised if I start seeing a screen trying to get me to pay. If so, I'll do what I currently do when asked to do a recaptcha, I fuck off and take my business elsewhere.
- Tijdreiziger 1y agoAre you behind a CGNAT?
- nosioptar 1y agoNot that I know of.
- nottorp 1y agoSo we used to have this company that did good things for the internet... like usable search... Now we have this company that does good things for the internet... like ddos protection, cdns, and now protecting us from "AI"... How long will the second one last before it also becomes universally hated?
- wewxjfq 1y agoGood things for the Internet? I stop visiting sites that nag me with their verification friction. They are the only reason I replaced Stack Exchange with LLMs.
- 9283409232 1y agoCloudflare isn't universally hated but I think most people are very nervous about the power Cloudflare holds. Bluesky puts it best "the company is tomorrow's adversary" and Cloudflare is turning into a powerful adversary.
- lofaszvanitt 1y agoYeah, time to kneecap them before it's too late. EU is sleeping...
- nosioptar 1y agoMost people I know in real life already hate cloudflare.
- 1dom 1y agoI really like the idea that crawlers who are profiting should have to pay content owners/creators per crawl. On principal though, I think Cloudflare doing this is just one more thing to create the perception that you can't put something on the internet unless it's through Cloudflare. This harms a transparent and decentralised web and makes selfhosting seem even less appealing to those who don't know any better. This should be implemented as a web protocol with crypto though so anyone can charge bots without having to be Cloudflare fronted. Not really a fanboi of 99% of crypto stuff, but IMO, a purely technical, open and decentralised solution to this sort of problem was the crypto dream. We can all guess the people who will make the most money off this, and one of them is Cloudflare. A bunch of the other winners probably also run some of the more aggressive crawlers.
- vbezhenar 1y ago> This should be implemented as a web protocol with crypto though so anyone can charge bots without having to be Cloudflare fronted. Not really a fanboi of 99% of crypto stuff, but IMO, a purely technical, open and decentralised solution to this sort of problem was the crypto dream. It's not just about payment. It's about refusing to serve content to bots, unless they paid. It might be hard to implement without Cloudflare, when bot developers specifically target your website. The whole point of Cloudflare is to let them decide whether it's bot or user that hits your website. It is complicated task. Unless you want to force all users to pay, both humans and bots.
- imglorp 1y agoYes, right, it should be an open protocol so any CDN or content provider can use it the same way. Hopefully it becomes a part of popular web servers so little guys can play along without a CDN. It needn't be crypto, but would be convenient. Lacking that, it would need some unforgeable presentation of identity that could be connected to a bank account. I shed not one tear for the crawlers - they had their chance to respect robots.txt on the honor system. Now we force them.
- 1dom 1y agoI can't work out how an open protocol implementation of this could work without crypto: ultimately if it's just fiat, a business entity needs to be the payment processor who aggregates microtransactions and pays them out to content owners, this is the role cloudflare is playing. The problem is microtransactions are not feasible in fiat, and to remove the aggregator role like Cloudflare means a huge amount of microtransactions from each potential crawler to each content owner. That's just too much expensive work compared to the current position. I agree though, I shed no tears for crawlers, but hopefully we're beyond the naivety of honour systems - again, the thing crypto was supposed to be solving. Forcing big evil crawler entities to bend the knee by hiding behind big evil CDN entities feels silly though.
- nialse 1y agoAlthough Cloudflare CEO Matthew Prince pre-launched their new offering with a compelling speech and numbers to boot, the mechanics does not add up. There is an assumption that AI companies need to scrape the web for content. This is certainly true for new AI companies and new content, but the vast majority of scraping useful content has already been completed. In addition new content will tend to be AI generated itself, which might not help training, and in the US training on purchased content has been deemed fair use recently. What problem is being solved? The perceived issues are twofold, increasing crawling by AI scraping bots is causing traffic and thus an additional cost, and content creators lack compensation for their work in terms of money or notoriety (according to Matt). Cloudflare obviously have traditionally focused on the first, and needing to grow they see the potential in being a middle man in the second. Where does this get us? Will Cloudflares service lower traffic volumes not generating revenue? Absolutely. Use of the service will be perceived as a success based on this metric and the revenue generating traffic will stay on similar and higher levels initially. Then, if the content indexed becomes more and more stale, as AI companies may or may not be willing to pay the associated costs, revenues will slide long term. Content creators seeking fame or fortune may then seek other avenues to promote and distribute their content as they perceive the alternatives as better. The sole hope for Cloudflare is that a couple of the large AI outlets "play ball", and make the payed for indexed content available based on subscription fees or, god forbid, ads. However, then they might would want their users to be able to access the full contents guarded by other paywalls, and not only previews offered. One would hope that this would lead to a future where creative humans are compensated more for their cognitive work. Unfortunately, with the trajectory we're on, that is a select few as the marginal cost of content is quickly approaching zero. https://x.com/carlhendy/status/1938465616442306871 https://x.com/carlhendy/status/1938465616442306871
- phillipcarter 1y ago> This is certainly true for new AI companies and new content, but the vast majority of scraping useful content has already been completed. For training a base model, yes, but there's a big category of AI use case: search engine. Those invocations of the model involve web searches, often during reasoning steps, and they will absolutely scrape for content.
- saddlerustle 1y agoThis ends up being pretty bad for competition because it does not block the largest AI scraper of them all: Googlebot.
- ukd1 1y agomeh. ads for ai content is really the answer - e.g "OpenAI ads" content creator puts a tag on their page / set their domain - when the crawler sees it, display an ad pass on $ as usual.
- adjfasn47573 1y agoomg what are you, a sadist?
- mattlondon 1y agoThis is where Google wins AI again - most people want the google-bot to crawl their site so they get traffic. There is benefit to both sides there, and Google will use it's crawl-index for AI training. Monopolistic? Perhaps. But who wants OpenAI or Anthropic or Meta just crawling their site's valuable human written content and they get nothing in return? Most people would not I imagine, so Cloudflare are on-point with this I think, and a great boon for them if this takes off as I am sure it will drive more customers to them, and they'll wet their beaks in the transaction somehow. Bravo Cloudflare.
- mysteria 1y agoEven before AI was a thing some websites would deny all crawlers in robots.txt except for the Googlebot for the same reason.
- Scaevolus 1y agoGoogle's "AI Overview" is massively reducing click-through rates too. At least there's a search intent unlike ChatGPT? > It used to be that for every 2 pages G scraped, you would expect 1 visitor. 6 months ago that deteriorated to 6 pages scraped to get 1 visitor. > Today the traffic ratio is: for every 18 pages Google scrapes, you get 1 visitor. What changed? AI Overviews > And that's STILL the good news. What's the ratio for OpenAI? 6 months ago it was 250:1. Today it's 1,500:1. What's changed? People trust the AI more, so they're not reading original content. https://twitter.com/ethanhays/status/1938651733976310151 https://twitter.com/ethanhays/status/1938651733976310151
- Workaccount2 1y agoPerhaps many people here live in tech bubbles, or only really interact with other tech folks, online, in person, whatever. People in tech are relatively grounded about LLMs. Relatively being key here. On the ground in normal people society, I have seen that people just treat AI as the new fountain of answers and aren't even aware of LLM's tendency to just confidently state whatever it conjures up. In my non-tech day to day life, I have yet to see someone not immediately reference AI overview when searching something. It gets a lot of hostility in tech circles, but in real life? People seem to love it.
- vasilzhigilei 1y agoMan, HN is sleeping on this right now. This is huge. 20% of the web is behind Cloudflare. What if this was extended to all customers, even the millions of free ones? Would be really amazing to get paid to use Cloudflare as a blog owner, for example
- DocTomoe 1y agoThe cynic in me says we'll be seeing articles about blog owners getting fractions of a tenth of a penny while Cloudflare pockets most of the revenue. And of course it will eventually be rolled out for everyone, meaning there will be a Cloudflare-Net (where you only can read if you give Cloudflare your credit card number), and then successively more competing infrastructure services (Akamai, AWS, ... meaning we get into a fractured marketplace kind of situation, similar to how you need dozens of streaming abos to watch "everything"). For AI, it will make crawling more expensive for the large guys and lead to higher costs for AI users - which means all of us - while at the same time making it harder for smaller companies to start something new, innovative. And it will make information less available on AI models. Finally, there’s a parallel here to the net neutrality debate: once access becomes conditional on payment or corporate gatekeeping, the original openness of the web erodes. This is not the good news for netizens it sounds like.
- vasilzhigilei 1y agoI worked at Cloudflare for 3 years until very recently, and it's simply not the culture to behave in the way that you are describing. There exists a strong sense of doing the thing that is healthiest for the Internet over what is the most profit-extractive, even when the cost may be high to do so or incentives great to choose otherwise. This is true for work I've been involved with as well as seeing the decisions made by other teams.
- seanw444 1y agoUnfortunately, even if it is as you describe, human nature is such that it will not stay that way forever. Likely not even for long.
- johnnyApplePRNG 1y agoThey're going to need to deal with camoufox being easily able to circumvent their bot detection before they start charging for this, imho.
- vpShane 1y agoHaven't heard of / seen camoufox; makes sense to detect browsers by running a modified firefox that renders things like WebGL/Canvas and makes it incredibly hard to stop them. This coupled with advanced proxy lists, which is what AI companies, or rather just random scrapers will go for if they're blocked and charged. They want the data for free.
- krunck 1y agoFirst a paywall for AI. Then a paywall for people - which is no different from an internet user license as the payment methods allowed would not be anonymous.
- aussieguy1234 1y agoIf cloudflare popularizes this type of pay per crawl setup, I'd expect to see an open source standard to be created for these types of internet payments.
- udev4096 1y agoClownflare strikes yet again, bloating the web one at a time!
- hubraumhugo 1y agoWe all agree that AI crawlers are a big issue as they don't respect any established best practices, but we rarely talk about the path forward. Scraping has been around for as long as the internet, and it was mostly fine. There are many very legitimate use cases for browser automation and data extraction (I work in this space). So how will Cloudflare detect bots and get them to pay? And how many humans and legitimate bots will get blocked as a side effect? We're somehow still stuck with CAPTCHAs, a 25 years old concept that wastes millions of human hours and billions in infra costs [0]. How can we enable beneficial automation while protecting against abusive AI crawlers? [0] https://arxiv.org/abs/2311.10911 https://arxiv.org/abs/2311.10911
- deleted 1y ago[deleted]
- RVuRnvbM2e 1y agoBy Cloudflare inserting themselves as a "HTTP market maker", is this step one in the enshittification of the web?
- Anamon 1y agoFirst step? I think we're a few thousand steps along that way already, and rapidly approaching our final destination.
- ryao 1y agoWhat happens when Google starts using the data scrapped by its search engine crawlers for AI training? What prevents another crawler from impersonating one that is part of this program and getting someone else to pay? What happens when people start using headless browsers as crawlers and they are undetectable?
- tzury 1y agoSeems like Google got a pass, since their search engine activity is not included. Once they cached the page, their AI can "crawl" internally. (Same applies to Bing most likely)
- dabbz 1y agoThe more I read this, the more I feel like web attestation is going to be suggested as a way to prove AI bots vs humans. They've been trying to push this through for a while now (to some moderate success). This may be the final push they are looking for to get it more thoroughly integrated in the web as a whole. I know it wasn't mentioned anywhere here, but it's the silent part that fits this puzzle piece really well.
- yonran 1y agoWhy do AI agents need to scrape so often, vs. aggressively caching or using archive.org or their own crawls of the internet?
- rralian 1y agoMy gut reactions… - I agree that something like this is necessary or the whole model of the internet will be broken, like Matthew Prince [explained in this video](https://www.youtube.com/watch?v=H5C9EL3C82Y https://www.youtube.com/watch?v=H5C9EL3C82Y). - Their approach seems very imperfect, but I understand that you have to start somewhere. - They are paying per crawl… but in fairness it should really be per usage. It’s like paying music artists once when they upload to Spotify rather than per-play -- even though one artist gets zero plays and another gets ten million. Sure, the idea is crawlers will bid more for the popular content author, but what if a nobody author has a one-hit-wonder piece of content. They’ll still just get a couple bips per crawl and then the cat is out of the bag. - One solution to this would be requiring a GDPR-style forget mechanism, where the author is granting a limited-duration license for the content (say… one week), after which it must be deleted and re-licensed. This would be a huge fix for the whole thing… and the more I think about it the more I think it’s essential for this to work. - The auction mechanics are biased to the crawler… if there is a spread between artist price and crawler max price, then the crawler pays the lower price set by the artist. It should be the average. - They will need to provide content authors with analytics about the pricing mechanics for the bids the crawlers are making. - If this whole thing works, then products that optimize bid mechanics on behalf of authors will be a big growth industry. - If Cloudflare are setting themselves up as the clearing mechanism for payments, that’s far too much power and profit for one company. It’s even worse than the Google monopoly. Somehow the payment mechanics need to be democratized.
- adjfasn47573 1y agoI see most people stating that the internet as we know it could be gone because of AI. I’m asking you: Why not? The internet is not even a typical human lifespan old. It’s crazy young on a large scale. Why would anyone assume that it will (and has to) stay the way it is today? There are so many downsides of the current web. Slob everywhere (even long before AI) because of all sorts of people trying to exploit it for money. I welcome a change. An internet with less ads, more genuine information. If AI will lead to this next phase of the internet, so be it. And this phase won’t be the last either.
- isodev 1y ago> all sorts of people trying to exploit it for money Because they could. In AI-first web, people can't really do anything about anything - only those in control of training the handful of "big popular AI models" are the gatekeepers of all knowledge. > with less ads, more genuine information That's orthogonal to AI. Models are already being trained to favour certain products/services and they already (re)produce factually incorrect information with no way to verify or correct them.
- NitpickLawyer 1y ago> only those in control of training the handful of "big popular AI models" are the gatekeepers of all knowledge. I think that's certainly the case now, and it will be for a while, but slowly we're getting closer to that "AI personal assistant" sci-fi inspired future, where everything runs on "your" infra and gathers data / answers questions locally. You'd still need "raw" data access for that. A way to micro-pay for that would certainly help, imo.
- c4wrd 1y agoYou're missing the bigger picture. It isn't free to put content on the Internet. At a bare minimum, you have infrastructure and bandwidth costs. In many cases, a goal someone may have is that if they publish content on the internet, they will attract people to return for more of the content they produce. Google acted as a broker, helping facilitate interactions between producers and consumers. Consumers would supply a query they want an answer to, and a producer would provide an answer or facilitate a space for the answers to be found (in the recent era, replace answer with product or store-front). There was a mostly healthy interaction between the producers and consumers (I won't die on this hill; I understand the challenges of SEO optimization and an advertisement-laden internet). With AI, Google is taking on the roles of both broker and provider. It aims to collect everyone's data and use it as its own authoritative answer without any attribution to the source (or traffic back to the original source at all!). In this new model, I am not incentivized to produce content on the internet, I am incentivized to simply sell my data to Google (or other centralized AI company) and that's it. A clearer picture to help you understand what's going on: the internet of the past few decades was a bazaar marketplace. Every corner featured different shops with distinct artistic styles, showcasing a great deal of diversity. It was teeming with life. If you managed your storefront well, people would come back and you could grow. In this new era, we are moving to a centralized, top-down enterprise. Diversity of content and so many other important attributes (ethos, innovation, aestheticism) go out of the window.
- orliesaurus 1y agoThere’s a tiny but important difference between scraping and getting work done. Scraping is associate with mindless extraction. Like a vacuum cleaner sucking-in any data without context, permission, or contributing value back. On the other hand AI agents aren’t here to scrape for the sake of it - I have seen it first hand. They are here to get work done, mostly researching, summarizing, assisting, building new products. You could argue this data is then use to train further a model, you would probably argue correctly, but that’s a topic for another day. I implemented the poor man's version demo of what a similar concept to this could be like: http://github.com/toolhouseai/fastlane-demo http://github.com/toolhouseai/fastlane-demo
- kumarski 1y agoFunny finding you here. ;) Hey Orlie.
- orliesaurus 1y agoHey datarade
- Zenul_Abidin 1y agoThis is cool but I don't like how this forces all crawlers to use Cloudflare. Google Chrome developers were proposing some Web Monetization API in Chromium a few years back, back when the Manifest V3 drama was still fresh, so maybe we should look into that instead. To allow decentralized payments to not be dependent on a single vendor.
- johnsbrayton 1y agoI distrust Cloudflare so much. I have been trying to get my RSS reader on their Verified Bots list for years, but their application form appears to go nowhere.
- blancotech 1y ago> An important mechanism here is that even if a crawler doesn’t have a billing relationship with Cloudflare, and thus couldn’t be charged for access, a publisher can still choose to ‘charge’ them. This is the functional equivalent of a network level block (an HTTP 403 Forbidden response where no content is returned) — but with the added benefit of telling the crawler there could be a relationship in the future. IMO this is why this will not work. If you're too small a publisher, you don't want to lose potential click-through traffic. If you're a big publisher, you negotiate with the main bots that crawl a site (Perplexity, ChatGPT, Anthropic, Google, Grok). The only way I can see something like this work is if a large "bot" providers set the standard and say they'll pay if this is set up (unlikely) or smaller apps that crawl see that this as cheaper than a proxy. But in the end, most of the traffic comes from a few large players.
- deleted 1y ago[deleted]
- vpShane 1y agoVery neat idea to charge the crawlers via headers. Wondering if it blocks things such as Infatica as well, where AI scrapers use that service to scrape data behind residential proxies, in which it's impossible to get the actual residential proxy for. App developers hide Infatica drone code in things; I haven't thought of a reasonable way to block it. IP addresses use the systems, humans use them in certain ways, Cloudflare must have a decent way to block these.
- lofaszvanitt 1y agoOh dear god, the gatekeepers are turning evil. Just watch as they devour the internet. Cloudflare is slowly losing its mask, and soon we’ll be staring into the jaws of the deadly spider. We protec you from evil depths of the inernet, like ddos. We protecc you from evil AI. Now we handle payments for you, of course you have to use our network. ... ... ...
- billy99k 1y agoWhat happened to the pro-piracy groups? Aren't we suppose to push for a more free Internet?
- xacky 1y agoI think Archive.org should have a "donate to scrape" policy in place as well.
- kampat 1y agoIs there a way to recognize the AI agents who mimic the http headers similar to a browser?
- KETpXDDzR 1y agoCan't (AI) crawlers just use a browser user agent? TL;DR: Yes, but with some more work that the big AI companies can surely manage to implement. To bypass detection, a crawler would need to mimic not just the user agent, but also the full set of browser behaviors, headers, and environmental fingerprints—a much more complex task that Cloudflare’s systems are designed to counter.
- gck1 1y agoAll that this will result in is Cloudflare starting to demand payment from humans to browse the web, eventually. It already became impossible to browse the Cloudflare-dominated web if you use VPN and/or any browser that resists fingerprinting. Companies with a lot of money will pay up, smaller companies that depend on crawling for legitimate purposes will struggle, Cloudflare will pocket most of the profits, website owners will get pennies on the dollar and end user experience will get another blow.
- cyounkins 1y ago> Cloudflare, along with a majority of the world's leading publishers and AI companies, is changing the default to block AI crawlers unless they pay creators for their content. It really seems that they did _not_ change the default, since this feature is in private beta.
- andreioros 1y agowhile pay per crawl might work for big publishers, I think a different model might be more interesting for small and medium content creators – pay per prompt response reference/citation https://andreioros.com/blog/ai-llm-search/ https://andreioros.com/blog/ai-llm-search/