14 ms·
A year of fighting scrapers on my 1.5 million-page website
- qiyi41789 1mo ago[dead]
- lazerg 1mo ago[flagged]
- Yiin 1mo agowhat gave you this idea?
- pixl97 1mo agoBecause the ubermensch on HN can detect AI text from 10 miles away with 100% accuracy... like that time they called documents from 2015 AI written.
- conartist6 1mo agoI'm just beyond thrilled that it's actually starting to be a public embarrassment to cough up a glob of AI text.
- deleted 1mo ago[deleted]
- chrisandchris 1mo agoTake a look at the dashes. It's all there. (i skimmed the whole post, there are none)
- rhdunn 1mo agoI knew it... Wuthering Heights was written by AI and Lucy Maud Montgomery, Edgar Allan Poe, et. al. were AI bots churning out content! Or maybe -- just maybe -- using dashes isn't a sign of content being AI written, just a style it picked up from the training data.
- splatter9859 1mo agoAs odd as this sounds, and as odd as the world gets, I still find a small comfort in the fact that there is enough stability in this world that I can count on one thing: someone always accusing a blog post on HN as being AI generated. Never fails. The OCD part of me can now go about my day.
- askl 1mo agoI didn't read the article, but the AI illustrations were off-putting enough to close the tab.
- runjake 1mo agoNot the OP, but because it appears structured just like AI output? (This isn't a condemnation. AI can often do a better job of representing thoughts than humans.) Examples: - The article organization - The general language flow - The bullet points with a bolded gist, a colon, and then elaboration (and bonus with details stats). - The images are almost certainly AI generated. They look AI generated.
- conartist6 1mo agoThat's the thing: even if you don't use AI, if most of what you read is AI, soon this is what you'll sound like, just because it's so much of the training data your own brain's model has. Using it slowly sucks the uniqueness out of you.
- throwaway219450 1mo agoUsing AI to write blogs doesn’t bother me in principle, but I hate the prose that gets left in. Even if it’s not AI writing, being human doesn’t give you a pass for writing like a self-help guru. This sort of grammar is straight out of Claude: > A human on a VPN sees one CAPTCHA and passes, but a headless browser fleet sees a wall. I do wish we’d stop complaining about em dashes though, that’s lazy criticism. The sentence structure is far worse, and a big tell is subheadings that are all variants of “The <adjective> <noun phrase>”.
- runjake 1mo agoUsing a global agent rule with something like the following makes agent output much more tolerable for me: "Use only ASD-STE100 Simplified Technical English when communicating with me or writing documents, comments, and other communications intended for humans." As a result, I've never seen Claude or Codex use it's weird terms like "load-bearing".
- duskdozer 1mo agoWell, I definitely never saw nearly as many emdashes a few years ago. I didn't even realize it wasn't just a typeface quirk as I'd used hyphens a lot. I don't even know how to type an emdash on my devices and would have to copy/paste it from somewhere. But yes your other points are absolutely right
- behole 1mo agoHN's blanket AI allegations are almost as annoying and tired as the thing they are rallying against. I think you are AI.
- bookofjoe 1mo ago>HN's blanket AI allegations are almost as annoying and tired as the thing they are rallying against. Best thing I've read on HN so far this year.
- mysterydip 1mo agoThat’s exactly what an AI prompted to reply to AI accusations would say!
- qbane 1mo ago> And yes, my site gets its data by scraping those public documents. So I'm a scraper writing a blog post complaining about scrapers. I'm aware of how that sounds.
- ethersteeds 1mo agoWho scrapes the scrapemen?
- ihuman 1mo agoThere's a difference between someone running a scraping tool occasionally and bots constantly and rapidly re-scraping the same site over and over again
- qbane 1mo agoThis is an important context: the site is more likely to be targeted by scrapers because it is a curated collection of scraped information.
- 0cf8612b2e1e 1mo agoDo the bots care? Seemingly very little intelligence in many of them. Could be as simple as the site has more pages, so more traffic. Loot first, ask questions later.
- everybodyknows 1mo agoA curated collection of people who give away money. Was ever sweeter honey ever found in a pot?
- Lalabadie 1mo agoThe author's website is responsible for storing its own data. AI services currently treat the entire web as their storage and cache layer.
- alansaber 1mo agoLive by the scraper, die by the scraper
- dzonga 1mo agoblocking by geo yeah might work - but what happens when someone is traveling abroad ? they've to use a VPN to access your site ? my take with all the bots - the web is gonna be a bunch of private walled gardens. with most sites set to no index. you will only discover them via referral from someone real.
- yoursred 1mo agoWhat happens when someone is from a shit country?
- RebelMonk 1mo agoAs someone living in Vietnam, I have a number of websites off limits. The one in this thread for one https://www.cam.ac.uk/ https://www.cam.ac.uk/ is an interesting one that I'm forbidden to see. VPN is a necessity. I spend some of my day turning it on to see sites blocking Vietnam, and then some of it turning it off because of sites that are blocking the VPN.
- mcraiha 1mo agoAFAIK you already have to use VPN if you are e.g. westerner visiting China or Belarus.
- landgenoot 1mo agoOther sites are blocking VPN's as well. It's a contineous hassle of turning VPN on and off. It feels like you are a third class citizen.
- dzonga 1mo agogood question ? I worry about that too since I'm from a poor/shit country that I travel to constantly & live for half the time !!
- zuzululu 1mo agopretty crazy that this article is seemingly written by an AI used a detector and it is 95% confident its generated
- r0b0tan 1mo agoMost AI detectors are dogshit, though. AI is trained on human writing and tries to imitate it, while these detectors basically flag certain writing styles as "AI-generated."
- storus 1mo agoIs there any cheap way to run personal websites without getting cooked by bots these days, degrading the performance? A $5/month VPS won't cut it anymore.
- yoursred 1mo agoSef-host with tailscale or something similar
- dmux 1mo agoExactly. I've been hosting a site from a spare M1 Macbook Pro with Ngrok. An equivalent bare metal server would be much more expensive.
- rglover 1mo agoAll of my boxes are cheap VPS. Highly recommend people throw Cloudflare in front of their stuff. I switched all of my load balancers over to there and all of that bot crap went away. That combined with a proper UFW setup keeps the weather clear for me.
- inigyou 1mo agoHighly do not recommend centralising the internet.
- esseph 1mo agoBuild a better service or better technology. If you can't, well then... We're stuck.
- inigyou 1mo agoUntil proven otherwise in your specific case, the better technology is just hosting directly on a VPS without Cloudflare. And if your 5$ VPS is maxed out, while serving any less than 10 requests per second, then you need to optimize your software before considering an upgrade.
- ddxv 1mo agoI'm in a similar boat. Probably 99.999% is bots. I have nearly 1m unique "visitors" according to cloudflare and my real users are in the dozens a day. That being said, I love the open internet and am holding on to keeping as much open as I can.
- hskalin 1mo agoBut what do these bots gain from this?
- esseph 1mo agoThe most recent info of every single fucking thing on the web. Everything. Data locusts.
- bredren 1mo agoPossibly content that appears only briefly? I'm not sure otherwise. It seems wasteful.
- ddxv 1mo agoIt's all OpenAI / Anthropic / Singapore/China based crawlers. I guess they gain data that I provide for free. I get very little crawling by Googlebot / Bing etc which barely even index my site, I only have a single page "indexed" by Google.
- AlienRobot 1mo agoHave you ever use an AI chatbot? It's basically a democratized scraper. Type something stupid, it "visits" 200 websites and regurgitates some random gibberish.
- fooey 1mo agoI shutdown all my little informational hobby projects this year that I've tinkered with for decades They all shifted from mostly paying for themselves (or being so cheap it didn't matter), to essentially producing zero income while resource usage leapt up in magnitudes I couldn't justify the stress and hassle of making sites, that a few dozen people a day might find useful, into some complex hyperscalable obligations just to feed the bots
- andai 1mo ago> There's a real conflict here. I want Google and Bing and DuckDuckGo to crawl my site and send me new readers. But I don't want everyone else strip-mining it. Kinda sounds like we're missing a peer to peer network here. Instead of downloading the same data over and over again we can just download it once and then share it. Wouldn't that be better for everyone involved? It would also function as a distributed WayBack Machine, in case anything ever happens to the Internet Archive. (Which I think is desperately needed, bot apocalypse aside.)
- amazingamazing 1mo agoThe problem there is trust
- econ 1mo agoYou can validate by repeating the work and sign your content. The domain name system is now just a rent seeking scheme. It was great as a temp solution but over time it has deleted more content than preserved. I might in theory be billed for having a country name but I don't pay for a city, street name, house number or postal code. Online you should be able to move your widget shop to widget street. Can bill people who want to live on real estate street or used car street and/or set some requirements.
- NicuCalcea 1mo agoCommon Crawl? https://commoncrawl.org/ https://commoncrawl.org/
- ccgreg 1mo agoThe author blocked CCBot even though CCBot isn't part of the high traffic problem -- apparently he trusted Cloudflare labeling us as an "AI Bot".
- inigyou 1mo agoAnyway, Google doesn't send traffic to your site any more. Only important sites and obvious scams seem to get indexed.
- hmokiguess 1mo ago[dead]
- righthand 1mo ago> Then in November 2025, four thousand "visitors" showed up over a few days. Each visited exactly one page with a bounce rate of 99%. More telling was that they had no referrer. That's usually the easiest way I spot a bot. > And they were only crawling my fund pages (like this one, this one, and this one), which only 10% of real visitors ever touch. > But those 4,000 bots were just the warm-up. I just hate this style of writing like you're on Twitter. Why does the above need to be 3 different paragraphs? A paragraph break indicates a separate thought but the author is still talking about the same data and still making their point. The sentence "But those 4,000 bots were just the warm-up." is effective when still the last line of a paragraph and it signals respect for your readers. I stopped reading after this because it's just a terrible reading experience. Here's a correct version that doesn't read like the author left for a week to think about what the next sentence would be or having some sort of anxiety-induced mental pause: > Then in November 2025, four thousand "visitors" showed up over a few days. Each visited exactly one page with a bounce rate of 99%. More telling was that they had no referrer, which is usually the easiest way I spot a bot. They were only crawling my fund pages (like this one, this one, and this one), which only 10% of real visitors ever touch. But those 4,000 bots were just the warm-up.
- nickgray 1mo agoHey! I'm the OP - thanks for feedback on my writing style. I went ahead and fixed this in the article. It should be updated by the time you read this: https://patronview.com/news/99-percent-of-my-website-traffic-is-bots/ https://patronview.com/news/99-percent-of-my-website-traffic... And you're totally right: I mostly post on X (nee Twitter) and I probably have ADHD or just a low attention span, so I prefer to read things broken up into paragraphs. But for a smarter audience like this, and that reads long-form blog posts, I should tighten it up. Thank you for the suggestion. LMK any other edits and I'll be happy to tighten it up.
- righthand 1mo agoGlad to hear you’re willing to accept feedback. You maybe don’t have ADHD and sorry for continued advice but you shouldn’t assume you have undiagnosed conditions IMO as it allows you to defer your mistakes from the self. Even if you do have ADHD you can still correct and understand good article structure. I highlighted the last sentence of that paragraph because what is clear from the split sentence style is that you’re writing for impact, this lends well to 140 characters but falls apart in longer form writing but as I stated the sentence is still impactful as you’re saying “but wait…there’s more to this!” Which is very intriguing. I think if you’re comfortable writing that way and it helps you split your ideas and sentences up so each one is impactful, that’s a good thing. But consider that style as a draft and then you can go back and group up your impactful ideas into paragraphs very easily. For other edits I think my qualm applies to other parts of the article but I found that specific paragraph the most impactful way to illustrate what I was talking about. I leave the rest to you as a challenge. Don’t lose sleep over it, there will be more writing in the future to apply it to. As for the ADHD stuff and the urge to self diagnose consider something less severe but similar symptoms. Have you considered VAST? Here is a good HN comment briefly detailing it and mentioning a book (titled ADHD 2.0 I believe) that may be more in line. I am not a doctor of course and VAST is rather new. https://news.ycombinator.com/item?id=49035436 https://news.ycombinator.com/item?id=49035436 Anyways I will finish reading your article now since you’re so wonderful to take a bit of feedback and be proactive.
- johnorourke 1mo agoAnubis[1] is a superb fix for sites not behind Cloudflare/Fastly/Bunny etc. We had millions of bot requests, on a site serving all countries so we couldn't block by country, with fake user-agents so we couldn't block using that. It uses 'proof of work' to detect real browser software. [1] https://anubis.techaro.lol/ https://anubis.techaro.lol/
- drum55 1mo agoWhich is trivially bypassed by an actual implementation of the proof of work in non-javascript, rendering it absolutely useless. The website is approximately 3800x times slower than native code, and hundreds of thousands of times slower than the CUDA kernel claude wrote. The "proof of work" is just non existent at that point, they're solved in milliseconds for what would take the browser version 10 minutes or more, it's security by obscurity being dressed up as something more. pow_server http://127.0.0.1:8080 backend avx512-x16 ────────────────────────────────────────────────────────── uptime 00:03:12 solver ● BUSY difficulty 9, 0.3s queue [####################............] 5/8 peak 12 ────────────────────────────────────────────────────────── accepted 1240 solved 1180 503 shed 48 504 timeout 2 4xx/5xx 10 ────────────────────────────────────────────────────────── last difficulty 5 nonce 645376 in 9 ms (101.6MH/s, avx512-x16) hashes 3.90GH total avg 65.3MH/s Ctrl-C to stop Claude even made a nice little API server for it after implementing midstate compression, AVX multi way hashing, and a CUDA kernel. This doesn't stop the literal LLM it's trying to block from solving the challenges, it's really annoying that everybody is using it and claiming that it's something that's usable in the real world as a result of it using proof of work. It's obscure, and obscure is fine so long as nobody is pretending that it is secure.
- gum_wobble 1mo agohow so, can you link to any sources?
- drum55 1mo agoThe prompt used for Opus 4.8 was: write a implementation of the anubis proof of work in native c code, optimized for speed above all else. use every trick available to make the proof of work as efficient and fast as possible, including modern processor tricks on the x86 platform. your code should avoid using external libraries where possible, include tests, and be readable and concise. a reference for what needs to be met is in this repository. https://github.com/TecharoHQ/anubis Then let’s develop this more. turn this solver into a local HTTP server that can be given work in the request, and it returns solved work. make an end to end tester that sends test work to the solver and waits for a valid response. add support for solving with a GPU using cuda. Then it was done more or less, it happily made a local server that supports solving the challenges given to it in bulk with priority based queue and can tolerate potentially tens of thousands of requests a second with no issue. The CPU time spent solving the challenges is less than the SSL setup for the connections. The GPU version does in excess of 20GH/s (but with high latency) though I didn't really test it, I'm not using this for anything but proving a point that the LLM itself can write the bypass tools and run them happily.
- deleted 1mo ago[deleted]
- Bender 1mo agoSeems about right. I rotated my logs this morning. Most humans go to access.log and most bots go to botpoop.log. This is the line count: 2 access.log [1] 40 botpoop.log [2] 2 is really 1 since a human will grab the CSS file. Most bots do not bother with the style-sheet so that's a 40:1 bots to humans. I could cut that down by blocking data-centers but then I inadvertently block a lot of VPN's which I really don't need to do for a static compressed blog served from ram. The bots just get a TCP Reset but it's still fun to log and study them. The most interesting one I've seen recently is ReadYou which may be a reader but it appears to be much more, possibly acting as a cell phone distributed bot collecting data for a centralized site. [1] - https://nochan.net/logs/access.log https://nochan.net/logs/access.log [2] - https://nochan.net/logs/botpoop.log https://nochan.net/logs/botpoop.log
- spockz 1mo agoIs there some existing mechanism already that counts how often an ip only scrapes the page and not the css and then block those origin IPs if it occurs “too often”? Unfortunately, the best practice is to make css cacheable so you need to keep long histories.
- Bender 1mo agoI thought about that but to your point CSS is cachable. In fact I made mine immutable. No I just visually spot patterns and use that to study other facets of the agent, other client headers or lack thereof, supported protocol, accepted encoding and so on.
- spockz 1mo agoMaybe it is enough to include some css/js which is served without cache and is loaded after all user visible css/js is loaded. Make it small enough to not cause too much bandwidth for the server and legitimate clients. Then anyone who doesn’t hit that CSS file gets banned.
- Bender 1mo ago
- Venn1 1mo agoI'm blocking the Amazon search crawler, anything coming from Googleusercontent, and limiting AI crawlers to search rather than allowing AI assistants. The residential proxy waves are something to behold, but Cloudflare does an okay job catching those in the AI labyrinth. Still, it's all a bit silly, and I can't imagine what large sites deal with when I'm tangoing with this much nonsense on a small tech blog.
- inigyou 1mo agoDon't. Just serve the page unless it's an unusually expensive page to serve (like search).
- ashu1461 1mo agoI wonder if the author tried out the recently released feature by cloudfare to block ai bots https://developers.cloudflare.com/bots/additional-configurations/block-ai-bots/ https://developers.cloudflare.com/bots/additional-configurat...
- smolder 1mo agoSome people don't want to use cloudflare on principle. Like that putting the whole internet behind cloudflare or AWS is a bad thing, in principle.
- mgbmtl 1mo agoI run scripts on my servers on an hourly basis to check which are the top 25 IPs visiting the server (aggregated by /24). If anyone in those top 25 IPs are from China, Vietnam, etc, or from Alibaba/Amazon/etc, the /24 gets blocked by iptables. It's far from perfect, but it was a quick way to get rid of bots, while not completely blocking people from countries such as Vietnam. However, on a Gitlab instance I manage (500 users), we have to restrict viewing of git logs and pretty much everything except issues. The bots were too aggressive. Chinese crawlers have access to a huge range of IPs and they often do only 10-20 requests per day, while generating in total over 50k requests per day. Our server load went from 99% down to 0.1% after that (and it's a fairly big server).
- IMSAI8080 1mo agoMy solution was similar. Anything coming from the ASN of a major cloud provider gets a CAPTCHA with a little nuance to allow Google and Bing to index. That seems to do a pretty good job. Also, anything coming out of China or Singapore also gets a CAPTCHA as my site is not popular in those regions and many Chinese bots seem to show up as a Chinese mobile provider. So far, the bots have never attempted to solve the CAPTCHA.
- tarr11 1mo ago> My normal bill for running this whole site is around $90 a month. During one bad spike month, it jumped about 500%. This is D1 - which has very surprising costs. you may just want to drop D1 and move to a static site. There’s no reason your site should cost this much.
- inigyou 1mo agoThere's a lot of people who host their site at extremely expensive places and then do everything they can to minimise unneeded traffic - instead of just moving to a cheaper host. Vercel is another popular extremely expensive host.
- creshal 1mo agoI'm always flabbergasted when I see what people pay and how much effort they need to invest to keep their cloud websites from eating them alive. My allegedly more complicated VPS stack needs an afternoon of attention every two years when a new Debian major release is necessary, and costs have been predictable for 15 years, no matter what happened traffic wise.
- inigyou 1mo agoThe predictability of costs is underrated too. You don't want your hosting solution to automatically scale up to $20,000. You want that if it's overloaded it's simply overloaded.
- tailscaler2026 1mo ago[dead]
- nickgray 1mo agoThank you! I need to tighten up my KV compression, which is actually carrying a lot of D1's load otherwise. We also had some bad queries some months, as the pages and database grew, that were counting the wrong things (or extremely inefficiently) and those have since been fixed.
- _tgxm 1mo agoThe annoying part is a large percentage of misbehaving bots (not obeying robots.txt for example) are via end user proxies across the world. However most of these aren't doing full-browser loop, so if you are behind cloudflare, you can do non-interactive challenge and that can help quite a bit.
- nickgray 1mo agoYes! OP here. I did the non-interactive challenge, and yet all those Chinese bots in my article got through (which surprised me).
- AdrianB1 1mo agoI checked the comments to see if anyone pointed to this: I can imagine so many memes with this line :)
- FerretFred 1mo agoSigh .. same here. I don't write blog posts often (enough) but the ones I do write are from personal experiences and I take a lot of care with them. I look at my logs snd see bots everywhere, but now I just let them get on with it. AI scrapers are different though; they get to read my content which, just for them contains a smsttering of finest digital toxin. A pox on your datasets!
- nromiun 1mo agoThis is a static website running on Cloudflare infra. What on earth costs $90 per month? First optimize your infra before throwing up rules in front of your visitors. I have several websites on Cloudflare too and I don't even check how many million requests I get. Because it does not cost me anything. > And yes, my site gets its data by scraping those public documents. So I'm a scraper writing a blog post complaining about scrapers. I'm aware of how that sounds. Being self aware does not make it okey. Either you are okey with scraping (like me) or against it. Don't use it yourself and block your site at the same time. These same people will be crying about how Cloudflare ruins the internet because they get these captchas.
- knuckleheads 1mo agoPreviously, I had done a fair amount of research into how Google's monopoly on web crawling further entrenches their monopoly in the search engine market. You can read more about this here, https://knuckleheads.club https://knuckleheads.club, there is a long report from ~2020 or so that explains how it worked at the time. The club is mothballed, I am doing other things with my life, and I'm happy to say that we played a very small role in the DOJ ordering Google to share their crawl data with qualified competitors (a work in progress, but it's progressing). Chatbots have super charged this dynamic though, to the point that it is showing up in the robots.txt data. The last few weeks I've been having Claude rerun some old analysis of Common Crawl from back then, when I have spare usage and time. What I've found is that you can see pretty clearly the rise in people outright blocking AI chatbot related crawlers likely because of how aggressive they have become. Quarter Crawl GPT Claude CC G-Ext Byte Bing Google 2023 Q1 2023-06 0.00% 0.00% 0.16% 0.00% 0.06% 0.47% 0.39% 2023 Q2 2023-14 0.00% 0.00% 0.18% 0.00% 0.06% 0.45% 0.38% 2023 Q3 none — — — — — — — 2023 Q4 2023-40 2.21% 0.00% 2.12% 0.04% 0.11% 0.39% 0.27% 2024 Q1 2024-10 0.53% 0.05% 0.31% 0.09% 0.18% 0.34% 0.31% 2024 Q2 2024-18 0.55% 0.09% 0.32% 0.11% 0.24% 0.32% 0.31% 2024 Q3 2024-30 0.68% 0.22% 0.36% 0.20% 0.38% 0.24% 0.33% 2024 Q4 2024-42 1.10% 0.50% 0.44% 0.32% 0.50% 0.25% 0.40% 2025 Q1 2025-05 1.14% 0.66% 0.54% 0.42% 0.66% 0.25% 0.44% 2025 Q2 2025-18 1.37% 0.93% 0.63% 0.70% 0.92% 0.29% 0.19% 2025 Q3 2025-30 1.42% 1.07% 0.74% 0.62% 1.01% 0.31% 0.27% 2025 Q4 2025-43 1.92% 1.51% 1.23% 1.15% 1.52% 0.27% 0.19% 2026 Q1 2026-04 2.13% 1.76% 1.68% 1.58% 1.77% 0.22% 0.15% 2026 Q2 2026-17 2.80% 2.38% 2.26% 2.13% 2.50% 0.22% 0.14% 2026 Q3 2026-30 3.45% 3.01% 2.89% 2.71% 3.16% 0.21% 0.14% GPTBot is OpenAI, ClaudeBot is Anthropic, CCBot is Common Crawl, Google-Ext is a way for website owners to indicate they don't want their content to be used for AI, Bytespider is Bytedance, Bing and Google are the last two. Take these numbers with a truck of salt, haven't had time to verify them. It's very clear that website owners do not like getting their content scraped and are indicating to GPTBot et al. that they are not welcome. It's a shame that CCBot is caught in the cross fire, but that's life. Bing and Google are doing just fine though, almost like having significant power in the search engine market gives you an advantage in other markets too. Who knew!
- szundi 1mo ago[dead]
- jwr 1mo agoThe worrying thing here is that so many people accepted outsourcing the decision on who can see their website to a large company (Cloudflare). If the company decides that a certain user should not see the website, the user will not see the website, and no one will know about it, and the user will have no recourse. That is not the open web that I would like to see. A second side effect of a knee-jerk reaction to bots crawling websites is that if you try to fight all bots, you also end up hurting real users that use "bots". If I run a local script or an LLM that needs to fetch a web page from your site and you block that script or LLM as being a bot, you hurt me, the user. I wanted the information and I have no way of getting at it without investing my own time and effort in going to your website. That might or might not be what you expected, but it's worth taking into account. And finally, something worth noting is that there are so many websites whose owners complain about bots, but the real problem is that the website is poorly built and should be improved anyway. Bot traffic is not necessarily bad.
- bluefirebrand 1mo ago> I wanted the information and I have no way of getting at it without investing my own time and effort in going to your website Yes, that's the social contract. Bots are not a part of it
- eigencoder 1mo ago> I wanted the information and I have no way of getting at it without investing my own time and effort in going to your website. I think you've hit the nail on the head here. I think that's a big reason why people want to ban bots.
- aomix 1mo agoThe social contract of accepting scraping for visibility was always tenuous and is now fully dead. But it’s shocking the number of people who are essentially victim blaming here. You should spend YOUR time to optimize your free website so MY use case is not impacted. How is that not incredibly selfish on its face?
- 1mo ago
- bediger4000 1mo agoI think that asymmetry is why scraping keeps getting worse. The scrapers' costs fell faster than everyone's defenses improved. There it is. Just like the fckn spammers who ruined SMTP email, scrapers externalize the costs. Who finances the effort to use residential proxies? That takes a lot of effort, even if it's shoddy
- throwaway63467 1mo agoMost people don’t know they’re acting as a residential proxy, lots of devices and apps and free tools install spyware which often includes a proxy script. It’s a really shady market.
- IMSAI8080 1mo agoThere was a popular pirate streaming stick sold on Amazon that ran a residential proxy by day and did ad fraud clicks by night.
- wandr 1mo agoI'm working on a web app right now, with the intention of it going to be 100% paywalled. It's 95% complete, but the remaining 5% is just implementing the paywall. In the meantime, the app is live and operational with a fully functional signup. I am deleting about 100 new bot signups per day right now, it is crazy out there.
- imthenitto 1mo ago[flagged]
- deleted 1mo ago[deleted]
- varenc 1mo agoCan someone help me understand the underlying motivation behind this? It makes sense that some crawlers, in the style of Google, would want to index the entire internet. But what is the point of the same crawler re-fetching a page they already fetched an hour ago? Or possibly all this traffic is just independent entities, each trying to cache the internet? The scale of bot traffic makes this seem unlikely. What's the motivation behind the same entity re-fetching a page it just fetched less than an hour ago?
- esseph 1mo agoThe most recent data on the internet for advertising, intelligence, etc. And a lot of bad scrapers.
- bigbuppo 1mo agoThey are poorly implemented by the "fuck you I got mine" crowd. They will get stuck doing things like trying to run through a calendar that could theoretically go back to the beginning of time and all the way to the end of time. And because that calendar might change, it gets scraped for every inquiry made to the poorly implemented AI system.
- phillmv 1mo agoSimilarly, I've never understood the economics behind the constant rescraping that is flooding the internet or really what's triggering it. It can't all be agents reacting to user queries. It's confounding how much CPU and bandwidth is getting flushed down the drain.
- Ekaros 1mo agoIf you have enough money to burn you don't care about efficiency. I personally blame VC model. Efficiency and profitability doesn't matter. Just make line grow up. And when same mindset is applied to those companies with money to waste... Yeah waste will happen.
- BLKNSLVR 1mo agoMy laymans understanding is that poorly written scripts get executed and then owner comes back occasionally to check that there is 'content' in their database. They're not sitting there troubleshooting their thing beyond 'it's getting data' because, for similar reasons to their scripts being poorly written, their strategy is 'get data'. Nuance, complexity, and an awareness of 'other people' do not exist in their worlds.
- r0b0tan 1mo agoI think it’s wild how AI and data bots are putting certain business models under pressure. We’ve already seen the same thing happen with Tailwind.
- oaw93j4oij 1mo agoI despise cloudflare. They've decided that my home IP address is bad, so I have to capchas for most websites. Sometimes on infinite loop and I never get to the website. I even reset my home IP address more than once, but it instantly continues. Especially if I use any VPN, even my work VPN.
- GodelNumbering 1mo agoI just checked Cloudflare for SignalBloom (https://www.signalbloom.ai https://www.signalbloom.ai, which I own and operate). Over the last 72 hours, Claude-searchbot [1] alone fetched ~205,000 pages. Sent exactly 1 referral. There is a lot of free financial data on the site, hoping for real users to benefit from it. It is hard to not feel a little cheated out that Claude gets to claim "Found it!" to its users without me getting no credits or compensation whatsoever. [1] Exact user agent `Claude-SearchBot/1.0; +searchbot@anthropic.com)` Proof: https://i.postimg.cc/Pqc3SS8T/Screenshot-2026-08-07-at-5-33-18-PM.png https://i.postimg.cc/Pqc3SS8T/Screenshot-2026-08-07-at-5-33-...
- markdown 1mo agoFYI, your proof is hosted on a NSFW page. You should have mentioned that.
- potamic 1mo agoReason #13843 for using an ad blocker.
- markdown 1mo agoChrome blocked it, but that's irrelevant. People don't have total control over their computers at work, which is why the NSFW tag is used.
- GodelNumbering 1mo agoApologies, I had no idea. I lookup host image and that site came up. I don't think the site itself is of NSFW nature.
- SahAssar 1mo agoIsn't your whole product scraping other sites and summarizing? I also don't see clear sources for your info where it is displayed, so it seems like you do the exact same thing.
- 1mo ago
- DaveZale 1mo agoThis is like technological cannibalism. I stopped posting to my website. Why should it be so much work to stop this theft? Would it be helpful to have geofencing and regulation?
- brownieman1325 1mo agoI recently got a surge from Singapore and heard a lot more peeps in my circle of friends saying the same thing...
- IMSAI8080 1mo agoIt'll be the Chinese. The big Chinese cloud providers have data centres in Singapore for some reason. My Chinese mobile phone sends all it's telemetry back to Singapore, not mainland China.
- Bender 1mo agofor some reason It's outside the great firewall. No requirements to hand over SSL keys to China. I worked for companies that ran into these challenges and solutions.
- thorsson12 1mo agoThe experience of browsing the web has really suffered lately. The mandatory 3-4 second "verifying that you're a human" block from Cloudflare seem to show up on more and more websites. Seems like a questionable choice from Cloudflare to teach everyone to associate Cloudflare's logo with high latency...
- duskdozer 1mo agoA 3-4 second cloudflare wait seems good to me at this point. I had to just block any cloudflare requests because some pages that used it would just rev up a cpu core indefinitely and hang the browser.
- tananaev 1mo agoI also see quite a bit of traffic from China and Singapore. I wonder if it's some scraping for AI training. It doesn't really bother me too much because traffic is still fairly low, but it skews all the analytics for me.
- falcor84 1mo agoIt's a sequence of thin lines, between using Chrome out of the box, to using something like Brave, to using a highly customized Zen Browser, to having ChatGPT look up a particular page for you, to using a small BeautifulSoup/scrapy script to scrape dozens of pages, to scraping the entire web. And at every point on this spectrum it's humans driving a "user agent" tool to make requests and process responses on their behalf.
- somebudyelse 1mo agoI visited someone's blog and didn't have to solve a cloudflare challenge! That's amazing work
- l72 1mo agoI don't look at the logs of my personal site very often as it is a static site, but just went to check, and yeah, it's almost all ai crawlers. Note sure what is going on here, but I hope this isn't really anthropic: 34.124.XXX.XXX - - [07/Aug/2026:07:55:31 -0700] "GET /secrets.json HTTP/2.0" 404 343 "-" "anthropic-ai" 34.124.XXX.XXX - - [07/Aug/2026:07:55:31 -0700] "GET /credentials.json HTTP/2.0" 404 343 "-" "anthropic-ai" 34.124.XXX.XXX - - [07/Aug/2026:07:55:31 -0700] "GET /secrets.yml HTTP/2.0" 404 343 "-" "anthropic-ai" 34.124.XXX.XXX - - [07/Aug/2026:07:55:31 -0700] "GET /service-account.json HTTP/2.0" 404 343 "-" "anthropic-ai" 34.124.XXX.XXX - - [07/Aug/2026:07:55:31 -0700] "GET /key.json HTTP/2.0" 404 343 "-" "anthropic-ai" 34.124.XXX.XXX - - [07/Aug/2026:07:55:31 -0700] "GET /config/.env HTTP/2.0" 404 343 "-" "anthropic-ai" 34.124.XXX.XXX - - [07/Aug/2026:07:55:32 -0700] "GET /service_account.json HTTP/2.0" 404 366 "-" "anthropic-ai" 34.124.XXX.XXX - - [07/Aug/2026:07:55:32 -0700] "GET /serviceAccountKey.json HTTP/2.0" 404 343 "-" "anthropic-ai" 34.124.XXX.XXX - - [07/Aug/2026:07:55:32 -0700] "GET /firebase-adminsdk.json HTTP/2.0" 404 343 "-" "anthropic-ai" 34.124.XXX.XXX - - [07/Aug/2026:07:55:32 -0700] "GET /Dockerfile HTTP/2.0" 404 343 "-" "anthropic-ai" 34.124.XXX.XXX - - [07/Aug/2026:07:55:32 -0700] "GET /.github/.env HTTP/2.0" 404 343 "-" "anthropic-ai" 34.124.XXX.XXX - - [07/Aug/2026:07:55:32 -0700] "GET /firebase-service-account.json HTTP/2.0" 404 343 "-" "anthropic-ai" 34.124.XXX.XXX - - [07/Aug/2026:07:55:32 -0700] "GET /.docker/config.json HTTP/2.0" 404 343 "-" "anthropic-ai" 34.124.XXX.XXX - - [07/Aug/2026:07:55:32 -0700] "GET /.npmrc HTTP/2.0" 404 343 "-" "anthropic-ai" 34.124.XXX.XXX - - [07/Aug/2026:07:55:32 -0700] "GET /.boto HTTP/2.0" 404 343 "-" "anthropic-ai" 34.124.XXX.XXX - - [07/Aug/2026:07:55:32 -0700] "GET /.s3cfg HTTP/2.0" 404 343 "-" "anthropic-ai" 34.124.XXX.XXX - - [07/Aug/2026:07:55:32 -0700] "GET /.svn/entries HTTP/2.0" 404 343 "-" "anthropic-ai" 34.124.XXX.XXX - - [07/Aug/2026:07:55:32 -0700] "GET /.htpasswd HTTP/2.0" 404 346 "-" "anthropic-ai" 34.124.XXX.XXX - - [07/Aug/2026:07:55:32 -0700] "GET /terraform.tfstate HTTP/2.0" 404 343 "-" "anthropic-ai" 34.124.XXX.XXX - - [07/Aug/2026:07:55:32 -0700] "GET /docker-compose.yaml HTTP/2.0" 404 343 "-" "anthropic-ai" 34.124.XXX.XXX - - [07/Aug/2026:07:55:32 -0700] "GET /.vscode/launch.json HTTP/2.0" 404 343 "-" "anthropic-ai" 34.124.XXX.XXX - - [07/Aug/2026:07:55:32 -0700] "GET /.ssh/id_rsa HTTP/2.0" 404 343 "-" "anthropic-ai" 34.124.XXX.XXX - - [07/Aug/2026:07:55:32 -0700] "GET /.ssh/id_ed25519 HTTP/2.0" 404 343 "-" "anthropic-ai" 34.124.XXX.XXX - - [07/Aug/2026:07:55:32 -0700] "GET /.ssh/id_ecdsa HTTP/2.0" 404 343 "-" "anthropic-ai" 34.124.XXX.XXX - - [07/Aug/2026:07:55:32 -0700] "GET /.ssh/authorized_keys HTTP/2.0" 404 343 "-" "anthropic-ai" 34.124.XXX.XXX - - [07/Aug/2026:07:55:32 -0700] "GET /.ssh/known_hosts HTTP/2.0" 404 343 "-" "anthropic-ai"
- luciana1u 1mo ago[flagged]
- travisgriggs 1mo agoI wonder what the web would be like if we priced bandwidth at the requester point and as you go.
- PeterHolzwarth 1mo agoYour comment kind of reminds me of the old slashdot spam-email-solution copypasta - the purpose of it is to highlight how so many sensible sounding ideas just can't work in practice.
- thomashabets2 1mo agoSomething suspiciously absent from this article is addressing whether some of these are in fact human visitors, but humans issuing chatgpt or similar queries instead of going directly. Is it really a bot if it's in response to a human asking for some aggregate information about charities, triggering a web search and then following the result links to get details for the human? Well, clearly yes it is, but it's a very different proposition from this article's implication that "they have no throttling on their scrapers"[1]. > Challenge 46 datacenter ASNs. Humans don't browse from AWS. People who have workstations in the cloud do. > The bots use 99% of the bill and I pay 100% of it. Running a site this way is always a wallet-DDoS risk. [1] though yes, by far most will be pure automation with no human in the loop. It's an assumption on my part, but feels like a safe one.
- bakugo 1mo ago> Something suspiciously absent from this article is addressing whether some of these are in fact human visitors, but humans issuing chatgpt or similar queries instead of going directly. That's not absent from the article, it's right there in the section titled "The Claude ratio". ChatGPT, Claude, etc. use different user-agents for scraping vs user-initiated requests, and the author notes that user-initiated requests were an absolutely miniscule fraction of the total traffic.
- thomashabets2 1mo agoOh, that's what "Anthropic's search crawler, had requested 420,680 pages in one week. That same week, Claude sent me 12 human visitors" meant? Yeah, googling it does seem like "Claude-User" is for user-initiated requests. By "Claude sent me" I thought the author meant referer header in real browser requests showed that they came from. I mean, that's what the section "Pages crawled per visitor referred" refers to, right?
- sethops1 1mo ago> Running a site this way What do you mean "this way". What other way is there to run the site?
- butz 1mo agoCloudflare's "Verify you are human" captcha is the new cookie banner.
- 1vuio0pswjnm7 1mo agoOriginal HN title: "99% of My Website Traffic Is Bots" Not clear how the author arrived at the precise 99% figure; perhaps "99%" is a figure of speech "And yes, my site gets its data by scraping those public documents. So I'm a scraper writing a blog post complaining about scrapers." "I'm trying to run a business here." What's the business (where "business" is defined as "buying and selling") From https://patronview.com/robots.txt https://patronview.com/robots.txt # As a condition of accessing this website, you agree to abide by the following # content signals: # (a) If a Content-Signal = yes, you may collect content for the corresponding # use. # (b) If a Content-Signal = no, you may not collect content for the # corresponding use. # (c) If the website operator does not include a Content-Signal for a # corresponding use, the website operator neither grants nor restricts # permission via Content-Signal with respect to the corresponding use. # The content signals and their meanings are: # search: building a search index and providing search results (e.g., returning # hyperlinks and short excerpts from your website's contents). Search does not # include providing AI-generated search summaries. # ai-input: inputting content into one or more AI models (e.g., retrieval # augmented generation, grounding, or other real-time taking of content for # generative AI search answers). # ai-train: training or fine-tuning AI models. # use: how AI systems may consume the content (immediate, reference, or full). # ANY RESTRICTIONS EXPRESSED VIA CONTENT SIGNALS ARE EXPRESS RESERVATIONS OF # RIGHTS UNDER ARTICLE 4 OF THE EUROPEAN UNION DIRECTIVE 2019/790 ON COPYRIGHT # AND RELATED RIGHTS IN THE DIGITAL SINGLE MARKET. # BEGIN Cloudflare Managed content User-agent: * Content-Signal: search=yes,ai-train=no,use=reference Allow: / Perhaps this could be construed as a license, e.g., permitting or prohibiting certain uses of the "content" If, for example, the website operator had enforceable intellectual property rights in the "content", such as copyrights, then perhaps the operator could restrict access to the "content" under the threat of litigation to enforce those rights Basic questions 1. Is the "content" protected by intellectual property rights, e.g., copyrights 2. Does the website operator have intellectual property rights in the "content", e.g., copyrights 3. Does the website operator have agreements with the rights holders, e.g., granting the operator authorization to restrict access to the "content"
- rsolva 1mo agoI made a small booking site for a local dutch canal boat, which has a calender function. A simple PHP app. I checked the Apache logs recently, and it had THOUSANDS of claudebot and other AI UserAgents flooding the logs every day, apparently because the scrapers keep hitting the 'next month' button on the calendar in a an infinite loop, all day, everyday! This is a small booking app without any useful information at all, it surprises me that the AI boots have no discernment about what the are scraping, just wasting their own and other peoples resources. And their own reputation! You would thing they could spare a few tokens on a classifier model to do a quick evaluation of their scraping efforts, but apparently they do not. Anyway, I have done my best to block these UAs and so far it seems to have improved the situation.
- Anamon 1mo agoIt's vibecoders all the way down. I'm sure Anthropic et al. would like to reduce their resource waste if they could, seeing the insane amounts of money they're bleeding. But I had to come to the conclusion that these LLM companies have not a single developer good enough to implement a decent scraper. Their tools produce garbage code and they don't know enough about programming to realise it, or do anything about it.
- DoesntMatter22 1mo agoI'm surprised that the solution wasn't to poison the data so they will train on junk data
- grigio 1mo agoLLM are the best way to read the web. right to the point, no ADS
- timbit42 1mo ago...yet.
- ticktock 1mo ago[dead]
- scotty79 1mo ago> My visitor stats got so polluted I couldn't trust my own numbers. This was the one that hurt. I'm trying to run a business here. I want to know what real people read on my site so I know what to build next. I couldn't see them through the bots. Are bots the solution to nosy websites that want to know things about their visitors?
- urrightimsorry 1mo ago[flagged]
- urrightimsorry 1mo ago[flagged]
- urrightimsorry 1mo ago[flagged]
- urrightimsorry 1mo ago[flagged]
- urrightimsorry 1mo ago[flagged]
- urrightimsorry 1mo ago[dead]
- deleted 1mo ago[deleted]
- urrightimsorry 1mo ago[flagged]
- Aeolun 1mo agoDid he really say that the VPS would be in trouble if it received 5 requests per second of mostly static pages? Something that currently costs $90 per month to serve through CloudFlare?
- joshspankit 1mo agoHow long until Cloudflare is the data broker for websites like this? “For a low $/GB, we’ll give you everything from this site and 1000 others as (structured data/a database)!” (yes there are lots of good counter arguments to this, but before you reply think ahead a couple extra steps)
- peter_d_sherman 1mo agoAn interesting article, and an interesting problem to have! The problem, to recap the title, is that "99% of My Website Traffic Is Bots". How to fight all of those bots, all of those web-scrapers... If the sole underlying issue is bandwidth for one's website (someone doesn't have enough of it and/or they pay too much for it), then one possible solution to this might be for a single tech company to create an online cache of web pages, accessible to all bots as an alternate route to fetch the desired web pages. If a tech company (or consortium thereof) decided to do this, they could certainly charge money to each AI / bot utilizing the service... Which means that it might be an investable idea as a for-profit service... Now, if the underlying issue is the privacy of individual web pages (or relative privacy, as the case may be!) from the public internet, then perhaps the solution is to simply put those behind user logins and/or a gauntlet of tests (which could also be tied to login!) designed to exclude AI/bot web traffic from human web traffic. Consider what some BBS'es and the online services of yesteryear used to do (Compuserve, The Source, Prodigy, QuantumLink, AOL, etc.)... basically they'd get a user's mailing address, and then physically snail mail them their login and password to their physical home address... Oh sure, user registration wasn't as fast back then as it is today... it might take several days to receive your username and password via postal mail -- but as a system operator you were 99.999999999% guaranteed that when you saw such a login on your BBS, that it was an actual live human user... There may be a market for a service like that, too... That is, validate that actual users are actual users, and give them some kind of credential that can be checked by an actual web site... find a way to do this at scale... Anyway, just thinking aloud... a very interesting article, and a very interesting problem to have!
- tomveber 1mo ago[dead]
- itake 1mo agoI live in Vietnam and was annoyed I can't access the website. You still allow Mullvad VPN users though
- rurban 1mo agoI had similar problems with my free movie festival ratings database, so eventually I had to get rid of the dynamic lookup, and I only dump now static sites, hosted on GitHub Pages. I could not fight the AI bots for free, the hosting CGI constantly ran out of memory.
- aorth 1mo agoThis will resonate with anyone who operates a public-facing website and doesn't work at a big tech company. In the past few years I have also been down this exact same road. I'm desperately trying to avoid resorting to Cloudflare, but I'm running out of time and patience to keep tweaking nginx and firewall rules every few weeks. I had initial success with https://git.gammaspectra.live/git/go-away https://git.gammaspectra.live/git/go-away as a more powerful and more reasonable self-hosted alternative to Cloudflare than Anubis, but it seems to have gone unmaintained. There is also https://github.com/dgl/haphash https://github.com/dgl/haphash if you run HAProxy, though I have not tried it. On that note, does anyone know what happened to Ted Unangst aka tedu? He was a prolific OpenBSD developer and blogger, and he had developed one of his own simple solutions https://humungus.tedunangst.com/r/anticrawl https://humungus.tedunangst.com/r/anticrawl, but all his web properties seem to have gone away recently...?
- astoilkov 1mo agoWhat happens when these Ohio IPs get their Chrome upgraded and you no longer can detect them?
- rozumem 1mo agoI run a site which gets around 60k real MAUs per month and have had to deal with this in some shape or form. I haven't given it as much attention as you have though, because it started to feel like whack-a-mole like you describe and other things took priority. I've concluded that the most viable long term solution is to move most pages behind a login wall. And only keep a page public if I see the ROI from the page being scraped by AI bots, Google, etc... for you this could be only the really high-profile donor pages, but not the long tail pages. Once pages are behind a login wall, we can track anomalous activity at a behavioral / account level and automatically suspend those accounts. For example, an account visiting every donor page will be easy to detect.
- landgenoot 1mo ago> Sorry, you have been blocked You are unable to access patronview.com Why have I been blocked? > This website is using a security service to protect itself from online attacks. The action you just performed triggered the security solution. There are several actions that could trigger this block including submitting a certain word or phrase, a SQL command or malformed data. Thank you, from Vietnam.
- aitchnyu 1mo agoI see popular blog posts, click it and see its blocked from India.
- pmdr 1mo agoThe only winner here is Cloudflare. Nowadays they have 30s wait screens for residential IPs they don't like. It's things like this that make me quit the internet altogether.
- zImPatrick 1mo agoRegarding the residental botnets: check their TLS fingerprint, they usually (at least from my experience) have the same few JA4 hashes that don‘t match with modern browsers
- aitchnyu 1mo agoAre they all LG TVs or similar?
- zImPatrick 1mo agoNot sure, I couldn‘t find those fingerprints anywhere and I couldn‘t replicate them in any browser I had access to (Safari iOS, Windows & Linux Firefox & Chromium), so I decided to block them. I haven‘t had a single person complain yet. Their UA are usually the most random ones you can find (who uses a PPC mac nowadays?)
- mohsinq227 1mo ago[flagged]
- RebelMonk 1mo agoAhh, the joys of accessing sites “protected” by Cloudflare when living in Vietnam: Sorry, you have been blocked You are unable to access patronview.com