7 ms·
It’s more economical at a compute level, but not at the developer level. The moment you start customizing your crawler to use protocol X for site Y your scale s
by pbronez 18d ago
It’s more economical at a compute level, but not at the developer level. The moment you start customizing your crawler to use protocol X for site Y your scale story collapses.
- paytonjjones 18d agoIt's a good point, but in practice it depends on how easy those customizations are to implement / maintain, and how much money and effort you save. At some point the compute cost can disrupt even the nicest scale story. I think the path forward is that websites offer one path for humans, and another for scrapers. But the huge catch is the path for scrapers must be _genuinely_ and _reliably_ the more economical and scalable path (either through something like PoW arms races, or through fear of litigation). Otherwise they will continue to ignore instructions and intrude on the human path.
- inigyou 18d agoWhy aren't we litigating against scrapers, anyway? DDoS is a felony.
- miki123211 18d agoBecause they're all in countries who would ignore such litigation.
- ekidd 18d agoLargely because they're residential botnets in places like Brazil (a real example from one of my sites that was crawled to near-destruction). Someone could probably do something about this, but it's out of reach for individual site owners.
- inigyou 18d agoIf you block Brazil, they'll find an alternative, maybe then you can sue them.
- afdbcreid 18d agoBut they operate from neither. You can at most sue the one renting them IP addresses. Which will do basically nothing. (Also, blocking a whole country is likely not what you do, but you probably know that).
- inigyou 18d agoWhy do you think you can only use the one who's renting them IP addresses?
- semiquaver 18d agoHow good are you personally at iteratively and sequentially jumping through the legal systems of dozens of countries over the course of years, interleaved with genuinely difficult technical investigation, to unmask successive onion-layers of identity in order to unmask one offender? Oh, and it also only takes a few minutes to reconfigure everything and invalidate those years of legal and investigatory work.
- inigyou 17d agoWhy do you think OpenAI, Anthropic, Bright Data, and Comcast aren't US companies?
- articulatepang 17d agoWho’s “we”? I don’t want to do that work. I don’t think kernel maintainers do, either. Do you?
- miki123211 18d agoThe temptation to offer different / inferior / limited content to scrapers will be too strong, so such solutions are doomed to fail. This would likely work as "Cloudflare SideChannel", a (hypothetical) Cloudflare product that would let scrapers download the pages that humans actually visit, as they are added to the CF cache. It wouldn't work for the non-Cloudflare part of the internet where humans connect directly to the servers that have their content.
- simonw 18d agoThe moment we establish a standard for offering a "optimized for scrapers" version of a site, people who do not want to be scraped will weaponise that to serve junk to scrapers... and scrapers will subsequently refuse to use it.
- eru 18d ago> It’s more economical at a compute level, but not at the developer level. Developers and compute are interchangeable now.
- nozzlegear 18d agoWho's proompting the machine to do it differently without a developer there to ask the right questions?
- eru 18d agoYou can have a high level prompt of: make our crawling cheaper and more reliable to run.
- deleted 18d ago[deleted]