5 ms·
If the scrapers are going to get it anyway, put the data up as a zip somewhere.
by hyperhello 2mo ago
If the scrapers are going to get it anyway, put the data up as a zip somewhere.
- jjgreen 2mo agoIt would not make any difference.
- codemonkey-zeta 2mo agoIndeed, the article mentions Wikipedia experiencing similar scraping pains, even though they already DO have bulk data available.
- HeatrayEnjoyer 2mo agoWho are running these bots? I presume developers at all of the frontier labs know (or at least would know to look for) Wikipedia has bulk APIs for automated access. Unnecessary scraping increases their workload/costs too, so why in 2026 is this still a problem?
- esseph 2mo agoBlack market and gray market data. All the firms want data. All the other firms want data. The banks want data. The other criminals also want data for their crimes and schemes. Oh insurance companies, and the ATS systems. Everybody wants as much data as they can get and they don't care how they get it.
- esseph 2mo agoThis data selling also happens with leaks of all kinds like medical data, often to current or future employers, health insurance companies, etc.
- antisthenes 2mo ago> Read the Docs, a non-profit that hosts documentation for open-source software, who watched a single crawler download 73 terabytes of zipped HTML in one month, costing it over $5,000 in bandwidth From the article. Not the same site, but an example of the same issue.