8 ms·
Understanding the recent DDoS attack against Read the Docs
- tescreal 7d agoI'm curious if anybody could speculate who would be attacking a documentation silo, and to what end?
- SoftTalker 7d agoCould be testing in preparation for attacking something more critical?
- lanyard-textile 7d agoMaybe just for the pleasure of doing it, too.
- davidfischer 7d agoI'm the author of the blog. I don't know. Internally, we were half joking that we were going to get ransom notice, but we never did. The only thing that sort of correlates with this attack is that before it started, we began rolling out some slightly more aggressive rate limits one by one. This was mostly because anytime any new "company" thinks they're going to catchup with Claude/OpenAI, they scrape us very aggressively (and they're not respectful about it). My guess is that the attackers behind this attack were already probing us (they were) and they thought the window of opportunity might be closing.
- SoftTalker 7d agoJust curious, if you're tolerant of scraping, do you make an archive of all your content available so that scraping is unnecessary, and if so do the scrapers prefer that?
- davidfischer 7d agoIt's terabytes of content and other than we're the host not really related to each other. However, for most projects, it's possible to download a zip file of all the HTML docs for that project. We have a lower rate limit to pull these, but a scraper can pull thousands of docs at once. We only host a few hundred thousand projects so pulling a zip of the latest docs for all of them could be done in a day or two at a very reasonable rate. It's also possible to request the docs already processed into markdown[1]. Lastly, basically all of the docs come from Git. A smart scraper could just clone a project's repo. [1] https://docs.readthedocs.com/platform/stable/reference/markdown-for-agents.html https://docs.readthedocs.com/platform/stable/reference/markd...
- simonw 7d agoCurrent evidence is that scrapers mostly aren't nearly considerate or sophisticated enough to take an "archive of all content" option if one exists. See https://people.kernel.org/monsieuricon/creepy-crawlies https://people.kernel.org/monsieuricon/creepy-crawlies which describes how the https://git.kernel.org https://git.kernel.org gets hammered by crawlers all the time even though you could run a single `git clone` and get the data that way instead.
- kees99 7d agoThis is exactly the problem, unfortunately. For somebody who knows a bit how things are set up, or is willing to spend 10 minutes researching, it's a no-brainer that you can just "git clone" entire linux kernel development history, or download entire wikipedia [0]. Alas, large number of scrapers are not willing to spend those 10 minutes, it would appear. So, here we are. [0] https://dumps.wikimedia.org/ https://dumps.wikimedia.org/
- nubinetwork 7d agoYou can tell Claude to clone from github for Linux stuff all you want... it's still going to try web, and fail, before doing what you asked it to do.
- 7d ago
- tescreal 7d agoGood to know. I use your site (with a manual transmission user-agent) often, and it's fantastic. Thanks for your work and the writeup!
- davidfischer 7d agoI've never seen the phrase "manual transmission user-agent". Using your own browser yourself is the new stick shift. Love it.
- gopher_space 7d agoIt feels like everyone's rebuilding their own desktop experience. Kind of Minecraft with folders and text files. The really interesting part of this is how little people talk about what they're doing, and it doesn't feel secretive in any way.
- tescreal 7d agoWhat do you mean exactly?
- Cilvic 7d agoI definitely fall into that, on linux it's just a lot of extensions, scripts etc. Thing is, it's also brittle and not really useful for anybody to talk about it? Not sure I get what you mean with the third sentence.
- gopher_space 6d agoJust like you say. It's not a secret business plan, it just doesn't really feel useful to talk about.
- kkapelon 7d agoEither testing for something bigger OR demonstrating their power to a 3rd party with minimal real disruption
- RobRivera 7d agoCould have been a live-fire exercise by a nation state. Edit: why the down vote? That is literally in the realm of possibility!
- mitxela 6d agoDue to mandatory scheduler maintenance, this test has been replaced with a live-fire CORS designed for military Androids. If you are an iPhone user, please proceed quickly to the chamber lock.
- gkoberger 7d agoI run a similar service, and we get almost daily attacks like this. Sometimes it's a specific high-profile customer, other times it's broader. I can't speak for RTD, but I think it's less "documentation site" and more just that we sit on the domains of high-profile products and the tools are just looking for any hole they can find? Often it's even the company themselves, for whatever reason (security research, etc).
- bennett_dev 7d agoInteresting that the Under Attack Mode wasn’t used at all here. I understand not wanting to break APIs but I feel temporarily challenging non-API usage could have at least helped without impacting users too much?
- davidfischer 7d agoI talked about that directly in the post. We didn't want to just challenge everyone. We use JS challenges but we try to use them sparingly. The rest of the ops team and I were fighting to stay up but it never got so bad that it was a choice between complete outage and using the Under Attack mode.
- account42 7d ago> without impacting users too much? You mean essentially locking out users with non-default Chrome configuration.
- Onavo 7d agoA more interesting question is, what exactly do the attackers gain from hitting read the docs? Most of their docs hosting is static/easily CDN cached. Unlike database bound sites, you would need a lot more traffic to overload pure/mostly static hosting. Maybe it's a malicious AI lab looking to deny their competitors training data? As far as infosec profiling goes, this is probably the oddest case I have heard of. I am thinking it's probably an AI lab that misconfigured their data scraper (made it too agentic) and it ended up looking like a DDoS. The new generation of scrapers are all agentic and self healing. (As an example see YC's https://parse.bot https://parse.bot)
- davidfischer 7d agoAuthor here. This was not a misconfigured data scraper. We see those every week[1]. This attack wasn't scraping useful content. It was almost entirely 404s and 302s and pulled virtually zero real docs. It specifically looked for URLs not served by the CDN and when it found a pattern, did millions of variations of it. Whether built by an AI or not, it was designed to cause outages and financial damage from autoscaling. However, as others have suggested, we may have been a test run for a real target. [1] https://about.readthedocs.com/blog/2024/07/ai-crawlers-abuse/ https://about.readthedocs.com/blog/2024/07/ai-crawlers-abuse...
- Bender 6d agoDo you have a list of the addresses that were hammering your site? Have you tried any of the techniques I list here? [1] Do sets of the IP's show up in here [2]? Are the bots mostly residential, VPS, Tor? What is the HTTP protocol breakdown? HTTP/1.1, 2.0, 3.0? Are they missing any expected client headers? Have you tried blackhole routing any of them from an out of band management console? # only useful if not behind a CDN for Ip in $(cat /dev/shm/list-of-attackers.ipset);do ip route add blackhole "${Ip}" 2>/dev/null;done [Edit] appears you are behind Cloudflare so the blackhole would be up to them. One could still return a 429 or 525 to the attackers. [1] - https://nochan.net/b/Internet-Crap/20260606-How-To-Block-Some-Of-The-Bots/ https://nochan.net/b/Internet-Crap/20260606-How-To-Block-Som... [2] - https://github.com/firehol/blocklist-ipsets/ https://github.com/firehol/blocklist-ipsets/
- gopher_space 7d agoMy naive take on a Cloudflare perspective wants to combine "three times is enemy action" with toddler-speed block dropping and manual clearance. What's the money reason this problem isn't handled at the ISP level?
- toast0 7d agoI don't quite understand your post, but is your question why don't the ISPs of the sources of the abusive traffic sort it out? The distributed nature of DDoS means each participating host isn't sending that much traffic, and there are often tens or hundreds of thousands of participating hosts. An ISP should verify claims of abuse before cutting off customers, and since most of the customers are presumably unaware of what their systems are doing, there will be a lot of unhappy customers and then you've got to spend a lot of customer support time on helping them clean up their systems so they can get back online. I spent a fair amount of time sending out abuse reports for phishing / malware senders about a decade ago, and most abuse reporting addresses are a black hole. Even if you do get to someone who will do something about abuse, they won't do it quickly. There's be a few high profile longer term DDoS attacks lately, but when I was running infra that got a lot of stuff, it was mostly people kicking the tires on DDoS as a service offerings and most attacks were 90 seconds long ... there's no way I'm convincing an ISP to drop a pwned customer over that. Starting from there, this DDoS sounds like layer 7 DDoS which is easy to track to the immediate senders, but a ton of DDoS is volumetric stuff, often volumetric reflection attacks where the senders spoof your address. If you're getting that, best you can do is get the reflectors kicked off (or cleaned up) ... tracing back to the sending hosts means getting a reflector (and their ISPs) engaged to do a lot of labor intensive work. All of that investigation stuff takes qualified people lots of time, that's your money reason it doesn't happen.
- lucb1e 7d ago> most attacks were 90 seconds long ... there's no way I'm convincing an ISP to drop a pwned customer over that. Every victim (such as readthedocs), or even people sharing blocklists to avoid becoming a victim, blocking that ISP's ranges until they do clean up their network could be a convincing argument? As you say, even at 90 seconds, it's clear to all involved parties that the customer is pwned or malicious. Such a reoccurring source of abuse needs to either clean up or find themselves a different ISP to spread harm onto the net I get what you're saying about that this won't solve an ongoing attack right this minute, or even by next week. But if we just let it all happen then the solution is going to be either (1) we all buy equipment that can handle something like a terabit per second and arm our infrastructure to the teeth or (2) centralize all traffic through a vetting entity who decides which client gets to visit the internet today. So far we're headed towards the latter and nobody really wants that. Abuse messages will have to slowly trickle down from victims to originating ISPs to users, and if users didn't willingly sign up, then to wherever users are getting this malware (Google's app store will be a big component). Stopping this at the source seems to me a much more desirable long-term solution
- fn-mote 7d agoThere’s an assumption that turning on Cloudflare’s “under attack” mode would mitigate the attack. Given how adaptive the rest of the attack was, I would be very curious to find out how it would approach that obstacle.
- Symbiote 7d agoThat has only partially mitigated much smaller attacks (residential proxy scraping etc) on my employer's site. We're currently on the "Business" plan, but I'm coming to the conclusion that we need to upgrade to the "Enterprise Advantage" plan for the JA3/4 fingerprinting and detection ID features. I get put off by "Contact Sales" pricing.
- davidfischer 7d agoHere's my take: * JA3s are mostly useless. JA4s supersede them entirely. * Using JA4s in rate limits is pretty useful and helps a lot against proxy scraping. It was not very helpful in this attack. * Bot detections are somewhat helpful but they don't solve scrapers/attacks by themselves. They're useful as a 2nd/3rd data point (eg. low bot score + bot detection + something else)
- cute_boi 7d agoisn't JA4 also useless because it is so easy to spoof tls. For eg. cycletls for nodejs etc..
- davidfischer 7d agoIt will probably be useless one day. In practice, it is still useful today though not for this attack.
- johneth 7d agoThere's also the JA4+ suite (https://github.com/FoxIO-LLC/ja4 https://github.com/FoxIO-LLC/ja4), in addition to standard JA4.
- deleted 7d ago[deleted]
- Surac 7d agoSure it wasn't a AI crawler?
- deleted 7d ago[deleted]
- ACCount37 7d ago> One decision we made is to always give real users an escape hatch. Read the Docs very rarely issues outright blocks or bans to specific IPs or user agents. Instead, our "worst" is a JavaScript challenge, and if a user solves a challenge, they are very unlikely to get challenged again for the next day or so. Finally, a competent response that doesn't leave the users hang out to dry. I'm so tired of seeing incompetents with measures like "blackhole 2 continents" deployed even outside active attacks.
- LoganDark 7d agoYippee, free load test!
- ezekiel68 7d agoIf automatic scaling were free, this would be true...
- Animats 7d agoI'd like to see more of a legal response. First, find out who's on the other end of a few hundred IP addresses. Start with ones in the US. Sue for damages. Use discovery to find out what's on the other end. Sue the maker of that device. If it turns out to be an appliance or smart TV, it may be possible to consolidate cases into one case against the manufacturer. Criminal negligence, tort interference with contract, harassment, Computer Fraud and Abuse act violation... Maybe a restraining order prohibiting the sale of "smart TV" known to be able to host attacks. Have imports seized by Customs and Border Protection. That would get a manufacturer's attention. The manufacturer's EULA will not help the manufacturer, because the plaintiff, the party being attacked, is not a party to the EULA at all.
- nubinetwork 7d agoI'd love to see China and Russia sued for hammering my personal websites for the past 20 years... will it ever happen? Hell no, LOL.
- OptionX 7d agoI agree they they should, but that would be hard before, now in the IoT-hell where even your lightbulbs and internet-facing and capable of being proxies seems like a herculean effort. Pretty sure I saw an article on HN a few days ago about, in part, how a bunch on seemingly innocuous apps for smart tvs, stuff like screen savers and the like, all ran proxy servers (in the users residential address) under the hood. I think it was in the GamerNexus investigation on the whole LG Tv spying on people IIRC.
- kevinbaiv 7d ago[flagged]
- bijowo1676 7d agothis might be a AI driven attack and readthedocs was just a test target. What surprised me is how easy it was to evade the cloudflare defenses. I know it was easy to evade CF, but I would expected CF to do a better job at blocking L7 DDOS. CF is really good in defending against the L4 DDOS, but not L7. this means that cloudflare is really not useful much in the era of Agentic DDOS driven by thousands agents across the globe
- mitxela 6d agoCloudflare is totally worthless actually. It is trivial to evade any part of it.
- rydersel 7d ago[dead]
- PinkSheep 7d agoSorry for the stress that the team had mitigating this attack. Yet I'm always excited to see signs of competence on the attacking side: a targeted application-specific and adaptive attack at L7? Wow! Here's another guess: a DDoS to steal your attention and mask other intrusion attempts. > Defenses need to have broader rate limits across more than just IPs (ASNs, hostnames, etc.). The description/approach seems static and limited? Why not maintain a leaky bucket that counts each request (tickets/points) with higher cost for expensive requests (404, redirects). As the IP's reputation deteriorates (IPv4/32), it begins to spill over to a broader subnet like IPv4/31 then /30 and so on. Fight adaptivity with adaptivity. // maybe I describe something totally obvious, I'm not involved in the web ddos protection side of things. > Attackers actively search for non-cacheable paths (e.g. dynamic redirects, search endpoints, and 404s). To continue with my previous point. As long as individual server's resources permit (memory, socket limit) stall requests before processing them. Low reputation IPs get stalled for longer and the requests that exceed the queue get dropped. The idea is graceful degradation: A good rep IP will not be stalled by sleep(). A poor rep IP will be stalled, but eventually receive its answer instead of some 429/403 (i.e. a user who opened many tabs at once). A bad rep IP will be slowed down by the wait times + rate-limits (queue exceeded) before getting completely banned for good. The point is to have more granularity before throwing errors at random users at the server-level.
- davidfischer 7d ago> The description/approach seems static and limited? Why not maintain a leaky bucket that counts each request (tickets/points) with higher cost for expensive requests (404, redirects). As the IP's reputation deteriorates (IPv4/32), it begins to spill over to a broader subnet like IPv4/31 then /30 and so on. Fight adaptivity with adaptivity. // maybe I describe something totally obvious, I'm not involved in the web ddos protection side of things. The penalty box strategy I described in the post is along these lines. It penalizes excessive expensive requests directly. Specifically, it does add those to a score and will rate limit more broadly as necessary.
- PinkSheep 7d ago> they were overwhelming a hardcoded Nginx redirect (a simple rewrite regex directive) 1. I wonder how much optimization ngx_http_rewrite_module has? Does it precompile the patterns? LLM said yes, this SO answer [1] says that a JIT config option must be on. I consider "Just in Time" to be a half measure when the config itself is static. 2. From looking at NGINX docs, it looks to me there are some pitfalls to writing these rules. Like you must manually make sure to short-circuit the rewrites to exit early? 3. The caveat of regex is that catastrophically backtracking regexes do look simple. I don't see this issue being talked about enough. See links, if you, the reader, haven't heard of it yet. [1] https://stackoverflow.com/questions/59284921/how-much-impact-will-nginx-rewrite-rules-have-on-performance https://stackoverflow.com/questions/59284921/how-much-impact... [3.1] https://joshua.hu/nginx-directives-regex-redos-denial-of-service-vulnerable https://joshua.hu/nginx-directives-regex-redos-denial-of-ser... [3.2] https://en.wikipedia.org/wiki/ReDoS https://en.wikipedia.org/wiki/ReDoS [3.3] https://infrafolks.com/blog/regex-backtracking-devops/ https://infrafolks.com/blog/regex-backtracking-devops/ [3.4] https://www.regular-expressions.info/catastrophic.html https://www.regular-expressions.info/catastrophic.html [3.5] https://gixy.io/plugins/regex_redos/ https://gixy.io/plugins/regex_redos/
- nirmeet011011 7d ago[flagged]