9 ms·
Incident Report: Railway Blocked by Google Cloud [resolved]
Subsequent thread: Incident Report: May 19, 2026 – GCP Account Suspension - https://news.ycombinator.com/item?id=48204770 https://news.ycombinator.com/item?id=48204770
- shevy-java 4mo agoDo not become dependent on Google. Ever.
- cindyllm 4mo ago[dead]
- rekabis 4mo agoTL;DR: putting all your eggs into one basket is bad, man.
- lfx 4mo agoThat’s true, however having only few eggs and shopping for several baskets does not make sense in early days. Not sure how big railway is, but usually you start small with one egg.
- christophilus 4mo agoYou’d think they wouldn’t have started with GCP. There are plenty of datacenters where you can buy racks and racks of servers, and talk to a human when something goes wrong, and even walk in and access your servers. That’s what I’d be using if I were to build a Rackspace today.
- tomschlick 4mo agoThey started on GCP and have been migrating to their own "Metal" DC doing exactly what you're describing. But GCP is still their overflow given how rapidly they are growing and holds some amount of networking that routes to their DC.
- wmf 4mo agoColo is worse than cloud when you're getting started. Sure, you can talk to a person but everything else is much lower quality. People are obsessed with having someone to yell at but yelling does not fix outages.
- iloveplants 4mo agoseems like it's every day
- throwaranay4933 4mo agoThis screenshot from Discord suggests the idea that the outage is caused by automated GCP account ban: https://x.com/acgfbr/status/2056866780866351323 https://x.com/acgfbr/status/2056866780866351323
- Alive-in-2025 4mo agoAutomated account bans are the bane of internet existence today. I was banned from reddit for "bad behavior", I appealed and both times it's oops, there was nothing there, some automated system thought your comment was rude even though it wasn't. Then they send you very strongly worded messages that says trying to work around the ban will lead to something bad happening. I've been worried my main email account provider would do this. The core issue is even if you pay, even if you are a company as shown here companies don't carefully enough have limits on banning. I can only imagine they ban lots of scammy things every day so "they think it's working great".
- faangguyindia 4mo agoGoogle cloud also locked out a Korean Goverment Organization recently. The guy posted on GCP subreddit. Google really need to improve their support team. It's strange such a big corp can't even afford to have proper support team.
- danpalmer 4mo ago> It's strange such a big corp can't even afford to have proper support team Railway say they are in touch with that support team.
- shooker435 4mo agogod help them
- danpalmer 4mo agoI had good experiences with their support, and bad experiences with AWS support. tldr: YMMV.
- King-Aaron 4mo ago> It's strange such a big corp can't even afford to have proper support team This seems to be by design.
- ndneighbor 4mo agoWe have a CSM, Head of Customer Support contact, and further contacts with GCP. Despite that, we still had this issue.
- add-sub-mul-div 4mo agoAutomating support, automating everything is the key to their whole deal. Tech giants leapfrogged the rest of the economy by innovating a company that can scale its customers without having to scale itself proportionally.
- 4mo ago
- mcontrerazCL 4mo agoall my fkn postgres bd in railways! what do i do now?
- cactusplant7374 4mo agoTake a walk. Breathe in the fresh air. It feels good.
- eoswald 4mo agoHahah at least you're not getting called every five minutes because you cant shut off the alerts, because its apparently deployed SOMEWHERE but good luck finding how to access it. Can't wait to see the bill from Twilio because of this lol
- gnabgib 4mo agoDupe - join the discussion started an hour ago instead of query string work (12 points, 4 comments) https://news.ycombinator.com/item?id=48200827 https://news.ycombinator.com/item?id=48200827
- aarondf 4mo agoI added the qs because it defaulted to a story from 3 months ago.
- ryanisnan 4mo agoYikes. I was wondering why my TLS certs were coming up as invalid.
- eoswald 4mo agoSorry, I have a hard time blaming Google for this, when Railway seems to be having increasing trouble keeping the platform stable. Something like this should NOT take down an ENTIRE service. There should be a backup when literally your business is about being the reliable backend. This just seems like poor planning to me.
- cactusplant7374 4mo agoDisaster recovery is pretty expensive, right? Especially for their size.
- ryanisnan 4mo agoI don't quite know what you mean. Do you really expect Railway to use a multi-cloud architecture to host all of their client's projects? I suspect that would lead to a lower availability, all things considered.
- deleted 4mo ago[deleted]
- impulser_ 4mo agoThey literally own their own data centers. That's whats surprising about this. They are lying to their customers when they say they operate their own data center because obviously they don't if everyone's apps are down with GCP blocking their account.
- ryanisnan 4mo agoOh, I see what you mean. Eh, it's possibly the same reason that AWS essentially goes down when us-east-1 goes down.
- brookst 4mo agoIs it not possible that they own their own data center and have an unfortunate Google dependency? Obviously a fiasco but I’m not prepared to call them liars when it could be an honest mistake.
- brokenodo 4mo agoI’m a new customer and have been falling in love with Railway over the last 2 weeks, but this is quite the wake up call.
- csw-001 4mo agoLiterally in the same boat. I've been really happy with it, but this is a major eye opener.... It's been done for a looooong time by provider standards.
- reelvideocap 4mo agosame
- TheAtomic 4mo agosame same
- choilive 4mo agoBeen a customer with them for over a year now, small incidents here and there but never anything this major.
- Mengkudulangsat 4mo agoThat explains why all my vibe-coded hobby projects are down. Thank God I'm not dealing with any public-facing sites! Would have been an expensive lesson for a newbie coder if my job depended on this.
- rekabis 4mo agoTL;DR: putting all your eggs into one basket is bad, man.
- canpan 4mo agoHow to handle domains? The rest is easy, but your domain registrar blocking you sounds like a pain. My current solution is to use a local small provider, just for the domain. Then if there is a problem with your play account it is out of any blast radius.
- FlamingMoe 4mo agoWhat do you mean by local small provider? A registrar on main street?
- truekonrads 4mo agoMarkMonitor
- Barbing 4mo agoAny changes since acquisition? Looks like they were sold at the beginning of the year to a company without a Wikipedia page whose parent company doesn’t have one either https://en.wikipedia.org/wiki/Markmonitor https://en.wikipedia.org/wiki/Markmonitor Acquired in November 2022 by Newfold Digital, it was later announced that the firm would be sold to Com Laude, a company owned by PX3 Partners. - Edit-Private equity apparently https://px3partners.com https://px3partners.com PX3 stands for purpose, passion, and performance. It is a pan-European private equity firm with headquarters in London. It invests behind transformative themes and targets companies operating within select segments of the business services, consumer and leisure, and industrials sectors with strong business fundamentals.
- rekabis 4mo agoWhat the deuce are you blathering on about. An account got blocked, this has nothing to do with a domain. And I’m talking about having disparate failovers that don’t rely on a single hosting provider. At that point, who cares what Google does to your cloud account… work with the hot failover and spin up another hot failover somewhere else.
- bshack0 4mo agoso....what are we switching to y'all? cloud-run ? ;P
- auxiliarymoose 4mo agofederated hardware (a bunch of raspberry pis networked into a high availability kubernetes cluster, hidden across various local coffee shops for free power and bandwidth)
- throwatdem12311 4mo agoraspberry-pi cluster in my closet
- frio 4mo ago16GiB Raspberry Pi 5s in my country are now going for ~$450USD, so I've gotta say that's out of reach for me now :(.
- enahs-sf 4mo agoI respect what railway is doing but also would never run my business on such a platform.
- dpark 4mo agoThat kind of sounds like you don’t respect what they are doing.
- enahs-sf 4mo agoI think it’s good people are making IaaS platforms, but have dealt with enough firefighter hero bullshit to have seen this coming a mile away. Uptime and redundancy are strongly correlated.
- eoswald 4mo agoToday changed my opinion on them completely. Was willing to give them the benefit of the doubt that they're growing fast, but now seeing that they've failed to scale properly, and are missing little things that become big things later. I can't take that risk.
- fjni 4mo agoWait… railway runs on GCP? Didn’t they make a whole thing about not “building a cloud on top of another cloud?” Or did they just mean that they’re not renting VPSs but only metal from the cloud provider? In my mind I was so excited that there was another provider not just paying one of the hyperscalars but at a minimum colocating and owning more of their stack. https://blog.railway.com/p/heroku-walked-railway-run https://blog.railway.com/p/heroku-walked-railway-run
- eoswald 4mo agoYep, and this is why I'm pissed. They lied. They're completely dependent on GCP. So, I gotta do some research, i need something a little more stable (and less dependent on one company's whims) than this. This is bad for them, because it really strikes at the heart of their 'big claim,' peacefull software deployments. This is chaos.
- ndneighbor 4mo agoYea, I mean, that's the whole MO of our platform and we failed at that. So yea, that's disappointing and more so for our customers. I can provide an explanation about the GCP dependency. Yes, we have host workloads off GCP, and we have been able to build a good business by performing a cloud exit. However, we were worried that we would have a circular dependency on our own cloud. I don't think we expected to get auto-modded out of our own account, hence we left our DB on CloudSQL. It was never our intent to deceive people that we didn't own our own destiny with our business. The last GCP issue, we were assured that this scenario wouldn't happen (when we got auto-ratelimited, which was bad, but survivable) - but it seems like we have further work to do. Apologies.
- fontain 4mo agoI’m very sympathetic and understand that decisions are easy to criticize in hindsight but leaving your database in GCP while moving everything else to your own data centres seems so backwards I can’t even begin to imagine how that could happen. Was this really an intentional design decision?
- Avicebron 4mo agoIsn't Railway the "the API key to delete the backups is in the prod database, because that's where the backups live duh" guys?
- trvz 4mo agoNo, this is the company that failed those guys. You should also read the story, as you're perpetuating a false version of it: https://x.com/lifeof_jer/status/2048103471019434248 https://x.com/lifeof_jer/status/2048103471019434248
- dangoodmanUT 4mo agoIt has been 0 days since GCP has taken down a startup (again). You see this at least once a year. Never heard of this from AWS or Azure. In all seriousness, this is why we don't use them. They have the most ergonomic cloud of the big three, then absolutely murder it by having this kind of reputation.
- tjpnz 4mo agoAWS normally contacts you first.
- cherioo 4mo agoThey better do. What is google doing?
- Gigachad 4mo agoIt's all AI powered
- kevin_nisbet 4mo agoDo they? The only anecdotal thing I've seen is we hired a vendor to do a pentest a few years ago, and they setup some stuff in an AWS account and that account got totally yeeted out of existence by AWS if memory serves.
- alchemism 4mo agoI’m fairly certain you are supposed to contact any vendor before attempting to penetrate hosts with authorization, not the other way around.
- coredog64 4mo agoHaving done this for both Azure and AWS, there's a specific ticket that needs to be filed with each provider that documents the scope of your pen test, where you're coming from, and a time frame over which you're doing it (which ISTR was "not more than 24 hours")
- TheTaytay 4mo agoI’ve seen a few smug “all your eggs in one basket” comments here. I’m aware of some companies hosting their own metal and infra, but I’m not aware of large companies mitigating risk by hosting on separate cloud providers as a fallback mechanism. We might disagree with cloud provider choice, or think they should have been hosting their own metal, but that’s still an “all your eggs in one basket” choice, right? Heck, they might even have multi-region fallback with GCP, but if GCP bans your account, that doesn’t matter. Are there good examples of running a company of railway’s size so redundantly that their host could nuke one of their accounts and they’d just keep on trucking?
- fontain 4mo agoThey do run their own metal. That’s their entire ethos. Railway is their own cloud.
- chradams 4mo agoJust google multi-cloud. Yes. It's a thing.
- wmf 4mo ago99% of multi-cloud is fake though. True multi-cloud is incredibly rare.
- TheTaytay 4mo agoI appreciate it. That's my belief as well. Very easy to write a post like, "Just use multiple clouds!" or to claim to have done it with a small project. But it's hard for me to imagine the benefits outweighing the extremely massive complexity costs at a certain scale.
- upnorthmedia 4mo ago[dead]
- rvz 4mo agoLet me guess… Googler running AI agent in production that blocked this startup’s account.
- dwa3592 4mo agoWait, I thought railway was a cloud provider like AWS, GCP but better and more agile. At least that's the impression i got from their website.
- whh 4mo agoThis could kill a startup. I really don't like Google's automated and silent account murder functionality.
- MrDarcy 4mo agoThere’s no way this was automated or silent. The only reasonable explanation is Railway lost control of their estate and something was happening that warranted a group of humans to decide flipping the kill switch was the best of a set of bad alternatives.
- macintux 4mo agoYou’re giving Google far more credit than they’ve earned.
- whh 4mo agoIt's almost certainly one of those Android Store related checks or YouTube account checks. It's why it's best to disable login for the services you don't want your staff messing with on Google Workspace.
- faangguyindia 4mo agoyou can go on google cloud subreddit and watch horror stories i actually built a good plan out of those horror stories for my companies.
- codegeek 4mo agoThis is bad. Even their own website is down at railway.com. Looks like total dependency on google cloud. Surprising for a company of their scale with all this VC money.
- choilive 4mo agoThey run a decent amount of their own compute/bare metal server for customer workloads. But likely still had some critical dependencies on GCP.
- cube00 4mo ago> Surprising for a company of their scale with all this VC money. Not sure too many VCs would be cool with deep redundancy when there's more features to build to bring in more customers instead.
- rmeara 4mo agoGoogle has a total dependency on it's own infra and does fine. Why do its customers need multicloud? Huge PITA unless you need an absurd number of 9s
- tux 4mo agoAt this point you can’t trust Google anymore, it keeps breaking things. Imagine having Google AI do this thins automatically. Will have apocalypse in in a day.
- deleted 4mo ago[deleted]
- Drew-Aetherwave 4mo agoIt is killing me...
- isninkhamiss 4mo agogithub got way more noise for less
- deleted 4mo ago[deleted]
- Osborn_Ojure 4mo agocompute recovered, get ready boys!
- r_lee 4mo agoseriously, is it possible to trust GCP with critical data/services at this point if you're not a billion dollar company? I'm exaggerating but someone said they got "auto banned" what if that happens to a small account which hosts some really important data/services there?
- Avicebron 4mo ago> what if that happens to a small account which hosts some really important data/services there? Pray to @dang that you will make the front page of HN?
- throwaway85825 4mo agoEven if you are a billion dollar company you still have problems like the Australian pension did. Google is just that bad.
- ttoinou 4mo agoRailway isnt far from being a billion dollar company, no ?
- intelVISA 4mo agoI don't want to believe this, lol.
- xyzzy_plugh 4mo agoI've managed several accounts with GCP over the years and I've always maintained a great relationship with our contacts there. Some of these accounts were quite small, on the order of <$20k/mo, and even then we were kept abreast of anything that might be cause for concern. I always maintain a standing biweekly meeting with at least someone on the other side (account exec, technical staff, whatever) and I've yet to be blindsided by anything. Is Google's communication good? No, not particularly. The only way something like TFA happens is if the relationship is neglected (by one or both parties). I'm not saying Railway did something wrong, but there are usually many flags and opportunities to correct long before drastic actions. I get the impression that Railway plays fast and loose with a lot of their limits and resources and that Google may not be a fan of that. Edit: would also like to say that if you put all your resources in one GCP project you are going to have a bad time. If you organize stuff over many projects it is very unlikely that they will ever take account wide action. I've had issues with, for example, a particular tenant's behavior, but it never jeopardized the other tenants.
- ChrisArchitect 4mo agoEarlier: https://news.ycombinator.com/item?id=48200827 https://news.ycombinator.com/item?id=48200827
- upnorthmedia 4mo ago[dead]
- jefborges 4mo agoRailway is back, but I’m not sure if I can trust keeping my projects there, so I’m going to migrate to another company.
- oofbey 4mo agoAfter reading about how their delete database API also deletes all the backups, I concluded they are not to be trusted.
- CodesInChaos 4mo agoDon't all major clouds do that by default? But at least they have additional protections you can configure, if you know about them.
- marknutter 4mo agoIt's not back.
- deleted 4mo ago[deleted]
- orliesaurus 4mo agoI wonder if someone has exploited a weird Google-safety automated process to report something on Railway which caused Google to block the whole thing.
- padolsey 4mo agoDoes anyone know how this even happens inside the walls of google? Is it an automated process? How is such a (presumably) high revenue account just magically blocked without human intervention? I'm quite perplexed.
- jpollock 4mo agoThere would have been efforts to contact them, but it would have been via their contact method, aka the email they set it up with. Common ways this happens? They are using a credit card to run their business with no backup payment method. Then the company's contact person is on vacation. Sign up for terms. It will get you payment terms!
- scratchyone 4mo agoHonestly still insane to nuke a high-volume client's business after a single payment issue. There would be no reason for Google to believe that a single hiccup like that is evidence that they won't get paid and have to cut account access immediately.
- antran22 4mo agoRailway might not be even in the realm of high-volume clients for Google. For all we know they might be efficient in utilizing Google infrastructure. But most likely, it's just automations in place without an appropriate human override coupled with gross negligence.
- 4lx87 4mo agoIt is insane, but my past experience with GCP is they suspended all service only days after a failed payment, after years of paying on time. It's a major factor in why I don't use them anymore. I'm not waking up to angry customers again because the CC is expired and I missed an email. I'd be curious to know why Railway's account was suspended. Was it a similar payment issue or something else?
- jpollock 4mo ago
- binarycleric 4mo agoHow the heck do these things happen, especially with companies with huge monthly spend? At my last job we had some suspicious workloads running on AWS and our TAM reached out to us before taking any action. Who wants to bet this was some AI automation gone wrong and because GCP seems to be allergic to actually contacting a human to get a response, this just sits in some support queue that outsourced workers look at after a few hours just to give a canned response?
- garciasn 4mo agoNothing surprises me with anything related to support on GCP. While we absolutely do not need them, I have been through no less than 12 different Account Executives over the last 6y and they're all ENTIRELY and COMPLETELY useless. They all introduce themselves, beg me to setup a meeting w/them and some sort of engineering resource(s), and they come to a meeting with a canned slide deck that is so absurdly unrelated to us that I just laugh, and then the next time I hear from them it's because we have a new AE. This is my most recent reply (right after Next '26): > I really appreciate you reaching out; however, we have met with, I dunno at this point, more than a dozen GCP Account reps, execs, technical teams, etc over the years and there's little to no value for us or you, now or in the future. Please do feel free to invest your time on your other clients. We're good; truly. I love GCP and its services; we have been very pleased with it over the years, but the human side of it? Fucking sucks and I just don't see why they even bother.
- OptionOfT 4mo agoIt's because they're measured on something, unsure which metric, but it's definitely not how helpful they are to you.
- YuriNiyazov 4mo agoDon't know about GCP, but our AE on AWS was also continuously rotating, and as best I can tell, their job was to figure out what we are planning to build, and to ensure that we should always use <INSERT AWS SERVICE DU JOUR> for that, rather than a competitor product or build it ourselves.
- bearjaws 4mo agoI will never leverage GCP in an enterprise setting, it's honestly amazing how hard they fumble the bag. Will be interesting to see when GCP support started working with them, from the updates there was an hour and change from when they identified the issue and GCP support was confirmed. In the cloud space it seems like AWS does nothing and wins.
- UrbanNorminal 4mo agoIs google allergic to humans or something? Cannot they just send an email or call the company before taking a wrecking ball to the entire company's infra? Are they stupid?
- lateral_cloud 4mo agoKeep the pitchforks at bay for now. No one knows what actually happened yet and we are only seeing one side of this outage.
- BarryMilo 4mo agoSurely this is automated. They wouldn't waste precious dollars on employing humans just to keep other humans happy.
- snypher 4mo agoIt surprises me there's not a manual review for $$$$ accounts. Speculation at this stage, but it's weird they would be put in the Recycle Bin like that.
- BitWiseVibe 4mo agoAs someone who runs some public APIs, the amount of spam from Railway IPs is insane. They have horrible abuse prevention. Hopefully this encourages them to improve their operations.
- nikcub 4mo agoThis is the conflict at the center of running a hosting company - make it easy to signup and you get a lot of new users but also a lot of abuse. Implement anti-abuse measures and you will hit some loud false positives (this may be the case with GCP here). I don't envy anybody running a hosting co - the internet is a really ugly place under the surface. edit: to add - AWS are really good here. Must be the ~30 years of retail fraud and abuse experience.
- edelbitter 4mo agoI continue to receive phishing via AWS pretending to be Amazon. And not even the Unicode-lookalike shenanigans that my spam filter refuses for excessive mixed scripts, no; literally claiming to be Amazon as in: the company that operates the relay.
- swyx 4mo agoi wonder if DID or World (various ways of Proof of Human) can help solve this issue.
- nikcub 4mo agoThis just incentivizes market for bio-mules, which already exists with world[0] - where prices stay low because it was rolled out to low-income countries. Then there's the platform game theory. If you adopt you add friction which reduces signups, and there will always be a competitor who would risk the 10x fraud increase in order to capture 100x the market. Railway has seen hyper-growth because it's so easy to run from, and is recommended by, coding agents[1]. The solutions are here already just not well implemented or understood - probabilistic fraud detection, resource limits, service and automation limits, standard gov identity verification as a signal, enterprise sales channels with human relationships, etc. There are tradeoffs with each platform choice that just aren't well understood. Most users shop on price and DX and don't see the abuse infra or problem until it hits them. Google and GCP have a problem where they completely cook users who get flagged in their automated fraud net (this isn't news - or shouldn't be) [0] https://www.coindesk.com/policy/2023/05/24/black-market-for-worldcoin-credentials-pops-up-in-china https://www.coindesk.com/policy/2023/05/24/black-market-for-... [1] and the problems that come with providing that simple interface, like sometimes dropping prod
- redanddead 4mo agoone of the many reasons companies are cloud agnostic and dont want to get locked in
- fh67 4mo agoYeah but until you find that the new cloud provider won't approve your compute quota or doesn't have enough capacity in the region or you hit fraud flags for stagnant account spinning up lots of compute.
- parineum 4mo agoThere's a lot of, what seems to me, unfounded blame being directed at Google for this. Isn't railway the company that just blamed Anthropic for deleting their prod database?
- mmmore 4mo agoNope, Railway was the company who was hosting PocketOS, which is the company that blamed Cursor for deleting their prod database. Railway is only involved insofar as their API allowed an instant delete of the prod database.
- oofbey 4mo agoRailway deserves a lot of blame here. Deleting backups along with the database is a lot like not having backups. Moronic design choice.
- Genego 4mo agoWhy does Railway deserve any blame here at all? It was an MCP with elevated infra access, that the user willingly connected through Cursor, which allowed an LLM Agent to manage infra on Railway. The user would first have gone through oAuth confirming the access level scope (I would have rejected the moment it indicates to me that it can delete critical infra and backups...). So obviously it has access to all commands the user would also have access to. From my perspective the blame is entirely on the user, and partly on Cursor for not enforcing HITL correctly across their agents.
- wmf 4mo agoPutting AI aside, people make mistakes. One of the most common mistakes people make is deleting the wrong thing. After they realize the mistake, people want to restore the thing they deleted from backups. Thus deleting the thing and deleting the backups of the thing should always be separate operations.
- 4mo ago
- brokenodo 4mo agoWell, as a 2 week tenured and very happy Railway customer until now, I am now a Render customer. Somehow DNS cut over within 1 min(!) and live after about 30 minutes of work. Not bad!
- DrewADesign 4mo agoIn my experience, DNS changes are a lot faster than they used to be. There’s some website that has a map that tries to resolve your domain with a bunch of name servers around the world that was pretty neat to look at last time I migrated something.
- nbarbettini 4mo agoI became so conditioned to waiting hours(!) for DNS propagation that I'm always pleasantly surprised when it takes <5 min these days.
- DrewADesign 4mo agoYeah way back in the day I’d be used to waiting overnight
- twostorytower 4mo agoI love pointing my name servers to Cloudflare so any DNS changes from that are practically instant.
- swyx 4mo agoas with many things, we say we like decentralization but quietly vote for centralization
- unit490 4mo ago[dead]
- usernametaken29 4mo agoI didn’t knew Railway so with this misleading headline I thought a Google Cloud data centre was being built in the way of a railroad. That’d been a funny story to read..
- astafrig 4mo agoHow is the title misleading?
- tauntz 4mo ago"Railway Blocked by Google Cloud" If you don't happen to know that "Railway" is referring to a company, then you might reasonably read that as "a GCP outage caused issues in the train network somewhere".
- Polizeiposaune 4mo agoAn elevated railroad once ran through one end of what is now a Google-owned building (Chelsea Market in Manhattan). It's now part of the High Line elevated pedestrian park.
- hnburnsy 4mo agoFrom their founder on X... "Absolutely. The Railway network is a mesh ring between AWS, GCP, and Metal So: - High availability interconnects - High availability path routing between clouds - Database itself is high availability However, Google's VPC itself is not. So we will add a shard to Metal and AWS"
- hnburnsy 4mo agoMore here... https://x.com/JustJake https://x.com/JustJake
- sammy2255 4mo agoThe 3-2-1 backup rule is pretty outdated in the world of cloud. You could have 3 complete copies of your data in different S3 buckets, but if they're all under the same account you've lost your blast radius protection
- rsync 4mo agoIf only there were a quick and easy way to replicate s3 buckets to an independent provider… … on the Unix command line … … to a cloud older than AWS… … if only …
- eclipticplane 4mo agoI don't think that technology exists. Sorry.
- oefrha 4mo agoWell having backups help, but I certainly can’t migrate my infra to rsync.net on moments’ notice (or ever since rsync.net does storage and nothing else) so my customers aren’t affected.
- funtech 4mo agoWish I could upvote this comment account more. Too many people look for something new and shiny when trusty ol tools are sitting right there. :)
- lemagedurage 4mo agoInflated egress costs might make this prohibitively expensive, $80 per TB at GCP and AWS
- whalesalad 4mo agoYou replicate data to different clouds.
- zootboy 4mo agoIt's not outdated, you just actually need to follow it. 3 copies of data in separate S3 buckets is ignoring the "2" in the 3-2-1 rule: 2 different mediums, and also the "1" rule: 1 copy offsite. In the cloud era, offsite means not on the same cloud provider. Different mediums ideally means a non-cloud provider (e.g. a NAS at your office under your control).
- mjy78 4mo agoAll in on cloud so we don’t need to worry about backups. Now your subscription is the single point of failure.
- chatmasta 4mo agoI thought Railway was building their own data centers? [0] > The fact of the matter is, you simply cannot build a cloud on someone else’s cloud. Indeed… [0] https://blog.railway.com/p/launch-week-02-welcome https://blog.railway.com/p/launch-week-02-welcome
- QuinnyPig 4mo agoVercel seems to be pulling it off. So does PlanetScale, albeit for databases only. But everything’s a database.
- jujube3 4mo agoIf you buy a cloud-on-a-cloud, you're a clown-on-a-clown.
- zelon88 4mo agoWild to me that any tech sector business would want to rent an operating environment to park their entire infrastructure into. This is the equivalent to traveling shoe salesmen setting up a tent in the parking lot of a strip mall.
- fnord77 4mo agowish I knew what "railway" is
- leventhan 4mo agoWhat's a good alternative to Railway?
- valgaze 4mo agoMay 2024 UniSuper incident: https://cloud.google.com/blog/products/infrastructure/details-of-google-cloud-gcve-incident https://cloud.google.com/blog/products/infrastructure/detail... https://www.unisuper.com.au/about-us/media-centre/2024/a-joint-statement-from-unisuper-and-google-cloud https://www.unisuper.com.au/about-us/media-centre/2024/a-joi... A joint statement from UniSuper CEO Peter Chun and Google Cloud CEO Thomas Kurian 8 May 2024 UniSuper and Google Cloud understand the disruption to services experienced by members has been extremely frustrating and disappointing. We extend our sincere apologies to all members. While supporting UniSuper to bring its systems back online, Google Cloud has been conducting a root cause analysis. Thomas Kurian has confirmed that the disruption arose from an unprecedented sequence of events, where an inadvertent misconfiguration during provisioning of UniSuper’s Private Cloud services ultimately resulted in the deletion of UniSuper’s Private Cloud subscription. This is described as an isolated, “one-of-a-kind occurrence” that has never before occurred with any Google Cloud client globally. This should not have happened. Google Cloud has identified the sequence of events and taken measures to ensure it does not happen again. Why did the outage last so long? UniSuper had duplication across two geographies as protection against outages and data loss. However, the deletion of the Private Cloud subscription triggered deletion across both geographies. Restoring the Private Cloud required significant coordination and effort between UniSuper and Google Cloud, including recovery of hundreds of virtual machines, databases, and applications.
- kvakvs 4mo agoThe instant cascading worldwide deletion upon closing or deleting a subscription sounds like a recipe for disaster. Why not mark it for deletion and delete say... a day or a week later?
- modernpacifist 4mo agoEither mark-for-delete has the same impact as deleting in terms of shooting all the Cloud resources associated with the subscription, at which point the outage still happens but maybe the recovery is smoother or you've just delayed the inevitable by a week because no one will look at it unless there is actual impact.
- eezing 4mo ago“Deletion of private cloud subscription…” Who deleted it?
- codepack 4mo ago[dead]
- codepack 4mo ago[dead]
- codepack 4mo ago[dead]
- bilalq 4mo agoBuilding a startup on GCP (or even Google Workspace) is an existential risk.
- codepack 4mo ago[dead]
- koolhead17 4mo agoLet's blame some rouge AI agent at GCP causing this.
- htrp 4mo ago[dead]
- deleted 4mo ago[deleted]
- dlcarrier 4mo agoThis is the kind of outage worthy of a Kevin Fang video.
- thrownthatway 4mo agoHuh. Railway dot com Has nothing to do with railways. I wish software people would get their own words.
- patrickmay 4mo agoI was also expecting a story about a physical railway being shut down.
- jaspanglia 4mo agoCloud platform dependencies are becoming a huge single point of failure
- steve1977 4mo agoLesson learned: don't rely on a single hyperscaler, even (or especially) as a startup.
- burnerRhodov3 4mo agoI just... I don't really understand why startups even use AWS, GC, or any other cloud hosted software? Hetzner, etc. Are all extremely cheap, and honestly scale so well... Code nowadays is cheaper for configs, and having full control over your compute is... liberating.
- steve1977 4mo agoOh absolutely... and many use architectures that have evolved out of the needs of really big companies and are not really a good fit for a startup. But I guess they want to be "ready for growth".
- dannersy 4mo agoLow cost to entry, easy to get scale from the beginning if you need it. The large cloud providers throw free credit at startups to lock them in all the time. I had a short lived stint trying to get my own startup off the ground and it was really easy to get free compute from Google with no strings attached. This was many years ago now, but I would be surprised if it is any different. I am with you entirely and would not have taken that route today, but it is really easy to see why people go that route.
- antran22 4mo agoA few years ago, when I was kinda active in the startup scene in my area, you have people selling access to cloud credits with penny-on-the-dollar price. The credits are given out liberally to big-corps, organization by AWS/GCP, through workshops, webinars, events. All in the hope of roping the departments into building MVPs, demos on AWS/GCP, but people also find a way to cheat on that system and make some quick bucks. I know a startup of my acquaintances that have been running on AWS for 5 years straight without paying a single dollar to AWS. When the credits almost run out, they started to migrate their data over to another account with credit. That happened twice already. It helps to have a portable, replicable IaC config. But also this is sustainable because they are a pretty small struggling shop. You will probably not be able to do this if you are trying to maintain more than 3 nines for an enterprise client.
- pavelevst 4mo agoAvoid vendor locking, have backups, make disaster recovery standby (or plan for quick recovery elsewhere)
- tardwrangler 4mo agoEveryone is eager to point a finger at Google, but I've been a user of Railway for a while now, and I've seen enough nonsense to want to hear what GCP has to say about this before drawing any conclusions. Let's just say Railway has had problems like this before, and the way their team handles them does not inspire any confidence. Regardless of how it happened, for me, this is the straw that broke the camel's back.
- prathamtharwani 4mo agoCould you point us to any specific past instances? I'd be interested to read about them.
- jeffreyq 4mo agohttps://blog.railway.com/p/incident-report-february-11-2026 https://blog.railway.com/p/incident-report-february-11-2026
- x0x0 4mo ago"we did not have the monitoring or controls to prevent our anti-fraud from hard killing 3% of workloads, including many instances of pg" Oof.
- rzmmm 4mo agoNeeds an anti-anti-fraud service which terminates malfunctioning anti-fraud services.
- x0x0 4mo agoWhen I've written similar services, there was a (low) hard cap on how many fraud decisions they could action before they quit and paged. If we were getting hit with a wave of something, a human had to temporarily bump that limit.
- locknitpicker 4mo ago
- jamwise 4mo agoThere goes a 9
- ksajadi 4mo agoWhen you signup for Railway, they have uncommon way of making sure you have read and understood their T&C regarding abuse of their systems, including crypto mining, etc. My guess is that many are abusing their free tier, causing them trouble with their service providers. I take no joy in seeing Railway take a hit like this, even as a competitor, but free compute attracts all sorts of strange users. We've been there and decided early on to avoid free compute even it costs us our top of the funnel.
- WhereIsTheTruth 4mo agoWhen your cloud depends on an other cloud All these companies are fraud
- cube00 4mo agoRailway "What we know so far: May 19th 2026": https://station.railway.com/community/what-we-know-so-far-may-19th-2026-86354cdd https://station.railway.com/community/what-we-know-so-far-ma...
- mattbee 4mo agoThe risk of an "upstream cloud provider" is not something you need to tolerate in your supplier of internet infrastructure!
- zx8080 4mo agoFor those who opened this link to read news about the real railway (with trains), it's not about it. Thank you for wasting my time!
- r721 4mo ago>We have resolved this incident and a post mortem is available here. >https://blog.railway.com/p/incident-report-may-19-2026-gcp-account-outage https://blog.railway.com/p/incident-report-may-19-2026-gcp-a... >May 20, 07:57 UTC https://status.railway.com/incident/I23M92U0 https://status.railway.com/incident/I23M92U0
- sschueller 4mo agoIt should be possible to sue Google for damages in such cases. This isnt a network outage or service failure which I would consider part of ToS.
- VladVladikoff 4mo agoWhat if the reason for their stuff being shut down was a payment issue like an expired credit card or maxed credit account? Unless I missed it skim reading their post I don’t see any information anywhere about their communications with Google.
- bastawhiz 4mo agoIf you have an account manager and a contract, there's zero excuse for automated suspension. That's literally the whole point of having a dedicated person. From the report: > May 19, 22:22 UTC - P0 ticket filed with Google Cloud. Railway's GCP account manager engaged directly.
- Cthulhu_ 4mo agoIt's always possible to sue, but Google has good terms of service and lawyers - I'm 99% confident that a lawsuit would end up nowhere.
- bastawhiz 4mo agoThey have every right to sue, and if they did sue they almost certainly would win. This is clear breach of contract. The only argument Google could make is "they did something to violate our agreement" but they'd have to prove that, and then have a damn good explanation for why they were in the right to suspend the account without any outreach. Unless Railway did something egregious, Google clearly made an error. But that's not what will happen. Google will offer an apology (perhaps even a public one), a giant pile of account credit, and a pinky promise not to do it again. Railway will accept it and hmmm and haw internally about whether to decrease their reliance on GCP, and then when they calculate the cost of going in on other clouds more heavily (or their own metal), they'll just think harder about weird failure modes.
- brunooliv 4mo agoHaving tried many of these hosting services to host/play with toy apps, DigitalOcean and Fly.io are both unparalleled GOATs.
- paganel 4mo agoApparently this has nothing to do with real-world trains and to the real-world rail system, at first, and reading the title alone, I had thought that some trains might have got stuck somewhere because of an IT (google cloud) failure. It's just another SaaS story.
- jkogara 4mo agoInterestingly, upon logging in this morning I was presented with a new terms and conditions banner that required me to agree to not deploy a list of, to varying degrees, nefarious things (bots, torrents, "anything illegal", etc.). Is it likely that some of these workloads resulted in the auto restriction from GCP?
- yomismoaqui 4mo agoRemember, the cloud is someone else's computer. If that person turns it off you're screwed.
- whh 4mo agoThere's that "automated action" again. Regardless of the architectural decision, it makes me incredibly uneasy relying on GCP if these types of things can happen.
- danpalmer 4mo ago7 minutes from bug filing to account restoration. This shouldn't have happened in the first place, but that's an excellent response time from the support team.
- AbstractH24 4mo agoDid anyone discover any unexpected tools/wesbties use railway during this outage?
- ryanshrott 4mo ago[flagged]
- mooktakim 4mo agoSo railway who provides cloud hosting uses another cloud hosting provider for vital services. Why use railway and not Google cloud directly?
- ryanshrott 4mo ago[flagged]
- daohieu91 4mo agoSmall Railway customer here ($5-10/mo, Spring Boot service). My service stayed up the whole time on May 19 — I only learned about the incident from this post-mortem. "Detailed post-mortem about an incident the customer didn't even notice" is the strongest trust signal Railway could have given. The broader concern though isn't Railway-specific: GCP's "we can suspend a downstream provider with no warning" policy is a structural risk for every IaaS-on-IaaS layer (Fly, Render, Northflank, etc.). Curious whether anyone here has seen contractual or technical mitigations beyond "have a second deploy target ready to switch to."