7 ms·
No tests? Just mess up some mundane detail [1] and voila! Wake-up calls and heart attacks for 100,000s of administrators? 1: "Oh, well, this is not a mundane d
by 01284a7e 2mo ago
No tests? Just mess up some mundane detail [1] and voila! Wake-up calls and heart attacks for 100,000s of administrators?
1: "Oh, well, this is not a mundane detail, Michael!" https://www.youtube.com/watch?v=3fGHaVn5rGo https://www.youtube.com/watch?v=3fGHaVn5rGo
- akdev1l 2mo agoNot even tests but just some basic anomaly detection lol. Like maybe if the bill amounts increase by like 10M% there should be someone that looks into it
- qurren 2mo agoYou overestimate how much people give shits at big techs like Amazon. When literally everything is driven with sticks instead of carrots, the work culture does not invite employees to proactively care about product quality. You'd be better off letting the heart attacks happen and take the 3am on-call and be the hero instead. It would be good promo doc material, and being a hero is extremely good insurance against getting kicked out of the country (via the PIP->H1B grace period expiry mechanism).
- tyre 2mo agoAre you speaking from experience or simply making things up? I know a fair number of former AWS engineers and managers. None of them think like this.
- geodel 2mo ago"Former" seems to an important detail here.
- gleenn 2mo agoIf someone quits their job, do all their opinions suddenly become suspect? You're kind of damned-if-you-do-damned-if-you-don't. Either you work for the company and you are biased one way, or you quit and now your bias is now suddenly the other way. I've joined and quit many jobs and my opinion may or may not have changed due to my change in status but it is clearly and ad hominem attack.
- FabCH 2mo agoNot the OP, but: The point was not that their opinion is suspect, the point was that they are former because people who care about the customer get fired and/or that everyone who cared is former, so nobody who is left cares.
- switchbak 2mo agoIf I worked at a place like that, I'd sure as hell work my butt of to get a job somewhere else. Or in my case, actively ignore any and all recruiting from that sesspool.
- zelphirkalt 2mo agoMaybe they are former AWS employees for a reason and now want things to go better than they were at AWS.
- nullsanity 2mo ago[dead]
- tmpz22 2mo agoAWS has had this reputation for over a decade. Every former AWS (including poached not fired) has relayed to me a verson of this. Every once in a while (~1:12) you get one that sounds like a Mormon missionary praising how its not that way and AWS is perfect.
- embedding-shape 2mo ago> Every once in a while (~1:12) you get one that sounds like a Mormon missionary praising how its not that way and AWS is perfect. The strange thing is that I only come across those missionaries online and never in person (although I did go to a AWS event once [never again] and met a bunch of them, so seemingly they are actually real, to my surprise).
- mendigou 2mo agoI am former AWS and this is pretty accurate. The other factor to add here is that, with some exceptions, the whole company feels like a Rube Goldberg machine and very few people care about what happens outside their cog (because they’re not incentivized to do so).
- qurren 2mo agoYes, I am a former AWS employee. I got put on Focus because my "contributions were not coming through" to leadership.
- nullorempty 2mo agoThis ^^^ amplified by indifference and not giving a shit caused by "AI Adoption". There is literally no fucking reason to try to improve your skill. Any IDIOT with AI will do an OK job. And no one is shooting for better than OK.
- hoppp 2mo agoYeah but this also negates the argument that people gotta use it or they get left behind. If it takes no skill or intelligence then nobody can get left behind because it's very fast to get back in.
- mightyham 2mo agoSpeaking from my experience at Amazon this is not the case. Any customer impact like this would necessitate a COE (correction of errors) report, which means a list of required action items to prevent such issues from happening again, which typically suck up at least man-month of labor. Not to mention the report itself, which has to be written by a manager. In fact, there are regular AWS-wide meetings where L10 technical staff will randomly pick and review reports from across the organization. Getting picked for one of these is not a fun experience. COEs are such a huge annoyance for teams that they create a strong incentive to be proactive in preventing issues like this from happening. One of the rules when it comes to writing COEs is that they are not the fault of individuals but processes; but in reality, no one wants to be the cause of one.
- eventualcomp 2mo agoAmazon is heterogeneous. So much so, that positive anecdotes and negative anecdotes are near worthless without specifying the org. Depending on if you're a cost cutting team, fixed expense team or organization, if you're a revenue driving team, or if you're a core team, or the very many other splits you can come up about the relationship between the expense/balance sheets and the team itself...there are very very different attitudes towards COEs and leadership principles.
- dlenski 2mo agoThis was very much my experience, having worked in two different sub-organizations at AWS, and on several different services, in two different countries. There's just extreme variation in the quality of the management, the quality of the engineers, the operational/development role split, the on-call schedule, and the development and testing methodology.
- awakeasleep 2mo agoHaving been the manager writing those reports, you can only practically find causes that are within a single team’s ability to resolve. If you find a problem like this thread’s hypothetical, the process stops being an annoyance just to line level managers, and something that directors and vice presidents need to handle by changing strategic priorities within their organizations. That entails a real loss of face for them, and because they are the ones who actually run the show, it would will only happen if you have one that is naïve or a masochist. In either case that moves them out of management.
- thewebguyd 2mo ago> take the 3am on-call and be the hero instead Ah yes, the good old ITism "Everything's good, what are we even paying you for?" followed by "Everything's on fire, what are we even paying you for?" I moved out of it largely for that reason, am now an infrastructure/IT project manager, quite refreshing actually.
- qurren 2mo agoThe trick to surviving under such management is to jump in and put out other peoples' fires but not spend time preventing them even if you know how to.
- rootsudo 2mo agoThis, exactly 100%
- dice 2mo agoHow did you swing that transition? Did you study for PMP before applying around, leverage network to get in the door then backfill skills, or what?
- thewebguyd 2mo agoYeah, I got my PMP before applying around, combined with some luck I suppose. My IT role was basically a solo sysadmin before where I basically was the technical PM + engineer in one, and I did that for about 8 years so I had a ton of experience I could spin on my resume.
- 01284a7e 2mo agoIf only there was some way to get anomaly detection services [1] inside of AWS... 1: https://aws.amazon.com/what-is/anomaly-detection/ https://aws.amazon.com/what-is/anomaly-detection/
- mcpherrinm 2mo agoWhile I didn't work on AWS, I did intern on the retail side of Amazon, and there's definitely this sort of monitoring in place. Surely somebody was paged. And even if not, this is "just" the cost explorer estimations, not what is ending up on folk's bills. I learned about <https://en.wikipedia.org/wiki/2011_T%C5%8Dhoku_earthquake_and_tsunami https://en.wikipedia.org/wiki/2011_T%C5%8Dhoku_earthquake_an...> from alarms like this, as sales in Japan almost entirely stopped. I've been told a tale of another incident where some customer ran some huge cpu-intensive workload that didn't do any networking. It caused various alarms to fire because it "looked like" a part of the network was idle (potentially indicating some sort of networking failure) It's generally (in the broad sense) easy to add alarms for things going wrong, but in my experience anomaly detectors are just as likely to fire from other weird things like that happening.
- oenton 2mo ago> there's definitely this sort of monitoring in place. Surely somebody was paged. Well you’re half right. Either there wasn’t monitoring for this or if there was monitoring in place that’s not what caught this, because the page originated from a customer support ticket.
- Groxx 2mo agoThey've already got anomaly detection: their users.
- el1s7 2mo agoI think billing is the only thing AWS doesn't really care about optimizing or putting enough tests to avoid anomalies lol.
- dclowd9901 2mo agoClearly some folks here have never worked at Amazon before. It's genuinely terrifying to see what most of the internet runs on.
- mvdtnz 2mo agoWhy would you think there are "no tests"?
- 27183 2mo agoWe have a pretty strong existence proof... the thing happened in production. Unless they have some means to override a failing test and scp broken shit to prod, there wasn't a test.
- nullorempty 2mo agoTechnically, there could be a test. It could just be wrong!
- 27183 2mo agoIf a tree falls in the forest and nobody hears it... [edit] Testing your tests, like testing your backups, is a good idea
- fragmede 2mo agoYes, test the negative case as well. eg if you get the system setup so you can log in, also make sure you get permission denied for bad login info.
- 27183 2mo agoYeah, negative tests are more important than the "test your tests" thing. Negative tests are themselves regression tests. Testing that a test works is something you do at dev time, it doesn't really live in CI. Negative tests, as well as positive tests, establish invariants on the code under test. They're effectively permanent, immortalized in the CI suite. Testing that the test actually works is a one time thing at the time the test is written. Permute the code under test, check that the test fails in the expected way, done. You just have to trust your colleagues won't edit the tests in such a way that vandalizes the invariant. That's why you can never trust an LLM to edit a test. They're notorious for tweaking tests such that they pass.
- deltaray3 2mo agoIt's just like in Superman III
- CobrastanJorji 2mo agoThere will have been tests, but there will have been missing end-to-end tests. Test 1 will verify that the new system/product emits billing entries in some expected way ("We did 100 bytes of operations and we see we called the billing system for 100 bytes of stuff, yay, test pass"). Test 2 will be in the billing system ("We provide an incoming bill for SKU#12345 for 100 gigabyte-units and we see it costs $17, yay, test passes"). But they won't test the two things together because it will be harder to do and the teams will have different management chains. Seen it happen several times at several companies. Somebody will have said at some point "we should actually have the tests charge money" and somebody else will have said "well we can't have the tests actually charge money, that's a legal/accounting problem, it might even be a crime" and then nobody would have asked what the next best thing was.
- qeternity 2mo agoBut these aren't the right services where the test should be, right? There's another service that says "ok we take the 100 bytes from A, and we take the $17 SKU from B, and this should equal $X". It's the third service that multiplies these things that failed. Where are the tests for that?
- DrewADesign 2mo agoTo me this sounds like a human-or-LLM-driven error. There must be a pretty limited set of factors that determine a pricing unit: I’m not really sure how a deterministic system could do that infrequently enough to not be a bigger story. Maybe a reeeeaaaallly rare race condition or something like that? To me this smells like having enough manual work involved in the process to fuck something up, but not nearly enough eyes on it to notice. I’m totally guessing though.
- lostlogin 2mo ago> human-or-LLM-driven error. In a computer system, dont those categories cover pretty much everything except a meteor strike?
- vasco 2mo agoI'm sorry but anyone that sees a multi million or billion dollar bill on an account that does nowhere near that should not be scared. It's obviously a mistake. Stories like this have happened with banks in my country. Check your account and you have billions in there. Guess what happened to those that withdrew money? The judge told them any reasonable person would know this is a bug. Had to give it back. Same thing here, any reasonable person doesn't get scared.
- simmerup 2mo agoI’d be concerned as it throws into doubt their entire accounting mechanism
- yawaramin 2mo agoThe problem with an anomalous AWS bill is that I don't know if it's because of a bug on their end or because of a goof on my end, eg did I accidentally leave on some giant compute for a month or something. With a billion in my bank account I at least know that I definitely didn't put it in there.
- znsnsksjiaja 2mo agoIf I experience a bug when transferring money (switching a digit or some such) they’ll shrug their shoulders and say there is nothing they can do. No amount of proof will push them to increment and decrement some integers over there. They transfer money to me. Their mistake: there is nothing I can do. Rules for me but not for thee?
- SlightlyLeftPad 2mo agoIt’s fractions of a penny Peter.
- bas 2mo agoWho needs tests when you have vibes?
- andai 2mo ago>no tests? Earlier this week my slopservant implemented several comprehensive changes to a codebase. It also wrote extensive tests to verify the correctness of the changes. A few days later I was working on something else and realized, everything had been implemented backwards, in a way that was nonsensical and also completely pointless. The many tests it had written were just confirming the LLM's idea of correctness, which turned out to be... completely incorrect. I laughed when I realized, if I had been using Rust, or indeed, formal verification, that wouldn't have helped at all — it would have just written a mathematical proof, proving the correctness of the wrong thing! Not sure what lesson to take from that (except read the damn diffs, obviously — it was a hobby project okay ;), but it seems like the more reliable this stuff gets, the more we expect it to work properly, the more risky it becomes. https://en.wikipedia.org/wiki/Normalization_of_deviance https://en.wikipedia.org/wiki/Normalization_of_deviance
- rtpg 2mo agowrite the tests yourself. Or at least the sketch of the test (to get filled in later). Make a commit of what you do so you can then look at the changes. If you write the tests the agents have an easier time writing the code!
- antihipocrat 2mo agoI'm struggling with this on a hobby project now. I'm torn between continuing the fast pace of development and taking a pause and checking the entire codebase. The more features I add now will likely make the inevitable refactor much harder, but adding new features is so easy that I want want extend this illusion of productivity just a little bit longer!
- andai 2mo agoYeah my logic was, we're doing so many refactors that it doesn't make sense for me to start memorizing how things work until things settle down a little. The argument in favour of "read each delta" (aside from catching them doing stupid shit) is to keep your mental model synced. But I already didn't understand my own code before AI, so that's moot too!
- hsbauauvhabzb 2mo agoWhat could a banana possibly cost Michael, ten dollars?
- roadhero 2mo agowe invoice clients monthly. if we ever sent one a bill 2^30 times too high, I don't think "estimated charges, no action required" would save the relationship :D
- DonHopkins 2mo ago...in the off chance that one in 100,000 will pay their bill instead of dropping dead from a heart attack (perhaps because there's a heartless OpenClaw running their accounting department), it's all worth it.
- nonameiguess 2mo agoThis reminds me of a discussion a few months back from a BSD maintainer who had done a lot of volunteer work for AWS over the years. I think he might have even been the person who alerted them to the insecurity of IMDSv1. There was a sense that AWS might have had great talented developers at the time, but they clearly didn't really understand the domain of running a hypervisor service exposed to the public. This feels like a similar situation, where I'm sure they "test" what they know to test, in the way they know, but compared to a bank, well, it wouldn't surprise me the least if no one at AWS ever even thought to ask banks how they handle things like this. Instead of testing that a process works the way you specified it, you need to make sure you even have the right specification in the first place by consulting with prior art and ensuring whatever you translate into software reflects the legal reality of the process you're trying to encode. There's a similar thing with physical processes. My wife works in geointelligence ground processing and encountered this when her system was expanded to support SAR collections instead of just visual spectrum. Software that passed all of its internal tests was passing nonsense collection parameters because the developers didn't understand the difference between energy collected to a sensor cell reflected from the sun versus energy reflected from your own active scan. The capabilities and limitations aren't the same. You can perform the same processing. The bits on disk won't care. But if the process you encode doesn't accurately correspond to the physics of reality, you're producing nonsense. Memory-safe, syntactically-valid nonsense, but nonsense nonetheless. This seems to happen quite a bit with software companies.
- spydum 2mo agoI have bad news if you think banks are the pinnacle of cyber security practices... Maybe some rare few are, but by and large they generally stink. They tend to give that impression because they hire tons of auditors and panjandrum, to hassle their suppliers, but internally they are winging it like the rest of us.
- antonvs 2mo agoThe difference is that at bigger banks at least, they tend to throw more money at the problem. They'll have defense in depth with e.g. DMZs, proxies for ingress and egress, locked down workstations and browsers, etc., and a different team for each one of those domains. So while the actual "on the ground" picture may look suboptimal at any given point, overall it does make for security that in practice is much better than average. This is reflected in the actual security compromise statistics. Your money in a bank is a lot safer than, say, your credit card deals on file with many retailers.