8 ms·
> This is also causing issues with Amplify, API Gateway, AppStream2, AppSync, Athena, Cloudformation, Cloudtrail, Cloudwatch, Cognito, DynamoDB, IoT Services, L
by rcardo11 6y ago
> This is also causing issues with Amplify, API Gateway, AppStream2, AppSync, Athena, Cloudformation, Cloudtrail, Cloudwatch, Cognito, DynamoDB, IoT Services, Lambda, LEX, Managed BlockChain, S3, Sagemaker, and Workspaces.
Well, this is a major outgage
- Schweigi 6y agoIndeed, we had the first AWS Kinesis issues already at 13:50 (UTC). Now it's still ongoing after two hours. The status page didn't even update in the first 45 min or so...
- jjoonathan 6y agoThat's typical. The AWS status page is a marketing gimmick whose job is to stay green, not a good faith attempt to assess and report status. If there's an outage, seeing it accurately reflected on the status page is the exception, not the rule.
- simlevesque 6y agoIsn't that fraud ? edit: not sure why my question deserved a downvote...
- WrtCdEvrydy 6y agoIf you're small, yes, if you're AWS, it's business as usual?
- tootie 6y agoAs of this moment, there are more non-green services than I've ever seen. And it's steadily getting worse. EDIT: 15 minutes later and the board is looking worse again.
- chizhik-pyzhik 6y agoUpdating the status dashboard is pretty low priority for operators trying to resolve this issue. It requires escalation up the management chain and careful wording.
- jjoonathan 6y agoBy design. If it was a good faith attempt to report status, it would be automatically updated from a flock of canaries instead of through a slow, political process.
- 35fbe7d3d5b9 6y agoEven that would be meaningless at the scale of AWS. "A top of rack switch let out the blue smoke and it'll be ~30 before we can re-rack it" would impact what fraction of a fraction of a percent of canaries? Irrelevant to me, unless of course my VM lives on a box backed by that switch. ;) The status dashboard exists for us to laugh at when things break and to convince C*Os that everything is fine. That's it.
- jjoonathan 6y agoEhhh... the ratio of "bump in the night" problems that affect just me to genuine outages that cross regions and affect others is about 1:1, and then about 1 in 4 or 5 of the cross-region problems blow up to the scale where they feel forced to update the dashboard. So I disagree, I think a canary flock would be both meaningful and useful. As you point out, though, the status dashboard isn't truly meant to be either of those things. I don't have any illusions about it ever changing.
- PKop 6y ago>for operators trying to resolve this issue It's a shame Amazon doesn't have thousands of employees to divide these tasks between different people, as it is only these busy operators who could update this status page. If you're right, why have the status page then? It is useless by your definition yes?
- booleanbetrayal 6y agoThis is also affecting Fargate (at least EKS) in that its scheduling system is broken. No way to get new pods.
- sethhochberg 6y agoSame story in ECS. Seems like virtually anything Fargate can't spawn new instances.
- zxcvbn4038 6y agoFargate console is reporting no capacity in us-east-1 which is a bummer because I've lost several services that got spun down apparently due to missing cloudwatch data. But EC2 appears to be working though its taking noticeably longer to create resources. I think the take-away for a lot of people is that multiple availability zones is not a substitute for proper BCP that encompasses multiple regions or cloud providers.
- jpp 6y agoWe're also seeing issues with FarGate ECS -- the task we had with auto-scaling scaled down to 0. The one we had with a fixed number of workers is fine.
- mcintyre1994 6y agoThanks for mentioning this - that's a nasty failure mode.
- holler 6y agoyep, iot in us-east-1 not working for me
- durkie 6y agoseeing issues with scaling up/down in elastic beanstalk too
- cpufry 6y agoand this is whats disclosed to the public
- manishsharan 6y agoThanks for this. My Lambda@Edge function was not working and I thought I broke something my permissions even though I had not touched that for atleast a month. This is the very "helpful" error message The Lambda function associated with the CloudFront distribution is invalid or doesn't have the required permissions. We can't connect to the server for this app or website at this time. There might be too much traffic or a configuration error. Try again later, or contact the app or website owner. If you provide content to customers through CloudFront, you can find steps to troubleshoot and help prevent this error by reviewing the CloudFront documentation.
- odiroot 6y agoIt's always a DNS issue.
- mdaniel 6y agoThat's a tough one -- I'm usually with you that it's always either DNS or cert expiry, but my go-to "it's always ..." when discussing AWS is: it's always security groups Heh, maybe they accidentally locked themselves out of IAM, since those are great fun to troubleshoot, also
- bart_spoon 6y agoI'm also seeing weirdness with Batch. Its working, but the dashboards aren't showing job statuses accurately and jobs aren't always terminating.