6 ms·
You must not have been spending million per year. A friend's company spends 10 million/year with GCP, which isn't huge, and can have an engineer from any group
by shiftpgdn 3y ago
You must not have been spending million per year. A friend's company spends 10 million/year with GCP, which isn't huge, and can have an engineer from any group in a meeting the next day after a high priority issue.
How frequently are you engaging your account reps? You should be able to get the ear of a PM within 48 hours in most cases.
- cornel_io 3y agoYeah, I spend a small fraction of that and have found GCP support through our account rep to be extremely good, at least on par with what AWS provides. Maybe the particular rep makes a big difference?
- Palmik 3y agoThe parent said "my company that spends many millions per year".
- shiftpgdn 3y agoI understand that and am calling that claim into question.
- freedomben 3y agoand the call out turned out to be right. they clarified that they don't spend those millions with Google. So GP was being honest but the skepticism led to an important clarification, because I read it the same way as parent
- Palmik 3y agoI see, my bad!
- pclmulqdq 3y agoThey said that they spend millions per year on GPU compute, not that they spend millions per year with GCP. They later said that they did not spend that money on GCP. That makes sense because GCP doesn't have all that much capacity for scale-out GPU compute (they offer TPUs for that).
- leetharris 3y agoWe were trying to get TPUs. We work in audio AI (ASR, TTS, translation, etc). When we saw whisper-jax, we wanted to test the viability of the TPU platform. From our napkin math, the cost/performance ratio seemed great as they always mention on these TPU blog posts. I posted this in another comment, but the default provisioning for V4 and V5 TPUs is ZERO. They don't tell you this anywhere. So when we'd try to allocate V5 TPUs on our GCP account, it would just fail with a generic error and a huge error number that led to nothing in a search. So I reach out to our GCP rep we had been working with. After about 10 days and 3 follow-up emails/calls that went unanswered, she replies, "you have to fill out this form." I click on it and it's a Google Form. The same type of Google Form you and I can make. I submit it. To this day I have heard nothing back. I reach out to various executives at GCP we had been talking with. They said, "you have to fill out a form." I tell them we did, and they say, "oh, it's usually pretty fast." I heard this so many times from so many people. It seems that not a single person who actually works in sales or account management knows how this process works. When I finally got a response, they told me that "my account has no billing account associated with it." I showed them yes we do, they replied that they were trying to provision for the wrong account. It took another couple weeks to get a follow-up response. Luckily they eventually connected us with one of their partnered consultants who was finally able to help us, but by then we just decided to go back to GPUs on other platforms because it was such a miserable experience and in that time, all of our providers came through with the volume we needed.
- leetharris 3y agoWe were trying to move to their platform because we needed high-scale TPU or GPU compute. We were not existing customers. We currently spend all those millions elsewhere (AWS, LambdaLabs, Coreweave, and many more) and will continue to do so. At my previous company, we even had an announced partnership with Google and did lots of co-marketing. That doesn't mean much when a GCP engineer changes something in GCP and breaks your production. I wish it worked out. I liked working in GCP outside of these problems. I really like GKE.
- WendyTheWillow 3y agoFWIW, my experience in AWS, as far as "engineer changes something and breaks your production," is not without its own set of issues. Perhaps the AWS version of this story hasn't impacted you, but AWS also constantly tweaks its services, sometimes to the detriment of its users.
- belter 3y agoReferences please.
- matthewcford 3y agoI've seen issues with RDS after an update. We were seeing CPU max out; adding more readers didn't help, and in the end, we had to rebuild the entire cluster. This somehow fixed the issue, and CPU use returned to normal levels; we didn't make any code changes to the app.
- WendyTheWillow 3y agoYou want me to provide references for my personal experience?
- belter 3y agoNot necessarily. I would like examples of where "...AWS also constantly tweaks its services, sometimes to the detriment of its users..." so I can be better informed, and make better decisions within my projects.
- usmannk 3y agoI used to work on cloud infra at an org that spent many millions on GCP (I'm certain ~everyone reading this comment would recognize it by name) and was working on adding more zeros to that number. We spent years working on this and still ended up switching to AWS in the end, largely because of support but also GCP's anti-customer business practices in general. We had weekly meetings with our account reps in person at our office but even they weren't able to get things done internally. No fault of their own though!