5 ms·
Wafer is making themselves synonymous with slop in the inference space. Exaggerated unfair comparisons in all their results, twitter hype posts with alarm emoji
by inferencecoder 2mo ago
Wafer is making themselves synonymous with slop in the inference space. Exaggerated unfair comparisons in all their results, twitter hype posts with alarm emojis etc.
> $2.50/GPU-hr for the MI355X, $6.00 for the B300, and $4.25 for the B200.
This is not an accurate price comparison for real terms.
- villgax 2mo agoI went the gpus.io website & it’s $2.95/hr right now, this is like comparing MSRP to actual market price. B300s are in demand & hence cost more, but these lazy editors at wafer.ai can't be bothered to do TCO on actual ownership nor share code to replicate their setups. Instead just relying on current market prices to win one row, which isnt even about per/$ on actual MSRPs.
- YetAnotherNick 2mo agogpus.io shows tensorweave pricing at $2.95/hr. Tensorweave just shows "Talk to sales".
- greyb 2mo agoThis is a company that launched a token subscription (WaferPass), before weeks later, rugpulling the plan for being unsustainable while simultaneously claiming they achieved incredible inference efficiency gains worthy of paying them mind.
- BoorishBears 2mo agoMost discourse around GPU prices is nonsense right now. Some people using Spot prices for providers who won't have Spot capacity, some people using hourly rates for instances that are never in stock, some people ignoring commitment discounts. Not to mention no one serious is serving this on 8xB200 instead of multiple nodes: the vast majority of Moonshot's inference work is focused on PD-disaggregation
- inferencecoder 2mo ago> Not to mention no one serious is serving this on 8xB200 instead of multiple nodes: the vast majority of Moonshot's inference work is focused on PD-disaggregation The GPU price discourse is absurd, but many are serving models on single node setups when the model fits
- BoorishBears 1mo agoYou can't beat current API pricing running Kimi K3 a single node, so as I said anyone serious is not doing that for Kimi K3. Not sure how you're getting "no one serves models on a single node" out of that.
- inferencecoder 1mo agoI meant serious players often do serve on a single node. They can beat API pricing as well. Multi-node can add gain, but also adds a lot of deployment complexity so its just not always possible or optimal.