9 ms·
The pricing is a little on the higher side. Working on a performance-sensitive application, I tried Mercury and Groq (Llama 3.1 8b, Llama 4 Scout) and the perfo
by asaddhamani 1y ago
The pricing is a little on the higher side. Working on a performance-sensitive application, I tried Mercury and Groq (Llama 3.1 8b, Llama 4 Scout) and the performance was neck-and-neck but the pricing was way better for Groq.
But I'll be following diffusion models closely, and I hope we get some good open source ones soon. Excited about their potential.
- tripplyons 1y agoGood to know. I didn't realize how good the pricing is on Groq!
- tlack 1y agoIf your application is pricing sensitive, check out DeepInfra.com - they have a variety of models in the pennies-per-mil range. Not quite as fast as Mercury, Groq or Samba Nova though. (I have no affiliation with this company aside from being a happy customer the last few years)
- asaddhamani 1y agoDeepInfra is amazing in terms of price, like really, they have the Qwen3 embedding model for $0.002 per mn tokens. That's an order of magnitude cheaper than most alternatives with better benchmark scores. But the performance P99 is slow and the variance is huge. For latency sensitive workloads it's problematic, if they can fix that it'll be a no-brainer to use them. DeepInfra does tend to have the lowest prices of any API provider.
- sexeriy237 1y agoYou're getting the savings by shifting the pollution of the datacenter onto a largely black community and choking them out.
- JimDabell 1y agoAre you confusing the AI company Groq with xAI, Elon Musk’s AI company that has a model called Grok?