7 ms·
Using gpt-5.4-mini in off-peak hours already feels like super-speed to me. That's probably no more than 100-150 tk/s. I can't imagine 750! I've always eyed Cer
by gandreani 3mo ago
Using gpt-5.4-mini in off-peak hours already feels like super-speed to me. That's probably no more than 100-150 tk/s. I can't imagine 750!
I've always eyed Cerebras but never had a use for it that would justify paying for the API directly. Although now that I think about it, trying out the API would probably cost less than a subscription for a month...
- kegs_ 3mo agoI have a pretty good use case for gpt-oss. The amount of time savings has actually been wild. Definitely worth a try. Just to be clear, it gets like 2000tok/s
- embedding-shape 3mo agoThe ChatGPT subscription gives you access to the -spark model(s) in Codex which are blazing fast (but pretty dumb) which I think runs on Cerebras hardware too.
- rrvsh 3mo agois this specifically in codex? have been trying to use the models for months on opencode then pi but it says chatgpt subscriptions don't have access to it - i was under the assumption that OpenAI doesn't lock down their models based on harness a la Claude Code
- cactusplant7374 3mo agoWhat plan are you on? It is only available to Pro users.
- rrvsh 3mo agoPlus. No wonder - i suspected this but i couldnt find any docs. Side note, how are you liking Pro? I have really been considering getting Pro recently, but not sure if its more worth to just switch to openrouter. I feel like my usage currently barely outstrips the plus usage limits and Pro would be too much, and using Openrouter by default would mean I would have a lot more leeway to run more random lighter workloads without worrying about using up my limit, but I'll really miss GPT 5
- cactusplant7374 3mo agoI find it to be excellent. I have three Pro subscriptions so I can build stuff 24/7. I only use GPT 5.5 xhigh. Before 5.5 I wasted a lot of time with bugs. I want to make sure that doesn't happen again -- if I can help it.
- jasonjmcghee 3mo agoTry gpt-5.3-codex-spark - it's 1000 TPS and from my experience more capable than 5.4 mini. If you have a subscription it's a different pool of usage.
- small_model 3mo agoUsed it, very fast but tiny context window and doesn't have good reasoning. (good for quick simple code changes)
- beering 3mo agoAgreed, 1000tok/s just fills up the context window (which is big by 2004 standards) super fast. But seems like 5.3-spark was just a taste of what’s to come.
- taneq 3mo ago2004 standards? O.o
- partsch 3mo ago1904
- bogeholm 3mo agoBack when we were kids, we would get 0 tokens/sec _if we were lucky_
- mlinsey 3mo agoIn 2004, I took a class where we trained "language models" that were bigram word models, on an archive of a couple years of the Wall Street Journal. I remember someone who literally announced they were dropping the class to the whole room at the end of a lecture, saying "This isn't AI!!!"
- trollbridge 3mo ago