6 ms·
I used their Coding Plan for a few months. It is genuinely difficult to keep up with the models. The output is so fast. Qwen 3.8 27B is likely one of the strong
by gardnr 13d ago
I used their Coding Plan for a few months. It is genuinely difficult to keep up with the models. The output is so fast. Qwen 3.8 27B is likely one of the strongest models they've hosted so far.
Edit: it looks like this is only available on a API token pricing. Does anyone know if they have rolled out prompt caching yet? It used to get pretty expensive for agentic coding tasks with no prompt caching.
- altertable 13d agoAgreed, but in our SAAS I can tell some UX will sky-rocket to next level with this
- jasongill 13d agoIt appears that they do support Prompt Caching: https://inference-docs.cerebras.ai/capabilities/prompt-caching https://inference-docs.cerebras.ai/capabilities/prompt-cachi...
- abtinf 13d ago> How are cached tokens priced? > There is no additional fee for using prompt caching. Input tokens, whether served from the cache or processed fresh, are billed at the standard input token rate for the respective model. Well, talk about flipping the narrative.
- the_duke 13d agoIt doesn't reduce the price though.
- jasongill 12d agoGood catch, I guess I got lost in the marketing speak of the page!
- deleted 13d ago[deleted]
- cute_boi 13d agoi believe they used to have monthly plan, what happened to that?
- eli 13d agoStrongest model that they host on the public endpoint. They do a super fast version of GPT 5.6 Sol for OpenAI and have bigger open models on dedicated endpoints.
- singpolyma3 13d agoThe coding plan is gone now right?
- gardnr 13d agoLast time I got one, I had to log into a Discord server and wait for "the drop" and IIRC Daniel Kim was giving them out based on who was there at the time. They were gone in less than a minute. This was ~8 months ago.