9 ms·
If it’s not zhipu then why is it returning errors that zhipu does for other models? Who else would return the exact same errors even if they took a lot of core
by gvkhna 23d ago
If it’s not zhipu then why is it returning errors that zhipu does for other models? Who else would return the exact same errors even if they took a lot of core infra like tokenizer from z?
- e9 23d agoSomeone could've trained model on top of GLM. Same way Cognition trained their SWE model on top of Kimi and Cursor did same with their Composer model.
- gvkhna 23d agoWhile possible the amount of variation in serving infrastructure is unlikely to land with actually giving the exact same errors zhipu does. It feels like glm flash, and there was a report zhipu had secured a huge new cluster suggesting they have the capacity. My guess anyway. https://www.tomshardware.com/tech-industry/artificial-intelligence/z-ai-powers-up-1gw-ai-data-center-built-entirely-on-chinese-chips https://www.tomshardware.com/tech-industry/artificial-intell...
- vitorgrs 23d agoThe reasoning levels are the same as GLM 5.3. GLM 5.3 is still not open... I believe it's GLM 5.3 Flash or Air.
- weiran 23d agoReasoning levels are often just injected system prompts so not a great way to fingerprint models.
- eli 22d agoBut it's an error, not a response.
- ggcr 23d agoZiphu has that many resources to be able to serve capacity for 1 quadrillion tokens per day on Nous portal? My bet is that it's a Composer model from Cursor running on xAI cluster, they already did a Composer based on Kimi-K2.5
- LaurensBER 23d agoThere's three options here: - The provider has a massive amount of (unused) hardware. Google or Cursor seem most likely - The model is extremely efficient, beyond anything we've seen so far - Whomever made the model has improved the cache efficiency in such a way that it's very cheap to serve. See e.g Deepseeks or Xiaomi caching (pre-price increase)
- re-thc 23d agoOption 4: the claimed capacity is not true. Real world usage hasn’t reached anywhere close to it.
- johndough 23d ago> 1 quadrillion tokens per day on Nous portal If you are referring to this number (https://xcancel.com/NousResearch/status/2090899914700054780 https://xcancel.com/NousResearch/status/2090899914700054780), they are either mistaken, or they mean that they can route 1 quadrillion tokens per day, but the provider behind Ox Alpha certainly can't provide that. Almost all of my requests have hit a rate limit so far.
- SyneRyder 23d agoI was hitting 429 overloaded regularly with Ox on OpenRouter yesterday, but a lot of that turned out to be problems with my harness. I fixed some bugs, improved the back-off, and I haven't hit a 429 error since (touch wood). OpenRouter says they're doing 6 Trillion tokens a day with Ox Alpha so far, and it has been their biggest launch of all time. OpenCode claimed they had capacity for 100T a day. https://x.com/OpenRouter/status/2091912024922177562 https://x.com/OpenRouter/status/2091912024922177562 https://x.com/opencode/status/2090544355824038300 https://x.com/opencode/status/2090544355824038300