7 ms·
> CS-4 delivers more than 1,000 tokens per second on models exceeding 10 trillion parameters Oops did they just out GPT-5.6 sol’s parameter count?
by reilly3000 29d ago
> CS-4 delivers more than 1,000 tokens per second on models exceeding 10 trillion parameters
Oops did they just out GPT-5.6 sol’s parameter count?
- whatever1 29d agoI mean we kinda know the frontier models are multi trillion parameter models. The only open weights that are close to the frontier are that size too
- verdverm 29d agosave qwen3.8 27B which is outclassing much larger models and is in spitting distance of the top 10 in https://artificialanalysis.ai/models#intelligence https://artificialanalysis.ai/models#intelligence
- Vax- 29d agoI wonder why they removed DeepSWE from their incorporates evaluations
- verdverm 29d agoThey didn't afaict https://artificialanalysis.ai/agents/coding-agents?coding-agents-performance-chart=deep-swe https://artificialanalysis.ai/agents/coding-agents?coding-ag... It seems it takes some time to run a new model on all the benchies, not sure they run all models on all of them either
- kanwisher 29d agocerebras model are different size then the original models
- sho 29d agoSol is supposed to be 5T according to rumour. The imminent Astra is allegedly 10
- nozzlegear 29d agoRumors and allegations aren't worth much. Why don't they just tell us mere mortals?
- brookst 29d agoWhy would they? What the upside, for them?
- eigenspace 29d agoYeah, its not like this js some sort of Open AI company. That'd be ridiculous.
- brookst 29d agoYou think their name means releasing competitive details would be good for them? I’ve got sone bad news about Federal Express.
- nozzlegear 28d agoIndeed, what is the upside of transparency?
- brookst 28d agoFor Linux? Assuring stakeholders that they can both control and observe development and how it works. For a charity? Assuring donors that funds are being managed appropriately. For Anthropic? No benefit at all. Turns out "transparency" is like "weight" or "velocity" in that it has no intrinsic value, and can be positive or negative depending on context.
- WinstonSmith84 29d agoBecause that would reveal their edge to investors, or the lack thereof. If Fable turns out to be a 10T or 20T model, there is little to boast vs Kimi at 3T. But the opposite is true: if Fable were to be e.g. a 500B model, that would show how far ahead they are from the open models. This isn't likely to be the case ...