5 ms·
Mixed signals, here it's performing below even GPT-5.4 Nano: https://livebench.ai/ https://livebench.ai/ while here it outperforms Fable by a significant marg
by WithinReason 21d ago
Mixed signals, here it's performing below even GPT-5.4 Nano:
https://livebench.ai/ https://livebench.ai/
while here it outperforms Fable by a significant margin:
https://oxalpha.com/ https://oxalpha.com/
but if the latter is true, will people still say it was "distilled" from Fable?
- sunbum 21d agothe 2nd website is not official, just something someone slopped together for some reason.
- Alifatisk 21d agoI have plenty of these websites, I can’t understand why someone is doing this.
- colesantiago 21d agoIt is called phishing and grifting. Many people and even software engineers fall for this all the time. Most of these people are from crypto pivoting to AI doing this. AI has made this easier and cheaper and it is going to get a LOT worse. Imagine lots of websites with typosquatting and looking exactly the same as another website, vibe coded and cloned within seconds. The public have no chance.
- Alifatisk 21d agoWhat is there to phish? These are simple vibe coded websites providing information for a certain topic, nothing else. In this case, that 2nd url is a website with information regarding the new model as well as a broken chat interface to try out.
- colesantiago 21d agoYou do realise there are hundreds of these types of 'sites'. This one that is listed is designed to rank on Google as an informational source (although unofficial and not from z.ai which is why I said it is phishing) Assuming you are technical you are able to discern this, imagine the average person. No chance.
- Alifatisk 20d ago> You do realise there are hundreds of these types of 'sites'. Yes, and its these sorts of websites I am asking about. > This one that is listed is designed to rank on Google as an informational source (although unofficial and not from z.ai which is why I said it is phishing) Again, what is there to phish?
- MrDrMcCoy 20d agoPhishing implies exploitable data collection. Is that happening here?
- yorwba 21d agoEven if it weren't slopped together, 65% vs 80% on 10 tasks just isn't a significant difference. For 80% power to distinguish at a significance level of 0.05, you'd need more like 140 samples, if those were the true success probabilities. The number one problem in LLM benchmarking is that people try to draw conclusions from sample sizes far too small to conclude anything but "it works sometimes, it fails sometimes, hard to say which is better." (The number two problem is that people run benchmarks blindly without checking that they measure something meaningful.)
- brotchie 21d agoI'm going steal "slopped together", great quip.
- epolanski 21d agoGLM 5.3 was a great model, so this would be strange to release a regressed model
- ImprobableTruth 21d agoIt's probably GLM 5.3 flash, so weaker but cheaper.
- re-thc 21d agoWith vision on top
- deleted 21d ago[deleted]
- re-thc 21d agothe outperform Fable was a mid (not completed) benchmark run. Real results were lower.
- woadwarrior01 21d agoThat benchmark is super sus. Until someone pointed it out, the top performing open weights model was a Kimi K3 fine tune from their sponsor (abacusai/Smaug-Agentic). Now, it's not on the list. Source: https://twitterwebviewer.com/?tweet=2091116504787935350 https://twitterwebviewer.com/?tweet=2091116504787935350
- tescreal 21d agoI really want to see hard evidence of distillation before I buy into it. Seems like a lot of sour grapes over not having the sort of lead assumed. In this field, it has been shown repeatedly that leaps in performance come swiftly and without notice.
- xienze 21d agoWhat would constitute evidence in your opinion?
- anon373839 21d agoHow about proof that black-box distillation can deliver these results without a very sophisticated RL pipeline doing the heavy lifting?
- dannyw 21d ago"Black-Box On-Policy Distillation of Large Language Models", Microsoft Research, https://aka.ms/GAD-project https://aka.ms/GAD-project > 'GAD consistently surpasses standard sequence-level distillation, delivering superior generalization and achieving performance that rivals the proprietary teacher. These results validate GAD as an effective and robust solution for black-box LLM distillation.' No RL, although I'm a little bit surprised to see MS Research publishing a paper on distilling GPT5?
- anon373839 21d agoThat's an interesting paper, but there is virtually no discussion of reasoning behaviors or optimization for long-horizon tasks (i.e., all of the recent advances in LLMs that people care about). The evaluation methodology also is pretty dated: > We reserve 500 samples of LMSYS-Chat-1M-Clean as the primary test set. We also include test datasets consisting of a 500-sample subset split from Dolly [6], the 252-sample SelfInst dataset [37], and the 80-question Vicuna benchmark [3] to evaluate out-of-distribution generalization. We report the GPT-4o evaluation scores [45, 10], where GPT-4o first generates reference answers and then scores the output of the student model against them. We also conduct human evaluations on the LMSYS-Chat-1M-Clean test set for qualitative assessment.
- Aurornis 21d agoClaims about Ox Alpha performing at Fable level were from the social media hype cycle. Everything new in the LLM space brings a wave of influencers hyping it up as a revolutionary leap forward. Don’t forget to like and subscribe to learn more. It is a capable small model, but it’s not frontier level. The interesting part will be seeing the model size, how it responds to quantization, and how fast it runs on the kind of non-server hardware that we can buy without selling a kidney.
- worldsavior 21d agoThis influencers are getting paid, it's not coincidential.
- nixon_why69 21d agoThey don't have to be getting paid. The natural bias of media is towards laziness and sensationalism (stolen from Jon Stewart, so maybe the same is true about comments).
- kkukshtel 21d agoI think so much of this is people greenfield-ing things as benchmarks, which is nearly always a success case for any AI these days.
- plumeria 21d ago> Claims about Ox Alpha performing at Fable level were from the social media hype cycle. They claim an "independent community benchmark" (pass-fail evaluation on tasks) here: https://oxalpha.com/ox-alpha-vs-fable-5 https://oxalpha.com/ox-alpha-vs-fable-5
- sunaookami 20d agoThat's a fake slop hype website.
- dpweb 21d agoKinda useless to compare simply based on model without considering harness. Different agents handle the context etc completely differently. I would like to start seeing these model vs model comparisons across different harnesses.
- daralthus 21d agoomp+0x-alpha beat both cc+fable and codex-sol in creating/refactoring a big eval setup. the former just knows where things should belong and completed the task all the way while the other two failed on both metrics.
- dannyw 21d agoto be fair, OMP/Pi is also just a better harness. e.g. https://www.databricks.com/blog/benchmarking-coding-agents-databricks-multi-million-line-codebase https://www.databricks.com/blog/benchmarking-coding-agents-d...
- deleted 21d ago[deleted]
- jauntywundrkind 21d agoI've been having oxa and sol do architecture design then compare notes. Sol is definitely still way ahead. But there's reliably some really good wins ideas and concepts that OxA throws out there that Sol is very happy to encorporate. One thing that I think matters a lot for the non developers, all three of us (sol, oxa, and me) usually agree that oxa's write up is far far better. It explains the situation very well, and has great structure for its write ups. Sol gets the job done, but it's terrible at re-explaining the problem for humans, at laying out information. It also doesn't show it's thinking, so it's imo a terrible peer to work with!