6 ms·
the 2nd website is not official, just something someone slopped together for some reason.
by sunbum 22d ago
the 2nd website is not official, just something someone slopped together for some reason.
- Alifatisk 22d agoI have plenty of these websites, I can’t understand why someone is doing this.
- colesantiago 22d agoIt is called phishing and grifting. Many people and even software engineers fall for this all the time. Most of these people are from crypto pivoting to AI doing this. AI has made this easier and cheaper and it is going to get a LOT worse. Imagine lots of websites with typosquatting and looking exactly the same as another website, vibe coded and cloned within seconds. The public have no chance.
- Alifatisk 21d agoWhat is there to phish? These are simple vibe coded websites providing information for a certain topic, nothing else. In this case, that 2nd url is a website with information regarding the new model as well as a broken chat interface to try out.
- colesantiago 21d agoYou do realise there are hundreds of these types of 'sites'. This one that is listed is designed to rank on Google as an informational source (although unofficial and not from z.ai which is why I said it is phishing) Assuming you are technical you are able to discern this, imagine the average person. No chance.
- Alifatisk 20d ago> You do realise there are hundreds of these types of 'sites'. Yes, and its these sorts of websites I am asking about. > This one that is listed is designed to rank on Google as an informational source (although unofficial and not from z.ai which is why I said it is phishing) Again, what is there to phish?
- MrDrMcCoy 20d agoPhishing implies exploitable data collection. Is that happening here?
- yorwba 22d agoEven if it weren't slopped together, 65% vs 80% on 10 tasks just isn't a significant difference. For 80% power to distinguish at a significance level of 0.05, you'd need more like 140 samples, if those were the true success probabilities. The number one problem in LLM benchmarking is that people try to draw conclusions from sample sizes far too small to conclude anything but "it works sometimes, it fails sometimes, hard to say which is better." (The number two problem is that people run benchmarks blindly without checking that they measure something meaningful.)
- brotchie 21d agoI'm going steal "slopped together", great quip.