8 ms·
They realized Astra having the same score as Sol was silly so they rushed to update the index so it fits what people expect. The old index was clearly bad (Ast
by redox99 12d ago
They realized Astra having the same score as Sol was silly so they rushed to update the index so it fits what people expect.
The old index was clearly bad (Astra is way better than Sol) but it's also unscientific to tweak it like this.
- paimapi 12d agois any of this 'scientific'? does AA allow peer review of its processes? are these published in journals of at least medium impact? what are the sample sizes? how grounded is their mechanistic reasoning? like they have words that are dressed in scientific language on their site like "We estimate a 95% confidence interval for Artificial Analysis Intelligence Index of less than ±1% - based on experiments with >10 repeats on certain models for all evaluation datasets included in Artificial Analysis Intelligence Index v4.2." but where's the outcome dataset justifying this? how did they get that probability? what was the specific methodology of the tests? what variables did they account for? there's a major difference between scientific sounding and being truly empirically rigorous. the 'research' in AI intelligence feels somehow even less trustworthy than supplement-funded studies because those are at least subjected to scrutiny by peers without profit motives
- elvin_d 12d agowhy keeping journals as an argument on the tech site that was always less formal with mandatory institutions but more open source. Journals discredited themselves multiple times. tech is expanding boundaries of scientific methods and AI will push it more.
- tancop 12d agoIt's not about journals. If they want to be 100 trustworthy they should release end to end reproducible pipelines for the whole process with everything but the private datasets included, and all design decisions fully documented. And let third party labs audit to confirm that the holdout questions are equal difficulty and similar task types to the public ones.
- deleted 11d ago[deleted]
- paimapi 11d ago1) open access is a rising trend in journal publishing. arxiv exists because of it but this also means that a lot of papers get 'published' before they are peer-reviewed which is often when significant issues are caught 2) there are definitely problems with journals and publishing but, like many similar absurdly reductive pronouncements, the argument that the current mode of scientific inquiry is bunk is both wrong and lacks nuance the problem with modern publishing is that private equity is buying up publishers [0]. these publishers are then giving peer-reviewers no time and zero pay to do the necessary work of review [1] while also charging exorbitant rates for access. this is leading to worsening quality of the published research along with highly overburdened researchers who are stuck between shrinking funding [2] and their myriad other professional obligations to just say that 'journals' are bunk is ignorant in a harmfully anti-empirical way. the process of empirical research and review is the entire reason why we see realworld results. foundations comprised of bullshit crumble fast but for some reason or another our economic system is highly driven to enshittifying everything it touches to have defenders who claim AA is 'pushing the boundaries of the scientific method' sounds like the screeching refrain of anti-intellectual cargo cults, apeishly mimicking the features of rigor and methodology while avoiding any real accountability [0] https://issues.org/how-academic-science-gave-its-soul-to-the-publishing-industry/ https://issues.org/how-academic-science-gave-its-soul-to-the... [1] https://www.insidehighered.com/news/faculty/books-publishing/2026/05/14/will-paying-reviewers-ease-peer-review-crisis https://www.insidehighered.com/news/faculty/books-publishing... [2] https://www.nature.com/articles/d41586-025-00754-4 https://www.nature.com/articles/d41586-025-00754-4
- elvin_d 11d agoYes, journals are harmful to freedom of knowladge which might seem ignorant especially to people coming from academia and mainstream who doesn't know better and can't imagine anything else besided wallen gardens they grew up with. Arxiv is dominated by CS [1] and nothing really changed over the years [2]. Saying about pushing the boundaries of scientific methods it was another rock into journals and the status quo on how it's done like p-hacking and other data manipulations. Being in open with community notes would would have a difference. Empirical research can exist without journals, science can be for masses and journals are antithetical for sharing making the knowledge exclusive. [1] https://info.arxiv.org/about/reports/submission_category_by_year.html https://info.arxiv.org/about/reports/submission_category_by_... [2] https://en.wikipedia.org/wiki/ArXiv?useskin=vector#/media/File:ArXiv's_yearly_submission_rate_plot.svg https://en.wikipedia.org/wiki/ArXiv?useskin=vector#/media/Fi...
- deleted 11d ago[deleted]
- Wheen 11d ago> but where's the outcome dataset justifying this? how did they get that probability? what was the specific methodology of the tests? what variables did they account for? It's just high school statistics: https://en.wikipedia.org/wiki/Confidence_interval https://en.wikipedia.org/wiki/Confidence_interval It's not a probability. It's essentially an assertion that if the test were run 100 times, the result would be within the interval 95 times.
- kingstnap 12d agoIn real life, there is always this feedback edge from the results to the methodology. Theoretically it's unscientific to do tweaks like this but in reality this is what actual science is because you need to see the results understand the problems in your experimental designs. Now the interesting thing is that you could repair the bias problem. The main issue is that you will do these tweaks after a bad result not after they look fine. Maybe we could commit ahead of time that the experiment will be reanalyzed after results regardless of what they are. Instead of only when the results disprove the hypothesis. So like artificial analysis committing to a fixed cadence of index updates instead of when the results start looking jank.
- x312 12d agoGiven how many private benchmarks they're using now, its likely they just tested different combos until they got the result they wanted. Completely discredits the index if it just gets modified to match social media vibes.
- villish 12d agoHopefully this wakes people up from this addiction to benchmarks when discussing various AI models. Different model families have strengths and weaknesses in various domains, but those are never discussed.
- weird-eye-issue 12d ago> but those are never discussed. They are literally frequently discussed and it's why there are different benchmarks for different domains.
- villish 11d ago> different benchmarks for different domains. Labs know the only thing the public even discusses on model releases are benchmarks, so they devote a majority of training on just benchmaxxing. It’s marketing. Muse Spark looks great in benchmarks. Everyone I know who has tried it (Rust & C++ projects) has determined it’s a resounding “meh”. That doesn’t mean it’s not a great tool for frontend devs, I wouldn’t know.
- throw10920 12d agoWhat evidence do you have that they tweaked the index to fit what they thought people expected?
- weird-eye-issue 12d agoIt's a benchmark for brand new technology and you admit the old score was silly but you think updating it is unscientific?
- andriy_koval 12d agoMaybe there is truth in it? I asked Astra to make some changes, it built crazy overengineered code, switched back to Sol, said code is too complicated, rewrite it, and it made lot cleaner code. I feel some models could chaise complicated benchmarks too much in expense of simpler tasks quality.
- cbg0 12d agoSol frequently over-engineers code.
- big-chungus4 12d agoWell, do you have any ideas about how to make it scientific? And they clearly needed to do something based on their unnatural ranking
- baq 12d ago> Astra is way better than Sol This needs to be evaluated per task, jagged frontier yadda yadda. I would be not at all surprised if sol was better at some things than Astra just like people still use opus 4.6 and for good reasons.
- cma 11d ago> Astra is way better than Sol For knowledge type questions is that necessarily true? In the past we've seen things like Google's models degrading on general knowledge after the preview releases while improving on code/tool use, presumably due to catastrophic forgetting from the additional training. Their preview would be free, get lots of agentic use from users, then additional training on that and probably additional automated RL. However Astra is on an entirely new base model so I also wouldn't expect it to be worse.