6 ms·
Yes, but benchmarks like this are often flawed because leading model labs frequently participate in 'benchmarkmaxxing' - ie improvements on ARC-AGI2 don't neces
by fishpham 7mo ago
Yes, but benchmarks like this are often flawed because leading model labs frequently participate in 'benchmarkmaxxing' - ie improvements on ARC-AGI2 don't necessarily indicate similar improvements in other areas (though it does seem like this is a step function increase in intelligence for the Gemini line of models)
- jstummbillig 7mo agoCould it also be that the models are just a lot better than a year ago?
- bigbadfeline 7mo ago> Could it also be that the models are just a lot better than a year ago? No, the proof is in the pudding. After AI we're having higher prices, higher deficits and lower standard of living. Electricity, computers and everything else costs more. "Doing better" can only be justified by that real benchmark. If Gemini 3 DT was better we would have falling prices of electricity and everything else at least until they get to pre-2019 levels.
- ctoth 7mo ago> If Gemini 3 DT was better we would have falling prices of electricity and everything else at least Man, I've seen some maintenance folks down on the field before working on them goalposts but I'm pretty sure this is the first time I saw aliens from another Universe literally teleport in, grab the goalposts, and teleport out.
- WarmWash 7mo agoYou might call me crazy, but at least in 2024, consumers spent ~1% less of their income on expenses than 2019[2], which suggests that 2024 is more affordable than 2019. This is from the BLS consumer survey report released in dec[1] [1]https://www.bls.gov/news.release/cesan.nr0.htm https://www.bls.gov/news.release/cesan.nr0.htm [2]https://www.bls.gov/opub/reports/consumer-expenditures/2019/ https://www.bls.gov/opub/reports/consumer-expenditures/2019/ Prices are never going back to 2019 numbers though
- gowld 7mo agoThat's an improper analysis. First off, it's dollar-averaging every category, so it's not "% of income", which varies based on unit income. Second, I could commit to spending my entire life with constant spending (optionally inflation adjusted, optionally as a % of income), by adusting quality of goods and service I purchase. So the total spending % is not a measure of affordability.
- WarmWash 7mo agoAlmost everyone lifestyle ratchets, so the handful that actually downgrade their living rather than increase spending would be tiny. This part of a wider trend too, where economic stats don't align with what people are saying. Which is most likley explained by the economic anomaly of the pandemic skewing peoples perceptions.
- twoodfin 7mo agoWe have centuries of historical evidence that people really, really don’t like high inflation, and it takes a while & a lot of turmoil for those shocks to work their way through society.
- layer8 7mo agoIsn’t the point of ARC that you can’t train against it? Or doesn’t it achieve that goal anymore somehow?
- theywillnvrknw 7mo ago* that you weren't supposed to be able to
- deleted 7mo ago[deleted]
- egeozcan 7mo agoHow can you make sure of that? AFAIK, these SOTA models run exclusively on their developers hardware. So any test, any benchmark, anything you do, does leak per definition. Considering the nature of us humans and the typical prisoners dilemma, I don't see how they wouldn't focus on improving benchmarks even when it gets a bit... shady? I tell this as a person who really enjoys AI by the way.
- deleted 7mo ago[deleted]
- WarmWash 7mo agoBecause the gains from spending time improving the model overall outweigh the gains from spending time individually training on benchmarks. The pelican benchmark is a good example, because it's been representative of models ability to generate SVGs, not just pelicans on bikes.
- D-Machine 7mo ago> Because the gains from spending time improving the model overall outweigh the gains from spending time individually training on benchmarks. This may not be the case if you just e.g. roll the benchmarks into the general training data, or make running on the benchmarks just another part of the testing pipeline. I.e. improving the model generally and benchmaxing could very conceivably just both be done at the same time, it needn't be one or the other. I think the right take away is to ignore the specific percentages reported on these tests (they are almost certainly inflated / biased) and always assume cheating is going on. What matters is that (1) the most serious tests aren't saturated, and (2) scores are improving. I.e. even if there is cheating, we can presume this was always the case, and since models couldn't do as well before even when cheating, these are still real improvements. And obviously what actually matters is performance on real-world tasks.
- XenophileJKO 7mo agohttps://chatgpt.com/s/m_698e2077cfcc81919ffbbc3d7cccd7b3 https://chatgpt.com/s/m_698e2077cfcc81919ffbbc3d7cccd7b3
- aleph_minus_one 7mo agoI don't understand what you want to tell us with this image.
- fragmede 7mo agothey're accusing GGP of moving the goalposts.
- olalonde 7mo agoWould be cool to have a benchmark with actually unsolved math and science questions, although I suspect models are still quite a long way from that level.
- deleted 7mo ago[deleted]
- gowld 7mo agoDoes folding a protein count? How about increasing performance at Go?
- fc417fc802 7mo ago"Optimize this extremely nontrivial algorithm" would work. But unless the provided solution is novel you can never be certain there wasn't leakage. And anyway at that point you're pretty obviously testing for superintelligence.
- optimalsolver 7mo agoIt's worth noting that neither of those were accomplished by LLMs.