5 ms·
Gemini 3 Pro Preview gets 96.8% on the same benchmark? That's impressive
by Donald 9mo ago
Gemini 3 Pro Preview gets 96.8% on the same benchmark? That's impressive
- bigyabai 9mo agoGPT-5.2 might be Google's best Gemini advertisement yet.
- outside1234 9mo agoEspecially when you see the price
- capitainenemo 9mo agoAnd performs very well on the latest 100 puzzles too, so isn't just learning the data set (unless I guess they routinely index this repo). I wonder how well AIs would do at bracket city. I tried gemini on it and was underwhelmed. It made a lot of terrible connections and often bled data from one level into the next.
- wooger 9mo ago> unless I guess they routinely index this repo This sounds like exactly the kind of thing any tech company would do when confronted with a competitive benchmark.
- rsanek 9mo agoI mean, the repo has <200 stars, it's not like it's so mainstream that you'd expect LLM makers to be watching it actively. If they wanted to game it, they could more easily do that in RL with synthetic data anyway.
- capitainenemo 9mo agoBelated update on this. Gemini reasoning did much better than quick on bracket city today (an easy puzzle but still). It only failed to solve one clue outright, got another wrong but due to ambiguity in the expression referenced and in a way that still fit the next level down making the final answer fairly cleanly solved. Still clearly has a harder time with it than the connections puzzle.