6 ms·
You should try the arc agi puzzles yourself, and then tell me you think these things aren't intelligent https://arcprize.org/blog/openai-o1-results-arc-prize h
by swaraj 2y ago
You should try the arc agi puzzles yourself, and then tell me you think these things aren't intelligent
https://arcprize.org/blog/openai-o1-results-arc-prize https://arcprize.org/blog/openai-o1-results-arc-prize
I wouldn't say it's full agi or anything yet, but these things can definitely think in a very broad sense of the word
- daveguy 2y agoLLMs don't do too well on those ARC-AGI problems. Even though they're pretty easy for a person.
- cubefox 2y agohttps://arcprize.org/blog/oai-o3-pub-breakthrough https://arcprize.org/blog/oai-o3-pub-breakthrough
- daveguy 2y agoLet me know when they can perform that well without a 300-shot. Or that well on unseen ARC-AGI-2.
- aoeusnth1 2y agoTwo years give or take 6 months.
- mrbungie 2y agoGive a group of "average human" two years, give or take 6 months, and they will also saturate the benchmark and probably some humans would beat the SOTA LLM/RLM. People tend to do so all the time, with games for example.
- refulgentis 2y agoDone (link says 6 samples?)
- daveguy 2y ago> OpenAI shared they trained the o3 we tested on 75% of the Public Training set. I'm talking transfer learning and generalization. A human who has never seen the problem set can be told the rules of the problem domain and then get 85+% on the rest. o3 high compute requires 300 examples using SFT to perform similarly. An impressive feat, but obviously not enough to just give an agent instructions and let it go. 300 examples for human level performance on the specific task, but that's still impressive compared to SOTA 2 years ago. It will be interesting to see performance on ARC-AGI-2.
- deleted 2y ago[deleted]
- gessha 2y ago[Back in the 1980s] You should try to play chess yourself, and then tell me you think these things aren't intelligent.
- jhbadger 2y agoWhile I agree that we should be skeptical about the reasoning capabilities of LLMs, comparing them to chess programs misses the point. Chess programs were specifically created to play chess. That's all they could do. They couldn't generalize and play other board games, even related games like Shogi and Xiangqi, the Japanese and Chinese versions of chess. LLMs are amazing at being able to do things they never were programmed to do simply by accident.
- gessha 2y agoAre they though? They’ve been shown to generalize poorly to tasks where you switch up some of the content.
- refulgentis 2y agoThat's not true
- gessha 2y agoThey don’t generalize well on logic puzzles https://huggingface.co/blog/yuchenlin/zebra-logic https://huggingface.co/blog/yuchenlin/zebra-logic
- refulgentis 2y agoApologies, I was a bit curt because this is a well-worn interaction pattern. I don't mean anything by the following either, other than, the goalposts have moved: - This doesn't say anything about generalization, nor does it claim to. - The occurrences of the prefix general* refer to "Can fine-tuning with synthetic logical reasoning tasks improve the general abilities of LLMs?" - This specific suggestion was accomplished publicly to some acclaim in September - To wit, the benchmark the article is centered around hasn't been updated since since September, because the preview of the large model accomplishing that blew it out of the water, 33% on all at the time, 71%: https://huggingface.co/spaces/allenai/ZebraLogic https://huggingface.co/spaces/allenai/ZebraLogic - these aren't supposed to be easy, they're constraint satisfaction problems, which they point out are used on the LSAT - The major other form of this argument is the Apple paper, which shows a 5 point drop from 87% to 82% on a home-cooked model