10 ms·
Human performance is 85% [1]. o3 high gets 87.5%. This means we have an algorithm to get to human level performance on this task. If you think this task is an
by obblekk 2y ago
Human performance is 85% [1]. o3 high gets 87.5%.
This means we have an algorithm to get to human level performance on this task.
If you think this task is an eval of general reasoning ability, we have an algorithm for that now.
There's a lot of work ahead to generalize o3 performance to all domains. I think this explains why many researchers feel AGI is within reach, now that we have an algorithm that works.
Congrats to both Francois Chollet for developing this compelling eval, and to the researchers who saturated it!
[1] https://x.com/SmokeAwayyy/status/1870171624403808366 https://x.com/SmokeAwayyy/status/1870171624403808366, https://arxiv.org/html/2409.01374v1 https://arxiv.org/html/2409.01374v1
- antirez 2y agoNNs are not algorithms.
- notfish 2y agoAn algorithm is “a process or set of rules to be followed in calculations or other problem-solving operations, especially by a computer” How does a giant pile of linear algebra not meet that definition?
- antirez 2y agoIt's not made of "steps", it's an almost continuous function to its inputs. And a function is not an algorithm: it is not an object made of conditions, jumps, terminations, ... Obviously it has computation capabilities and is Turing-complete, but is the opposite of an algorithm.
- raegis 2y ago> It's not made of "steps", it's an almost continuous function to its inputs. Can you define "almost continuous function"? Or explain what you mean by this, and how it is used in the A.I. stuff?
- taneq 2y agoWell, it's a bunch of steps, but they're smaller. /s
- janalsncm 2y agoIf it wasn’t made of steps then Turing machines wouldn’t be able to execute them. Further, this is probably running an algorithm on top of an NN. Some kind of tree search. I get what you’re saying though. You’re trying to draw a distinction between statistical methods and symbolic methods. Someday we will have an algorithm which uses statistical methods that can match human performance on most cognitive tasks, and it won’t look or act like a brain. In some sense that’s disappointing. We can build supersonic jets without fully understanding how birds fly.
- antirez 2y agoLet's see that Turing machines can approximate the execution of NN :) That's why there are issues related to numerical precision, but the contrary is also true indeed, NNs can discover and use similar techniques used by traditional algorithms. However: the two remain two different methods to do computations, and probably it's not just by chance that many things we can't do algorithmically, we can do with NNs, what I mean is that this is not just related to the fact that NNs discover complex algorithms via gradient descent, but also that the computational model of NNs is more adapt to solving certain tasks. So the inference algorithm of NNs (doing multiplications and other batch transformations) is just needed for standard computers to approximate the NN computational model. You can do this analogically, and nobody would claim much (maybe?) it's running an algorithm. Or that brains themselves are algorithms.
- deleted 2y ago[deleted]
- zeroonetwothree 2y agoWe don’t have evidence that a TM can simulate a brain. But we know for a fact that it can execute a NN.
- necovek 2y agoComputers can execute precise computations, it's just not efficient (and it's very much slow). NNs are exactly what "computers" are good for and we've been using since their inception: doing lots of computations quickly. "Analog neural networks" (brains) work much differently from what are "neural networks" in computing, and we have no understanding of their operation to claim they are or aren't algorithmic. But computing NNs are simply implementations of an algorithm. Edit: upon further rereading, it seems you equate "neural networks" with brain-like operation. But brain was an inspiration for NNs, they are not an "approximation" of it.
- deleted 2y ago[deleted]
- mvkel 2y ago> continuous So, steps?
- necovek 2y ago"Continuous" would imply infinitely small steps, and as such, would certainly be used as a differentiator (differential? ;) between larger discrete stepped approach. In essence, infinite calculus provides a link between "steps" and continuous, but those are different things indeed.
- necovek 2y agoI would say you are right that function is not an algorithm, but it is an implementation of an algorithm. Is that your point? If so, I've long learned to accept imprecise language as long as the message can be reasonably extracted from it.
- CooCooCaCha 2y agoEach layer of the network is like a step, and each token prediction is a repeat of those layers with the previous output fed back into it. So you have steps and a memory.
- benlivengood 2y agoDeterministic (ieee 754 floats), terminates on all inputs, correctness (produces loss < X on N training/test inputs) At most you can argue that there isn't a useful bounded loss on every possible input, but it turns out that humans don't achieve useful bounded loss on identifying arbitrary sets of pixels as a cat or whatever, either. Most problems NNs are aimed at are qualitative or probabilistic where provable bounds are less useful than Nth-percentile performance on real-world data.
- KeplerBoy 2y agoRunning inference on a model certainly is a algorithm.
- drdeca 2y agoHow do you define "algorithm"? I suspect it is a definition I would find somewhat unusual. Not to say that I strictly disagree, but only because to my mind "neural net" suggests something a bit more concrete than "algorithm", so I might instead say that an artificial neural net is an implementation of an algorithm, rather than or something like that. But, to my mind, something of the form "Train a neural network with an architecture generally like [blah], with a training method+data like [bleh], and save the result. Then, when inputs are received, run them through the NN in such-and-such way." would constitute an algorithm.
- necovek 2y agoNN is a very wide term applied in different contexts. When a NN is trained, it produces a set of parameters that basically define an algorithm to do inference with: it's a very big one though. We also call that a NN (the joy of natural language).
- scotty79 2y agoStill it's comparing average human level performance with best AI performance. Examples of things o3 failed at are insanely easy for humans.
- FrustratedMonky 2y agoThere are things Chimps do easily that humans fail at, and vice/versa of course. There are blind spots, doesn't take away from 'general'.
- deleted 2y ago[deleted]
- Matumio 2y agoWe can't agree whether Portia spiders are intelligent or just have very advanced instincts. How will we ever agree about what human intelligence is, or how to separate it from cultural knowledge? If that even makes sense.
- FrustratedMonky 2y agoI guess my point is more, if we can't decide about Portia Spiders or Chimps, then how can we be so certain about AI. So offering up Portia and Chimps as counter examples.
- noobermin 2y agoThe downvotes should tell you, this is a decided "hype" result. Don't poo poo it, that's not allowed on AI slop posts on HN.
- FrustratedMonky 2y agoYeah, I didn't realize Chimp studies, or neuroscience were out of vogue. Even in tech, people form strong 'beliefs' around what they think is happening.
- 2y ago
- phillipcarter 2y agoAs excited as I am by this, I still feel like this is still just a small approximation of a small chunk of human reasoning ability at large. o3 (and whatever comes next) feels to me like it will head down the path of being a reasoning coprocessor for various tasks. But, still, this is incredibly impressive.
- qt31415926 2y agoWhich parts of reasoning do you think is missing? I do feel like it covers a lot of 'reasoning' ground despite its on the surface simplicity
- phillipcarter 2y agoI think it's hard to enumerate the unknown, but I'd personally love to see how models like this perform on things like word problems where you introduce red herrings. Right now, LLMs at large tend to struggle mightily to understand when some of the given information is not only irrelevant, but may explicitly serve to distract from the real problem.
- KaoruAoiShiho 2y agoo1 already fixed the red herrings...
- zmgsabst 2y agoThat’s not inability to reason though, that’s having a social context. Humans also don’t tend to operate in a rigorously logical mode and understand that math word problems are an exception where the language may be adversarial: they’re trained for that special context in school. If you tell the LLM that social context, eg that language may be deceptive, their “mistakes” disappear. What you’re actually measuring is the LLM defaults to assuming you misspoke trying to include relevant information rather than that you were trying to trick it — which is the social context you’d expect when trained on general chat interactions. Establishing context in psychology is hard.
- 2y ago
- ALittleLight 2y agoIt's not saturated. 85% is average human performance, not "best human" performance. There is still room for the model to go up to 100% on this eval.
- cryptoegorophy 2y agoWhat’s interesting is it might be very close to human intelligence than some “alien” intelligence, because after all it is a LLM and trained on human made text, which kind of represents human intelligence.
- hammock 2y agoIn that vein, perhaps the delta between o3 @ 87.5% and Human @ 85% represents a deficit in the ability of text to communicate human reasoning. In other words, it's possible humans can reason better than o3, but cannot articulate that reasoning as well through text - only in our heads, or through some alternative medium.
- 85392_school 2y agoI wonder how much of an effect amount of time to answer has on human performance.
- yunwal 2y agoYeah, this is sort of meaningless without some idea of cost or consequences of a wrong answer. One of the nice things about working with a competent human is being able to tell them "all of our jobs are on the line" and knowing with certainty that they'll come to a good answer.
- unsupp0rted 2y agoIt's possible humans reason better through text than not through text, so these models, having been trained on text, should be able to out-reason any person who's not currently sitting down to write.
- hamburga 2y agoAgreed. I think what really makes them alien is everything else about them besides intelligence. Namely, no emotional/physiological grounding in empathy, shame, pride, and love (on the positive side) or hatred (negative side).
- 6gvONxR4sf7o 2y agoHuman performance is much closer to 100% on this, depending on your human. It's easy to miss the dot in the corner of the headline graph in TFA that says "STEM grad."
- tim333 2y agoA fair comparison might be average human. The average human isn't a STEM grad. It seems STEM grad approximately equals an IQ of 130. https://www.accommodationforstudents.com/student-blog/the-subjects-with-the-highest-iqs https://www.accommodationforstudents.com/student-blog/the-su... From a post elsewhere the scores on ARC-AGI-PUB are approx average human 64%, o3 87%. https://news.ycombinator.com/item?id=42474659 https://news.ycombinator.com/item?id=42474659 Though also elsewhere, o3 seems very expensive to operate. You could probably hire a PhD researcher for cheaper.
- jeremyjh 2y agoWhy would an average human be more fair than a trained human? The model is trained.
- hypoxia 2y agoIt actually beats the human average by a wide margin: - 64.2% for humans vs. 82.8%+ for o3. ... Private Eval: - 85%: threshold for winning the prize [1] Semi-Private Eval: - 87.5%: o3 (unlimited compute) [2] - 75.7%: o3 (limited compute) [2] Public Eval: - 91.5%: o3 (unlimited compute) [2] - 82.8%: o3 (limited compute) [2] - 64.2%: human average (Mechanical Turk) [1] [3] Public Training: - 76.2%: human average (Mechanical Turk) [1] [3] ... References: [1] https://arcprize.org/guide https://arcprize.org/guide [2] https://arcprize.org/blog/oai-o3-pub-breakthrough https://arcprize.org/blog/oai-o3-pub-breakthrough [3] https://arxiv.org/abs/2409.01374 https://arxiv.org/abs/2409.01374
- usaar333 2y agoSuper human isn't beating rando mech turk. Their post has stem grad at nearly 100%
- tripletao 2y agoThis is correct. It's easy to get arbitrarily bad results on Mechanical Turk, since without any quality control people will just click as fast as they can to get paid (or bot it and get paid even faster). So in practice, there's always some kind of quality control. Stricter quality control will improve your results, and the right amount of quality control is subjective. This makes any assessment of human quality meaningless without explanation of how those humans were selected and incentivized. Chollet is careful to provide that, but many posters here are not. In any case, the ensemble of task-specific, low-compute Kaggle solutions is reportedly also super-Turk, at 81%. I don't think anyone would call that AGI, since it's not general; but if the "(tuned)" in the figure means o3 was tuned specifically for these tasks, that's not obviously general either.
- dyauspitr 2y agoI’ll believe it when the AI can earn money on its own. I obviously don’t mean someone paying a subscription to use the AI I mean, letting the AI lose on the Internet with only the goal of making money and putting it into a bank account.
- hamburga 2y agoDo trading bots count?
- 1659447091 2y agoNo, the AI would have to start from zero and reason it's way to making itself money online, such as the humans who were first in their online field of interest (e-commerce, scams, ads etc from the 80's and 90's) when there was no guidance, only general human intelligence that could reason their way into money making opportunities and reason their way into making it work.
- concordDance 2y agoI don't think humans ever do that. They research/read and ask other humans.
- 1659447091 2y agoWhich AI already has stored in spades, even more so since people in the 80's 90's weren't working with the information available today. The AI is free to research and read all the information stored from other humans as well, just like the humans who reasoned their way into money making opportunities--only with vastly more information now, talk about an advantage. But is it intelligent enough do so without a human giving direct/step-by-step instructions; the way humans figure it out?
- creer 2y agoYou don't think there are already plenty of attempts out there? When someone is "disinterested enough" to publish though, note the obvious way to launch a new fund or advisor with a good track record: crank out a pile of them, run them one or two years, discard the many losers and publish the one or two top winners. I.E. first you should be suspicious of why it's being published, then of how selected that result is.
- lastdong 2y agoCurious about how many tests were performed. Did it consistently manage to successfully solve many of these types of problems?
- dmead 2y agoThis is so strange. people think that an llm trained on programming questions and docs can do mundane tasks like this means intelligent? Come on. It really calls into question two things. 1. You don't know what you're talking about about. 2. You have a perverse incentive to believe this such that you will preach it to others and elevate some job salary range or stock. Either way, not a good look.
- javaunsafe2019 2y agoThis