9 ms·
One fundamental challenge to me is that if each training run because more and more expensive, the time it takes it to learn what works/doesn't work widens. Half
by t_serpico 2y ago
One fundamental challenge to me is that if each training run because more and more expensive, the time it takes it to learn what works/doesn't work widens. Half a billion dollars for training a model is already nuts, but if it takes 100 iterations to perfect it, you've cumulatively spent 50 billion dollars... Smaller models may actually be where rapid innovation continues simply because of tighter feedback loops. O3 may be an example of this.
- ramesh31 2y agoBut if the scaling law holds true, more dollars should at some point translate into AGI, which is priceless. We haven't reached the limits yet of that hypothesis.
- threeseed 2y agoa) There is evidence e.g. private data deals that we are starting to hit the limitations of what data is available. b) There is no evidence that LLMs are the roadmap to AGI. c) Continued investment hinges on their being a large enough cohort of startups that can leverage LLMs to generate outsized returns. There is no evidence yet this is the case.
- ComplexSystems 2y ago"There is no evidence that LLMs are the roadmap to AGI." - There's plenty of evidence. What do you think the last few years have been all about? Hell, GPT-4 would already have qualified as AGI about a decade ago.
- idiotsecant 2y agoHave you ever heard of a local maxima? You don't get an attack helicopter by breeding stronger and stronger falcons.
- lolinder 2y agoFor an industry that spun off of a research field that basically revolves around recursive descent in one form or another, there's a pretty silly amount of willful ignorance about the basic principles of how learning and progress happens. The default assumption should be that this is a local maximum, with evidence required to demonstrate that it's not. But the hype artists want us all to take the inevitability of LLMs for granted—"See the slope? Slopes lead up! All we have to do is climb the slope and we'll get to the moon! If you can't see that you're obviously stupid or have your head in the sand!"
- zmgsabst 2y agoYou’re implicitly assuming only a global maximum will lead to useful AI. There might be many local maxima that cross the useful AI or even AGI threshold.
- eru 2y agoAnd we aren't even at a local maximum. There's still plenty of incremental upwards progress to be made.
- lolinder 2y agoI never said anything about usefulness, and it's frustrating that every time I criticize AGI hype people move the goalposts and say "but it'll still be useful!" I use GitHub Copilot every day. We already have useful "AI". That doesn't mean that the whole thing isn't super overhyped.
- int_19h 2y agoSo far we haven't even climbed this slope to the top yet. Why don't we start there and see if it's high enough or not first? If it's not, at the very least we can see what's on the other side, and pick the next slope to climb. Or we can just stay here and do nothing.
- gwervc 2y agoNo, GPT-4 would have been classified as it is today: a (good) generator of natural language. While this is a hard classical NLP task, it's a far cry from intelligence.
- falcor84 2y agoGPT-4 is a good generator of natural language in the same sense that Google is a good generator of ip packets.
- coldtea 2y ago>What do you think the last few years have been all about? Next token language-based predictors with no more intelligence than brute force GIGO which parrot existing human intelligence captured as text/audio and fed in the form of input data. 4o agrees: "What you are describing is a language model or next-token predictor that operates solely as a computational system without inherent intelligence or understanding. The phrase captures the essence of generative AI models, like GPT, which rely on statistical and probabilistic methods to predict the next piece of text based on patterns in the data they’ve been trained on"
- thrwthsnw 2y agoEverything you said is parroting data you’ve trained on, two thirds of it is actual copy paste
- coldtea 2y ago>Everything you said is parroting data you’ve trained on "Just like" an LLM, yeah sure... Like how the brain was "just like" a hydraulic system (early industrial era), like a clockwork with gears and differentiation (mechanical engineering), "just like" an electric circuit (Edison's time), "just like" a computer CPU (21st century), and so on... You're just assuming what you should prove
- mrbungie 2y agoHe probably didn't need petabytes of reddit posts and millions of gpu-hours to parrot that though. I still don't buy the "we do the same as LLMs" discourse. Of course one could hypothesize the human brain language center may have some similarities to LLMs, but the differences in resource usage and how those resources are used to train humans and LLMs are remarkable and may indicate otherwise.
- shwouchk 2y agoNot text, he had petabytes of video, audio, and other sensory inputs. Heck, a baby sees petabytes of video before first word is spoken And he probably cant quote Shakespeare as well ;)
- n144q 2y ago> GPT-4 would already have qualified as AGI about a decade ago. Did you just make that up?
- OtomotO 2y agoThe last four years? ELIZA 2.0
- aantix 2y agoHave we really hit the wall? Do they use GPS based data? Feels like there’s data all around us. Sure they’ve hit the wall with obvious conversations and blog articles that humans produced, but data is a by product of our environment. Surely there’s more. Tons more.
- threeseed 2y agoWe also could just measure the background noise of the universe and produce unlimited data. But just like GPS data it isn't suited for LLMs given that you know it has no relevance what so ever to language.
- aantix 2y agoYou’re thinking of language in the strictest of sense. GPS data as it relates to location names, people, cultures, path finding.
- eru 2y agoWhat does culture and names and people have to do with the Global Position System? You are right that we can have lots more data, if you are willing to consider other modalities. But that's not 'GPS'. Unless you are using an idiosyncratic definition of GPS?
- eru 2y agoIgnoring the confusion about 'GPS' for a moment: there's lots and lots of other data that could be used for training AI systems. But, you need to go multi-modal for that; and you need to find data that's somewhat useful, not just random fluctuations like the CMB. So eg you could use YouTube videos, or even just point webcams at the real world. That might be able to give your AI a grounding in everyday physics? There's also lots of program code you can train your AI on. Not so much the code itself, because compared to the world's total text (that we are running out of), the world's total human written code is relatively small. But you can generate new code and make it useful for training, by also having the AI predict what happens when you (compile and) run the code. A bit like self-playing for improving AlphaGo.
- thrwthsnw 2y agoPrivate data is 90% garbage too
- eru 2y ago> c) Continued investment hinges on their being a large enough cohort of startups that can leverage LLMs to generate outsized returns. There is no evidence yet this is the case. Why does it have to be startups? And why does it have to be LLMs? Btw, we might be running out of text data. But there's lots and lots more data you can have (and generate), if you are willing to consider other modalities. You can also get a bit further with text data by using it for multiple epochs, like we used to do in the past. (But that only really gives you at best an order of magnitude. I read some paper that the returns diminish drastically after four epochs.)
- zifpanachr23 2y agoI agree, these are good points.
- unshavedyak 2y ago> which is priceless This also isn't true. It'll clearly have a price to run. Even if it's very intelligent, if the price to run it is too high it'll just be a 24/7 intelligent person that few can afford to talk to. No?
- pbhjpbhj 2y agoComputers will be the size of data centres, they'll be so expensive we'll queue up jobs to run on them days in advance, each taking our turn... history echoes into the future...
- unshavedyak 2y agoYea, and those statements were true. For a time. If you want to say "AGI will be priceless some unknown time into the future" then i'd be on board lol. But to imply it'll be immediately priceless? As in no cost spent today wouldn't be immediately rewarded once AGI exists? Nonsense. Maybe if it was _extremely_ intelligent and it's ROI would be all the drugs it would instantly discover or w/e. But lets not imply that General Intelligence requires infinitely knowing. So at best we're talking about an AI that is likely close to human level intelligence. Which is cool, because we have 7+ billion of those things. This isn't an argument against it. Just to say that AGI isn't "priceless" in the implementation we'd likely see out of the gate.
- dkobia 2y agoAGI is the Sisyphean task of our age. We’ll push this boulder up the mountain because we have to, even if it kills us.
- h0l0cube 2y agoThere's no doubt been progress on the way to AGI, but ultimately it's still a search problem, and one that will rely on human ingenuity at least until we solve it. LLMs are such a vast improvement in showing intelligent-like behavior that we've become tantalized by it. So now we're possibly focusing our search in the wrong place for the next innovation on the path to AGI. Otherwise, it's just a lack of compute, and then we just have to wait for the capacity to catch up.
- idiotsecant 2y agoAnd when we get it there, it kills us.
- hex3 2y ago[dead]
- anothernewdude 2y ago[flagged]
- falcor84 2y agoIt seems to me that given how AI is likely to continuously increase capitalism's efficiency, your argument actually supports the claim you're trying to dispute.
- sourcepluck 2y agoCapitalism is not efficient, it's grabby. Read Bullshit Jobs. Moreover, capitalism isn't interested in efficiency, it's interested in grabbing more stuff. It's relatively effiicient at centralising power and resources into the pockets of shareholders, but that's probably not what you meant. I think this is borne out even moreso in recent years, as environmental degradation continues, and we watch as capitalist systems are unable to do anything but continue to efficiently funnel money into the pockets of shareholders. The word "efficient" can only plausibly be applied to overly simplified models in fantastical economic theories which don't reflect reality. The kind of AI offered by companies like OpenAI may very well be an effective tool at grabbing more stuff though, sure. Or, rather, at convincing everyone they simply must move to this new area, that they control, effectively grabbing that newly created space.
- fny 2y agoO3 is not a smaller model. It's an iterative GPT of sorts with the magic dust of reinforcement learning.
- falcor84 2y agoI'm pretty sure that the parent implied that o3 is smaller in comparison to gpt5
- deleted 2y ago[deleted]
- dyauspitr 2y agoUntil you get to a point where the LLM is smart enough to look at real world data streams and prune its own training set out of it. At that point it will self improve itself to AGI.
- bloodyplonker22 2y agoI am working at an AI company that is not OpenAI. We have found ways to modularize training so we can test on narrower sets before training is "completely done". That said, I am sure there are plenty of ways others are innovating to solve the long training time problem.
- gerdesj 2y agoPerhaps the real issue is that learning takes time and that there may not be a shortcut. I'll grant you that argument's analogue was complete wank when comparing say the horse and cart to a modern car. However, we are not comparing cars to horses but computers to a human. I do want "AI" to work. I am not a luddite. The current efforts that I've tried are not very good. On the surface they offer a lot but very quickly the lustre comes off very quickly. (1) How often do you find yourself arguing with someone about a "fact"? Your fact may be fiction for someone else. (2) LLMs cannot reason A next token guesser does not think. I wish you all the best. Rome was not burned down within a day! I can sit down with you and discuss ideas about what constitutes truth and cobblers (rubbish/false). I have indicated via parenthesis (brackets in en_GB) another way to describe something and you will probably get that but I doubt that your programme will.
- icpmacdo 2y agoThis is literally just the scaling laws, "Scaling laws predict the loss of a target machine learning model by extrapolating from easier-to-train models with fewer parameters or smaller training sets. This provides an efficient way for practitioners and researchers alike to compare pretraining decisions involving optimizers, datasets, and model architectures" https://arxiv.org/html/2410.11840v1#:~:text=Scaling%20laws%20predict%20the%20loss,%2C%20datasets%2C%20and%20model%20architectures https://arxiv.org/html/2410.11840v1#:~:text=Scaling%20laws%2....
- cma 2y ago>the time it takes it to learn what works/doesn't work widens. From the raw scaling laws we already knew that a new base model may peter out in this run or the next with some amount of uncertainty--"the intersection point is sensitive to the precise power-law parameters": https://gwern.net/doc/ai/nn/transformer/gpt/2020-kaplan-figure15-projectingscaling.png https://gwern.net/doc/ai/nn/transformer/gpt/2020-kaplan-figu... Later graph gpt-3 got to here: https://gwern.net/doc/ai/nn/transformer/gpt/2020-brown-figure31-gpt3scaling.png https://gwern.net/doc/ai/nn/transformer/gpt/2020-brown-figur... https://gwern.net/scaling-hypothesis https://gwern.net/scaling-hypothesis
- merizian 2y agoBecause of mup [0] and scaling laws, you can test ideas empirically on smaller models, with some confidence they will transfer to the larger model. [0] https://arxiv.org/abs/2203.03466 https://arxiv.org/abs/2203.03466
- ciconia 2y agoWhen you think about it it's astounding how much energy this technology consumes versus a human brain which runs at ~20W [1]. [1] https://hypertextbook.com/facts/2001/JacquelineLing.shtml https://hypertextbook.com/facts/2001/JacquelineLing.shtml
- concerndc1tizen 2y ago20w for 20 years to answer questions slowly and error-prone at the level of a 30B model. An additional 10 years with highly trained supervision and the brain might start contributing original work.
- vbezhenar 2y agoMultiply that by billion, because only very few individuals of entire populations can contribute original work.
- rwyinuse 2y agoAnd yet that 20w brain can make me a sandwich and bring it to me, while state of the art AI models will fail that task. Until we get major advances in robotics and models designed to control them, true AGI will be nowhere near.
- sekai 2y ago> Until we get major advances in robotics and models designed to control them, true AGI will be nowhere near. AGI has nothing to do with robotics, if AGI is achieved it will help push robotics and every single scientific field further with progression never seen before, imagine a million AGIs running in parallel focused on a single field.
- onlyrealcuzzo 2y agoWe already have that. It's called civilization. Maybe you mean quadrillions of AGIs?
- 2y ago
- soheil 2y agoIt's like saying bacteria reproduction is way faster than humans so that's where we should be looking for the next breakthroughs.