8 ms·
I must say, understanding how transformers work is arguably the most important research problem in history, assuming that AGI can be achieved by just scaling up
by xcodevn 3y ago
I must say, understanding how transformers work is arguably the most important research problem in history, assuming that AGI can be achieved by just scaling up current LLM models on text, video, audio, etc.
- kindking 3y agoThat is a very big assumption.
- ImHereToVote 3y agoWhich might turn out to be correct. Might be wrong also. We have no priors to AGI developing. Only NGI, and we know preciously little about how to achieve NGI too, except the bedroom way.
- joaogui1 3y agoIn vitro fertilization too!
- furyofantares 3y agoWe have a lot of priors - everything we've ever done has not produced AGI. Maybe scaling transformers is the best way forward. I'm hopeful. But it's a big assumption that it will produce AGI.
- kromem 3y agoIt really depends on the definition. "Better than the average human at most profitable tasks" is a much lower bar than most people on HN might think. I have vendors who instead of filling out a web form which remembers their inputs and eventually even fills everything out for them instead print it out and fax it back in. We're probably only about 2-3 years away from transformers being self-optimizing enough in prompts and evaluations to outpace the average worker in most tasks in most roles. (It won't necessarily be that much cheaper after the multiple passes and context windows required, and crucially probably won't be better at all tasks in most roles.) If you define AGI as "better than any human at profitable tasks" or "better than average at all tasks" then yes, we're a long ways off and transformers alone probably won't get us there.
- shafyy 3y ago> "Better than the average human at most profitable tasks" This is not the definition of AGI. You can't just make up a random definition to fit your argument, lol.
- exe34 3y agoNo that's the economically dominating definition. The philosophical one will happen much later or may never happen, but human society may change beyond recognition with the first one alone.
- hnben 3y ago"The philosophical one" seems to get updated with every new breakthrough. 20 years ago, GPT3 would have been considered AGI (or "strong AI", as we called it back then). https://en.wikipedia.org/wiki/Artificial_general_intelligence#History https://en.wikipedia.org/wiki/Artificial_general_intelligenc...
- exe34 3y agoDennett describes it as real magic. The magic that can be performed is not considered real magic (it's merely a trick of confidence), whereas real magic is that which couldn't possibly be done.
- falcor84 3y agoI'm actually all in on people making up new definitions for vague terms at the start of an argument as long as they're explicit about it. And I particularly like this one, which is much more clearly measurable. If you feel AGI is taken, maybe we should coin this one as APGI or something like that
- worldsayshi 3y agoI don't think the main intention was to define AGI but to zoom in on an interpretation of AGI that would provide enough value to be revolutionary.
- logicchains 3y agoWe already understand how transformers work: their architecture can learn to approximate a very large class of functions-on-sequences (specifically continuous sequence-to-sequence functions with compact support: https://arxiv.org/abs/1912.10077 https://arxiv.org/abs/1912.10077). It can do it more accurately than previous architectures like RNNs because transformers don't "forget" any information from prior items in the sequence. Training transformers to predict the next item in sequences eventually forces them to learn a function that approximates a world model (or at least a model of how the world behaves in the training text/data), and if they're large enough and trained with enough data then this world model is accurate enough for them to be useful for us. If you're asking for understanding the actual internal world model they develop, it's basically equivalent to trying to understand a human brain's internal world model by analysing its neurons and how they fire.
- theGnuMe 3y ago>Training transformers to predict the next item in sequences eventually forces them to learn a function that approximates a world model There is absolutely no proof for this statement.
- H8crilA 3y agoHave you used ChatGPT? It can do (at least) simple reasoning, for example simple spatial reasoning or simple human behavior reasoning.
- theGnuMe 3y agoI suggest you start here: https://ahtiahde.medium.com/limits-of-turing-machines-and-algorithmic-finger-prints-f33803ecbb43 https://ahtiahde.medium.com/limits-of-turing-machines-and-al... If you don't have a CS background, I would suggest reading the wikipedia entries referenced in the medium article as well.
- seydor 3y agoWe shouldn't assume that either of those tasks is impossible
- reexpressionist 3y ago[dead]
- infecto 3y agoI hope you are not part of the founding team but if you are, you truly are doing your startup a disservice. Sharing your startup/ideas is great but doing it in the form of an advertisement "underlying approach introduced in Reexpress as among the more significant results of the first quarter of the 21st century" is just weird.
- 12345hn6789 3y agoRe: this is an advertisement for a product by its own employees.
- xxs 3y ago>most important research problem in history That has to be some extremely narrowed version of all research that has happened (or will happen?)
- seydor 3y agoConsidering that knowing how knowing works is at the top of the ordo cognoscendi, it s not that narrow.
- resource_waste 3y ago>assuming that AGI can be achieved by just scaling up current LLM models on text, video, audio, etc. Is any sane person actually trying this? I can't imagine an LLM ever going AGI.
- otabdeveloper4 3y ago> the most important research problem in history Probably not. > assuming that AGI can be achieved by just scaling up current LLM models Lmao.
- xcodevn 3y ago> Probably not. Lmao.