7 ms·
Exciting to see this so soon after Anthropic's "Mapping the Mind of a Large Language Model" (under 3 weeks). I find these efforts really exciting; it is still c
by andreyk 2y ago
Exciting to see this so soon after Anthropic's "Mapping the Mind of a Large Language Model" (under 3 weeks). I find these efforts really exciting; it is still common to hear people say "we have no idea how LLMs / Deep Learning works", but that is really a gross generalization as stuff like this shows.
Wonder if this was a bit rushed out in response to Anthropic's release (as well as the departure of Jan Leike from OpenAI)... the paper link doesn't even go to Arxiv, and the analysis is not nearly as deep. Though who knows, might be unrelated.
- jerrygenser 2y ago> but that is really a gross generalization as stuff like this shows. I think this research actually still reinforces that we still have very little understanding of the internals. The blog post also reiterates that this is early work with many limitations.
- thegrim33 2y agoFrom the article: "We currently don't understand how to make sense of the neural activity within language models." "Unlike with most human creations, we don’t really understand the inner workings of neural networks." "The [..] networks are not well understood and cannot be easily decomposed into identifiable parts" "[..] the neural activations inside a language model activate with unpredictable patterns, seemingly representing many concepts simultaneously" "Learning a large number of sparse features is challenging, and past work has not been shown to scale well." etc., etc., etc. People say we don't (currently) know why they output what they output, because .. as the article clearly states, we don't.
- surfingdino 2y agoNot holding my breath for that hallucinated cure for cancer then.
- ben_w 2y agoLLMs aren't the only kind of AI, just one of the two current shiny kinds. If a "cure for cancer" (cancer is not just one disease so, unfortunately, that's not even as coherent a request as we'd all like it to be) is what you're hoping for, look instead at the stuff like AlphaFold etc.: https://en.wikipedia.org/wiki/AlphaFold https://en.wikipedia.org/wiki/AlphaFold I don't know how to tell where real science ends and PR bluster begins in such models, though I can say that the closest I've heard to a word against it is "sure, but we've got other things besides protein folding to solve", which is a good sign. (I assume AlphaFold is also a mysterious black box, and that tools such as the one under discussion may help us demystify it too).
- surfingdino 2y ago[flagged]
- wg0 2y ago[flagged]
- bigyikes 2y agoIs your argument that because AI can’t currently do the arbitrary things you wish it would do, it is therefore bullshit? This perspective discounts two important things: 1. All the things it can obviously do very well today 2. Future advancements to the tech (billions are pouring in, but this takes time to manifest in prod) I’m trying not to be one of the “guys” you’re talking about, but I just can’t comprehend your take. Do you not recognize that there is utility to current models? What makes it all bullshit?
- vitus 2y ago> 1. All the things it can obviously do very well today I'm curious what those things are. At least to me, it isn't obvious that LLMs solve any of their many applications from the past year "very well". I worry about failures (hallucinations, misinterpretation of prompts, regurgitation of incorrect facts, violation of copyright, and more). I don't have a good sense of when they fail, how often this happens, or how to identify these failures in scenarios where I'm not a domain expert. But maybe some subset of these are solved problems, or at least problems that are actively being worked on for the next generation of models. Yes, there are many other kinds of AI. Stockfish is better at chess than any human. But when you start talking about emergent behavior from machine learning, the failure modes are much harder to reason about.
- TrainedMonkey 2y agoI read this as "we have not built up tools / math to understand neural networks as they are new and exciting" and not as "neural networks are magical and complex and not understandable because we are meddling with something we cannot control". A good example would be planes - it took a long while to develop mathematical models that could be used to model behavior. Meanwhile practical experimentation developed decent rule of thumb for what worked / did not work. So I don't think it's fair to say that "we don't" (know how neural networks work), we don't have math / models yet that can explain/model their behavior...
- gradus_ad 2y agoChaotic nonlinear dynamics have been an object of mathematical research for a very long time and we have built up good mathematical tools to work with them, but in spite of that turbulent flow and similar phenomena (brains/LLM's) remain poorly understood. The problem is that the macro and micro dynamics of complex systems are intimately linked, making for non-stationary non-ergodic behavior that cannot be reduced to a few principles upon which we can build a model or extrapolate a body of knowledge. We simply cannot understand complex systems because they cannot be "reduced". They are what they are, unique and unprincipled in every moment (hey, like people!).
- deleted 2y ago[deleted]
- baxtr 2y agoPhysicists would probably argue that the system might be understood but that we don’t have the model for it yet. Many natural phenomena look chaotic at best without a model. Once you have a model things fall into place and everything starts looking orderly. Maybe it cannot be reduced. But maybe we are just observing the peripherals without understanding the inner workings.
- MrsPeaches 2y agoIf I can speak in aphorisms, Creation is downhill, analysis is uphill. Profound ideas often seem simple once understood.
- submeta 2y agoScary actually. Because how can we asses the risks when we don’t know what the system is capabale of doing.
- ein0p 2y agoWe know exactly what the system is capable of doing. It’s capable of outputting tokens which can then be converted into text.
- reducesuffering 2y agoAnd social media manipulation is just registers and bytes, wait no, sand and electrons.
- isaacremuant 2y agoJust because you can do something with technology doesn't mean the problem is technology itself. It's like newspapers. Printing them I technology and allows all kind of things. If you're of the authoritarian mindset, you'll want to control it all out of some stated fear, but you can do that for everything.
- friendzis 2y agoIf you can do something with technology, that something is part of risk assessment. Authoritarianism is irrelevant, that's engineering.
- ben_w 2y agoWhich is so broad as to be unhelpful. We also know that petroleum mixed with air may be combusted to release energy; we needed to characterise this much better in order for the motor car to be distinguishable from a fuel-air bomb.
- ein0p 2y agoAnd that's exactly my point. Regulating the underlying tech is utterly pointless in this case - it's utterly harmless by itself.
- realPtolemy 2y agoCould there also be a “legal hedging” reason for why you would release a paper like this? By reaffirming that “we don’t know how this works, nobody does” it’s easier to avoid being charged with copyright infringement from various actors/data sources that have sued them.
- ben_w 2y agoI'd be surprised if doing so had any impact on the lawsuits, but I'm not a lawyer.
- dimitrios1 2y ago"I'm sorry officer, I didn't know I couldn't do that"
- icandoit 2y agoIf you know how it works, you can make it better,faster,cheaper. Without the 300k starting salaries. I imagine that is a stronger incentive. It's the users of the LLMs that want to launder repsonsibility behind "computer said no".
- deleted 2y ago[deleted]
- imjonse 2y agoBoth Leike and Sutskever are still credited in the post.
- swyx 2y ago> Wonder if this was a bit rushed out in response to Anthropic's release too lazy to dig up source but some twitter sleuth found that the first commit to the project was 6 months ago likely all these guys went to the same metaphorical SF bars, it was in the water
- leogao 2y agoThis project has been in the works for about a year. The initial commit to the public repo was not really closely related to this project, it was part of the release of the Transformer debugger, and the repo was just reused for this release.
- swyx 2y agoha thank you Leo; i myself felt uneasy pointing out commit date based evidence and you just proved why. mild followup question: any alpha to be gained from training the same SAEs on two different generations of GPT4, eg GPT4 on march 2023 vs june 2023 vintage, whatever is most architecturally comparable, and diffing them. what would be your priors on what you’d find?
- nicce 2y agoVisualizer was added 18 hours ago: https://github.com/openai/sparse_autoencoder/commit/764586aedaeb469c3adea166859259a437b8353f https://github.com/openai/sparse_autoencoder/commit/764586ae...
- pininja 2y agoIt’s hard to believe it was written overnight.. this seems more like a public stable dump of what they’ve been working on without saying when they started. Some clues could come from looking at when all the deps it uses were released. They’re also calling this version 0.1.67, though I’m not sure that means anything either.
- szvsw 2y ago> likely all these guys went to the same metaphorical SF bars, it was in the water It also is coming from a long lineage of thought no? For instance, one of the things often thought early in an ML course is the notion that “early layers respond to/generate general information/patterns, and deeper layers respond to/generate more detailed/complex patterns/information.” That is obviously an overly broad and vague statement but it is a useful intuition and can be backed up by doing some various inspection of eg what maximally activates some convolution filters. So already there is a notion that there is some sort of spatial structure to how semantics are processed and represented in a neural network (even if in a totally different context, as in image processing mentioned above), where “spatial” here is used to refer to different regions of the network. Even more simply, in fact as simple as you can get: with linear regression, the most interpretable model you can get- you have a clear notion that different parameter groups of the model respond to different “concepts” (where a concept is taken to be whatever the variables associated with a given subset of coefficients represent). In some sense, at least in a high-level/intuitive reading of the new research coming out of Anthropic and OpenAI, I think the current research is just a natural extension of these ideas, albeit in a much more complicated context and massive scale. Somebody else, please correct me if you think my reading is incorrect!!
- 3abiton 2y agoBut even with current efforts so far, I don't think we have an understanding of how/why these emergent capabilities are formed. LLMs are still a black box as ever.
- choppaface 2y agoThe Deep Visualization Toolbox from nearly 10 years ago is solid precedent for understanding deep models, albeit much smaller models than LLMs. It’s hard to say OpenAI’s “visualization” released today is nearly as effective. It could be that GPT-4 is much harder to instrument. https://github.com/yosinski/deep-visualization-toolbox https://github.com/yosinski/deep-visualization-toolbox
- leogao 2y agoWe were planning to release the paper around this time independent of the other events you mention. I think it is still predominantly accurate to say that we have no idea how LLMs work. SAEs might eventually change that, but there's still a long way to go.
- joaquincabezas 2y agoit makes sense that the leaders are building around similar ideas in parallel, for me it's a healthy sign
- darby_nine 2y ago> Mapping the Mind of a Large Language Model The fact that a paper is implying a LLM has a mind doesn't exactly bode well for the people who wrote it, not to mention the continued meaningless babbling about "safety". It'd also be nice if they could show their work so we could replicate it. Still, not shabby for an ad!
- castigatio 2y agoWell - what is a mind exactly? We don't really have a good definition for a human mind. Not sure we should be claiming domain over the term. It's not a terrible shorthand for discussing something that reads and responds as if it had some kind of mind - whether technically true or not (which we honestly don't know).
- darby_nine 2y ago> It's not a terrible shorthand for discussing something that reads and responds as if it had some kind of mind I really don't see it like that—it has very little memory, it has no ability to introspect before "choosing" what to say, no awareness of the concept of the coherency of statements (i.e. whether or not it's saying things that directly contradict its training), seems to have little sense of non-pattern-driven computation beyond what token patterns can encode at a surface level (e.g. of course it knows 1 + 1 = 2, but does it recognize odd notation/can it recognize and analyze arbitrary statements? of course not). I fully grant it is compelling evidence we can replicate many brain-like processes with software neural nets, but that's an entirely different thing than raising it to a level of thought or consciousness or self-awareness (which I argue is necessary in order to appropriately issue coherent statements, as perspective is a necessary thing to address even when attempting to make factual statements), but it strikes me as a lot closer to an analogy for a potential constituent component of a mind rather than a mind per se.
- realPtolemy 2y agoIndeed, and the very last section about how they’ve now “open sourced” this research is also a bit vague. They’ve shared their research methodology and findings… But isn’t that obligatory when writing a public paper?
- lanceflt 2y agohttps://github.com/openai/sparse_autoencoder https://github.com/openai/sparse_autoencoder They actually open sourced it, for GPT-2 which is an open model.
- realPtolemy 2y agoThanks, I must have read through the document to hastily.
- throw46365 2y ago> that is really a gross generalization It's really not though, and on multiple levels. At the shit-tier level, the majority of people building applications on this technology are projecting abilities onto it that even they can't really demonstrate it has in a reliable way. At the inventor level, the people who make it are dependent on projecting the idea that magic will happen when they have more compute. At every level, the products are so far ahead of the knowledge that it's actually unethical.
- deleted 2y ago[deleted]