6 ms·
Mamba Explained: The State Space Model Taking On Transformers
- fancyfredbot 3y agoSee also: https://jackcook.com/2024/02/23/mamba.html https://jackcook.com/2024/02/23/mamba.html
- Der_Einzige 3y agoFirst it was longformer, and linear attention models. Then it was RWKV and now it's Mamba. So many bombastic claims of improved architectural performance - and no open source models that beat the thing they purport to beat. The proof is always in the pudding, and these models will remain a curiosity for most until their weights are being benchmarked favorably on LLM leaderboards.
- digdugdirk 3y agoYes, that's technically accurate. But I prefer to think of the entire LLM space as a new scientific field that started when OpenAI released ChatGPT. In that context, all new research directions are valuable simply for the fact that they're expanding the foundation of the field. 5 years from now, who knows what the most effective models will use under the hood, but the more we can learn about them in general, the better.
- CityOfThrowaway 3y agoThe field of research here is far older than ChatGPT's release. Neural network research has been going on for at least 50 years. Most of the research that enabled ChatGPT was also already known. "Attention is all you need" was a 2017 paper. It still is a fast evolving field, but not one that just kicked off.
- lettergram 3y agolol I think in general, LLM research traces its origins back to all the standard deep learning techniques: NNs, CNNs, LSTMs, RNNs, etc. In 2018, with the release of transformers (via google) it enabled much more rapid training of models and more generalization with less data. 100% of the LLMs (as you’d probably thing of them)trace their origins to BERT. That said, my team was working with hundred million to low billions of parameter LSTMs & CNNs back in 2016-2017 that were comparable to some lighter weight LLMs today. In my opinion, the greatest strides in the space has less to do with the underlying architecture, and more to do with improved data formatting, accessibility and compute improvements.
- sigmoid10 3y agoTrue, but bear in mind the Mamba preprint is less than three months old. A lot of people are probably experimenting with these ideas right now and training a completely new, large foundation model with a different architecture will take a significant amount of time.
- imjonse 3y agoMost (all?) open-ish 7B+ models today are finetunes of proprietary/semi-closed/bigbudget LLMs. There is no such foundation model for Mamba yet.
- nickpsecurity 3y agoGPT3-176B cost $30 million dollars in compute plus millions in design, preprocessing, and operations. Then, it was able to perform as much better than prior architectures as it does today. You might want to include that in your challenge for competing models. Let’s rephrase it. If their architecture is superior, and they have $30 million dollars, and similar preparation for training, and similar operational teams during training, then we can see if they can beat the model they’re comparing themselves to. Except, the alternatives don’t have tens of millions of dollars with the best support teams. So, the proof you seek hasn’t had a chance to happen due to severe lack of resources. Hence, comparisons to GPT2 and small versions of GPT3. Even that might not be fair given the money and teams behind even small GPT3’s. Execution of the project is as critical for success as the model architecture.
- thecolorgreen 3y agoWhy doesn't Equation 1b use the h' defined in Equation 1a?
- atlacatl_sv 3y agoI believe h' is for the next state. y(t) is to predict the next word so it uses the current hidden state h(t).
- deleted 3y ago[deleted]
- koayon 3y agoHey! OP here Great question - h' in Equation 1a refers to the derivative of h with respect to time (t). This is a differential equation which we can solve mathematically when we have x in order to get a closed-form solution for h. We would then plug in that h (the hidden state) into equation 1b. In our case, we don't actually wait for a closed-form solution but instead compute the discrete representation (Equation 2) Hope that helps!
- deleted 3y ago[deleted]
- CrypticShift 3y ago> In other words, you can drag and drop downloaded states into your model, like literal plug-in cartridges The same could be said of "control vectors" [1]. Both ideas are still experimental, but is seems to me IINM that they could replace "system prompts" and "RAG" respectively. [1] https://news.ycombinator.com/item?id=39414532 https://news.ycombinator.com/item?id=39414532
- refulgentis 3y agoCan control vectors replace RAG? i.e. if I want the model to give me a summary of the news today, and the model was trained before today, can control vectors help?
- p1esk 3y agoNo technique can get you the news other than actually searching for and then parsing the published news.
- refulgentis 3y agoCan a control vector replace system prompts? i.e. can it do in-context learning without the context?
- jncfhnb 3y agoIt more or less is the same as a system prompt
- refulgentis 3y agoSo, no
- jncfhnb 3y agoSo, yes, but not in a meaningfully different form
- behnamoh 3y agoCan the low adoption of Mamba be attributed to what is being discussed today on HN (https://news.ycombinator.com/item?id=39491863 https://news.ycombinator.com/item?id=39491863)? Basically, Nvidia et al. don't want the AI research to move in a direction that requires less GPU compute, less training data, and less inference compute. Someone on HN (I don't remember the name) mentioned that the idea of deep learning is backed by big tech because it benefits them the most as they are the only players in town with huge amounts of data. If the AI community would find entirely different approaches to AGI (maybe not even learning), who do you think would suffer the most from the implications?
- p1esk 3y agoThis doesn’t make sense - there are literally thousands of academic AI research labs who are severely limited by compute resources. If anything could work better than transformers and require less compute they would be all over that.
- behnamoh 3y agoI guess the argument is that most AI research is supported by the big tech, and they have heavily invested in the deep learning approach. If the fundings were funneled to research groups working on alternative approaches, maybe we'd see the same amount of progress in AI only using another approach.
- kettleballroll 3y agoAs a member of the research community: that's nonsense. Like already pointed out: academic groups (who by no means are dependent on big tech) would jump all over that. Mamba has been out long enough that you'd already see tons of papers at arxiv showing mamba dominating transformers in all sorts of applications. But that's not happening, despite the ton of hype. That doesn't mean that mamba is nonsense. Just that it isn't the immediate transformer killer. It remains to be seen if something comes from it, eventually.
- 3y ago
- deleted 3y ago[deleted]
- imjonse 3y agoExplaining Mamba is a rite of passage, like the monad tutorials of yore.
- SkyMarshal 3y agoMamba is like a burrito...
- deleted 3y ago[deleted]
- kekebo 3y agoIt gets soggy and disintegrates when not consumed swiftly?
- hyperbovine 3y agoSimilar market share too.
- sja 3y agoOr Balks[0]: BALK RULES! IMPORTANT! 1. You can’t just be up there and just doin’ a balk like that. 1a. A balk is when you 1b. Okay well listen. A balk is when you balk the 1c. Let me start over 1c-a. The pitcher is not allowed to do a motion to the, uh, batter, that prohibits the batter from doing, you know, just trying to hit the ball. You can’t do that. 1c-b. Once the pitcher is in the stretch, he can’t be over here and say to the runner, like, “I’m gonna get ya! I’m gonna tag you out! You better watch your butt!” and then just be like he didn’t even do that. 1c-b(1). Like, if you’re about to pitch and then don’t pitch, you have to still pitch. You cannot not pitch. Does that make any sense? 1c-b(2). You gotta be, throwing motion of the ball, and then, until you just throw it. 1c-b(2)-a. Okay, well, you can have the ball up here, like this, but then there’s the balk you gotta think about. 1c-b(2)-b. Fairuza Balk hasn’t been in any movies in forever. I hope she wasn’t typecast as that racist lady in American History X. 1c-b(2)-b(i). Oh wait, she was in The Waterboy too! That would be even worse. 1c-b(2)-b(ii). “get in mah bellah” – Adam Water, “The Waterboy.” Haha, classic… 1c-b(3). Okay seriously though. A balk is when the pitcher makes a movement that, as determined by, when you do a move involving the baseball and field of 2. Do not do a balk please. [0]: https://justinbee.tumblr.com/post/15309101943/best-explanation-of-a-balk-ive-ever-seen/amp https://justinbee.tumblr.com/post/15309101943/best-explanati...
- AndrewKemendo 3y agoSomeone is going to re-invent Bellman's equations and call it Learnformer
- Straw 3y agoThe SSMs papers and blogs always have unnecessarily complicated explanations. At this point I almost wonder if its to hide how simple the underlying algorithms are, or to make them seem fancy. SSMs are doing exponentially weighted moving averages (EMA). That's it- to summarize the past, you take an EMA of a variable output at each time step. Mamba changes one key thing- instead of decaying the past by a fixed amount each step as in a constant-time EMA, we have another output which decides how much to forget, or equivalently, how much 'time' has passed since the last observation in our EMA. All of the matrix equations, continuous time, discretization, etc, will end up with a dynamic-forgetting EMA as I describe above. This also makes the benefits and limitations clear- finite state size, has to decide at a given layer what to forget before it sees the past at that layer.
- logicchains 3y agoAre there any fundamental differences between Mamba, Retnet and RWKV, or are they all variants of this same architecture?
- Straw 3y agoNo, all of these use the same fundamental architecture with minor tweaks, such as the dynamic gate for mamba or an outer product paramterization of the values for RWKV-v5
- pama 3y agoA dynamic gate is a pretty distinct feature from previous SSM architectures in my opinion. In a sense, the overall fundamental architecture of mamba is still that of the transformer but with attention replaced by an SSM with dynamic gating. All of deep learning uses closely related ideas, but the SSM class of models took advantage of stability guarantees from integrators in control theory and created a class of RNN that don’t have to worry about exploding gradients. Mamba is one of the ways to make these SSM models much more expressive.
- Straw 3y ago
- givemeethekeys 3y agoTransformers, Rise of the Mambas, coming to a theater near you!
- kken 3y agoThere is also this: https://jackcook.com/2024/02/23/mamba.html https://jackcook.com/2024/02/23/mamba.html
- vanjajaja1 3y agoI really enjoyed this article, thanks
- kgeist 3y agoIf Mamba selectively forgets "unnecessary" details, can it repeat the input verbatim (if asked)?
- hackerlight 3y agoIf you ask at the end of the prompt then it may have already deliberately tossed the information it deemed irrelevant prior to the question. These aren't transformers. In general the recall for arbitrary information will be worse.
- kgeist 3y agoSo the questions should come before the content and it might work? I think that's how also RWKV works.
- hackerlight 3y agoIt's known to help although I wouldn't expect it to be perfect recall unless the network is big enough. The network will read the data token by token. So if you put the question at the beginning it will know what information it needs to pay attention to inside the rest of your context. Of course, if the network is too small, it still won't be perfect recall for a sufficiently complicated/large question/context.
- alok-g 3y agoA potentially naive question. Isn't this modeled like a Kalman Filter? Edit: Sounds like it is. https://openreview.net/pdf?id=AL1fq05o7H https://openreview.net/pdf?id=AL1fq05o7H