Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
koayon
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
Art Criticism for Software
(paloaltoreview.co)
1 points
by
koayon
2y ago
|
0 comments
2.
▲
by
koayon
2y ago
And given that the compute is O(n^2) with context window, it's a very real tradeoff, at least in the short term
3.
▲
by
koayon
2y ago
This is a very fair point! If we had infinite compute then it's undeniable that transformers (i.e. full attention) would be better (exactly as you characterise it) But that's the efficiency-effectiveness tradeoff that we have to m
4.
▲
by
koayon
3y ago
Hey! OP here Great question - h' in Equation 1a refers to the derivative of h with respect to time (t). This is a differential equation which we can solve mathematically when we have x in order to get a closed-form solution for h. We
5.
▲
by
koayon
3y ago
Another interesting one is that the hardware isn't really optimised for Mamba yet either - ideally we'd want more of the fast SRAM so that we can store more larger hidden states efficiently
6.
▲
by
koayon
3y ago
Definitely agree that a lot of work going into hyperparameter tuning and maturing the ecosystem will be key here! I'm seeing the Mamba paper as the `Attention Is All You Need` of Mamba - it might take a little while before we get every
7.
▲
Mamba Explained: The State Space Model Taking On Transformers
(kolaayonrinde.com)
270 points
by
koayon
3y ago
|
93 comments
8.
▲
The Frontier of Adaptive Computation in Machine Learning
(github.com)
1 points
by
koayon
3y ago
|
0 comments
9.
▲
DeepSpeed's Bag of Tricks for Training Large Models
(kolaayonrinde.com)
1 points
by
koayon
3y ago
|
0 comments