Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
korbip
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
1.
▲
Unlocking the Working Memory of Large Language Models for Latent Reasoning
(arxiv.org)
2 points
by
korbip
4mo ago
|
0 comments
2.
▲
by
korbip
8mo ago
This was done already here as well: https://arxiv.org/abs/2507.04239
3.
▲
Show HN: CompoConf – modular configuration for modular systems
(korbi.ai)
2 points
by
korbip
1y ago
|
0 comments
4.
▲
by
korbip
1y ago
I can share a similar PhD story (the result being visible here: https://github.com/NX-AI/flashrnn ). Back then I didn't find any tutorials that cover anything beyond the basics (which are still important). Once you
5.
▲
by
korbip
2y ago
Test it out here: https://github.com/NX-AI/mlstm_kernels https://huggingface.co/NX-AI/xLSTM-7b
6.
▲
by
korbip
2y ago
There is a LOT of effort in the research community currently: 1. Improving the Self-Attention in the Transformer as is, keeping the quadratic complexity, which has some theoretical advantage in principle[1]: The most hyped one probably Deep
7.
▲
by
korbip
2y ago
Thanks! I don't see any implementation there. In any case, we are planning a code release soon.
8.
▲
by
korbip
2y ago
You mainly got it right. Usually one does have many scalar 'c' cells, that talk to each other via memory mixing. For the sLSTM, you group them into heads, talking only to cells within the same head. The reason that we referred to
9.
▲
by
korbip
2y ago
This was formulated a bit unclear. It is not possible to parallelize in the sequence dimension for training as it is possible for Transformers. In the batch dimension you can always do it.
10.
▲
by
korbip
2y ago
For language in general it seems fine. But there might be specific tasks where it is necessary indeed.
11.
▲
by
korbip
2y ago
Thank you! I can say that it is not really a diminishing factor at the scales reported in the paper. So, xLSTM[7:1] is pretty much on par with xLSTM[1:0] in speed. We show that it is helpful on toy tasks, and it shows even better sequence e
12.
▲
by
korbip
2y ago
Disclaimer: I'm shared first author of this paper. As a clarification: The speed for training will be on par with FlashAttention-2, when fully optimized and only including the mLSTM. For decoding/inference both are very close to M
13.
▲
Show HN: OpenPV – Photovoltaic Potential in Bavaria (and Beyond?)
(openpv.de)
6 points
by
korbip
3y ago
|
0 comments
14.
▲
Self-Expanding Neural Networks
(arxiv.org)
4 points
by
korbip
3y ago
|
1 comments
15.
▲
DeepMind AlphaCode: AI learns to solve programming competition tasks [pdf]
(storage.googleapis.com)
3 points
by
korbip
5y ago
|
0 comments
16.
▲
Transformers Connected to Hopfield Networks
(arxiv.org)
3 points
by
korbip
5y ago
|
0 comments
17.
▲
Norona – We make rapid Covid tests visible
(noronatest.me)
2 points
by
korbip
6y ago
|
0 comments
18.
▲
I Can. We Want
(korbip.de)
1 points
by
korbip
6y ago
|
0 comments
19.
▲
by
korbip
6y ago
Feel free to add them. :)
20.
▲
AutoRail – Autonomous, Simple, Energy-Efficient Mobility – A Vision
(korbip.de)
2 points
by
korbip
6y ago
|
0 comments
21.
▲
PyTorch-ProbGraph – Hierarchical Probabilistic Graphical Models in PyTorch
(github.com)
25 points
by
korbip
6y ago
|
4 comments