Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
delicious_apple
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
9 ms
·
1.
▲
by
delicious_apple
1mo ago
I'm currently running it on an RTX 3090 (street price ~$1000 USD) with a long context and getting pretty good performance. Prefill: ~1000 tok/s Decode: 75-100 tok/s It'll be far faster on a 5090, but I find the above per
2.
▲
by
delicious_apple
1mo ago
I am running it on a single RTX 3090 (24GB VRAM). Some folks on Reddit are having the same experience: https://www.reddit.com/r/LocalLLaMA/comments/1vkm42m/muse_gl... It uses an order of magnitude less V
3.
▲
by
delicious_apple
4mo ago
You can always try openSUSE Slowroll (in beta), which is a rolling release that updates less frequently than Tumbleweed. It advertises better stability. https://en.opensuse.org/Portal:Slowroll