Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
kiraaa
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
1.
▲
by
kiraaa
2y ago
when there are two commands in a prompt example do A and then do B. the model completely ignores the second task B.
2.
▲
by
kiraaa
2y ago
on gpu that is still huge.
3.
▲
by
kiraaa
3y ago
super easy to install and use
4.
▲
by
kiraaa
3y ago
mistral 7b v0.2 supports 32k
5.
▲
by
kiraaa
3y ago
could be true, we can only speculate.
6.
▲
by
kiraaa
3y ago
maybe they are using ring attention, on top of their 128k model.
7.
▲
by
kiraaa
3y ago
really like the font and great article btw.
8.
▲
by
kiraaa
3y ago
the paper does not live up to the quality of model lol
9.
▲
by
kiraaa
3y ago
you need 94gb, does not matter which RAM.
10.
▲
by
kiraaa
3y ago
https://github.com/sysid/sse-starlette makes token streaming so much easier in python
11.
▲
by
kiraaa
3y ago
and also easy to deploy
12.
▲
by
kiraaa
3y ago
its not alway about the size, but yeah its really good!
13.
▲
by
kiraaa
3y ago
installing it is a nightmare
14.
▲
by
kiraaa
3y ago
you could but it will be very slow. oh and make sure the code is cpu compatible.