Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
phazonoverload
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
1.
▲
by
phazonoverload
14d ago
I will give it a look!
2.
▲
by
phazonoverload
14d ago
Coming back to this a few hours later, I've decided to add a section to explain this to the blog post. Thank you for flagging it.
3.
▲
by
phazonoverload
14d ago
I haven't, but that's just because it really isn't my personal usage pattern.
4.
▲
by
phazonoverload
14d ago
A literal manual typo. Good catch, will fix.
5.
▲
by
phazonoverload
14d ago
The beauty of this is that you can just swap out the platform and everything remains as it's the same backend. You make a really good point, one that I haven't really considered, but I also only have so many hours in the day to be
6.
▲
by
phazonoverload
14d ago
Hehe I really am just working it out as I go along - I promise it is fairly painless. Hugging Face allow you to specify your machine and then browse models that fit. And then you can just vibe out 'oh this one is a bit slow let me try
7.
▲
by
phazonoverload
14d ago
It is not
8.
▲
by
phazonoverload
14d ago
Running it depends on RAM, which is what I wrote, bandwidth is important for speed. I chose my words carefully, but you are absolutely right.
9.
▲
by
phazonoverload
14d ago
My perf sucks compared to yours. Added it to the post - same model averages 325 tok/s in processing prompts, and 34 tok/s in token generation. What am I doing wrong..?
10.
▲
by
phazonoverload
14d ago
I'm the author - hello! Added to the post! Qwen averages 325 tok/s in processing prompts, and 34 tok/s in token generation. That isn't instant, but it's quick enough that I never really think about it.
11.
▲
by
phazonoverload
14d ago
I'm the author - hello! I talk about it in the blog post - knowing what's being run, knowing where it's being run, and not having anyone else control it.