Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
carloslfu
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
1.
▲
Ask HN: What is the best local model that runs on your Mac at decent speed?
1 points
by
carloslfu
3h ago
|
0 comments
2.
▲
by
carloslfu
13d ago
It wasn't either/or, the N-gram table is part of Qwen itself and stays on disk. I’ve now added its 1.5GB MTP draft head too, it gets 86% acceptance and about 1.24× faster decoding on my 48GB Mac.
3.
▲
by
carloslfu
14d ago
thanks! in part, I was wondering how you got the code into files. I guess you copy pasted it inside a file, am I right?
4.
▲
by
carloslfu
14d ago
This is next! in the works rn.
5.
▲
by
carloslfu
14d ago
thanks! > ditched oLlama" yeah! this is interesting. > 8.1GB per slotserve process is a lot! Is that in your control? yes, it is hard, but I agree the smaller the better. I'll work on that > If it's local and open-we
6.
▲
by
carloslfu
14d ago
great idea!! a native app would be awesome
7.
▲
by
carloslfu
15d ago
Both projects are different in scope. Think of slotstream as optimizing for memory and for this specific model for now, my intention is not to build an inference engine the same as oMLX
8.
▲
by
carloslfu
15d ago
Interesting! I'll check it out
9.
▲
by
carloslfu
15d ago
true
10.
▲
by
carloslfu
15d ago
for the record, I'm almost 35
11.
▲
by
carloslfu
15d ago
interesting! Yes, thermal is important. Pretty cool project man! Starred and checking it out!
12.
▲
by
carloslfu
15d ago
This is the best I could find: https://huggingface.co/Qwen/Qwen3.8-Flash-Next?utm_source=ch... About the specifics, I have only anecdotal evidence, but I guess this info can be found somewhere
13.
▲
by
carloslfu
15d ago
I hope not! this is a new macbook lol!
14.
▲
by
carloslfu
15d ago
I don't know actually. I'll check haha. My best guess is it isn't.
15.
▲
by
carloslfu
15d ago
I see! yes, downloading the weights part is painful. I tried a couple fixes and it is as fast as it can get downloading from HuggingFace. I think the field is heading toward smaller, more capable models soon, so you won't have to wait
16.
▲
by
carloslfu
15d ago
yes! I'm bullish on this. there is a lot of work to do. I've been experimenting with pruning, distillation, and retraining too. I'm sure your 32gb m6 will run a badass local model!
17.
▲
by
carloslfu
15d ago
I feel you! fix incomming
18.
▲
by
carloslfu
15d ago
thanks! I'll do!
19.
▲
by
carloslfu
15d ago
interesting!
20.
▲
by
carloslfu
15d ago
yes! I guess future hardware designs will have something like that!
21.
▲
by
carloslfu
15d ago
thanks!
22.
▲
by
carloslfu
15d ago
Ah! Yeah, I didn't invent anything (yet!). The goal is to see how far I can take it in terms of speed without consuming that much RAM.
23.
▲
by
carloslfu
15d ago
Good one! I haven't measured this. I'll include it!
24.
▲
by
carloslfu
15d ago
I agree with the sentiment, but have you seen those videos in which all men say other men are gay? This feels like the same, so much AI paranoia! I genuinely want to contribute. And hey! I was doing oss this since 2014 so waay before AI was
25.
▲
by
carloslfu
15d ago
I'm sorry this makes it seem like I didn't do my research. I did a TON. To fix it I'll add a benchmark/comparison table. Also, I wouldn't call it market research since this is not commercial AT ALL.
26.
▲
by
carloslfu
15d ago
Sorry, I don't get "NIH". what's that?
27.
▲
by
carloslfu
15d ago
Thanks for the feedback! I'll create a section with a benchmark and comparisons. This will hold the project accountable and speed things up imo
28.
▲
by
carloslfu
15d ago
I see your point. As an oss defender myself, I agree, however, the spirit of this is to see how fast I can make it. I'm sharing this with the community, which I think is aligned with the original oss spirit. It's an experiment for
29.
▲
by
carloslfu
15d ago
isn't the system prompt and tools part of the harness?
30.
▲
Show HN: Running 104GB Qwen3.8-Flash-Next on 48GB Mac with at ~12 tok/s
(github.com)
240 points
by
carloslfu
15d ago
|
118 comments
More ›