Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
macwhisperer
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
8 ms
·
1.
▲
by
macwhisperer
14d ago
lmao *qwen3.8-27b runs in the background*
2.
▲
by
macwhisperer
14d ago
honestly boggles my mind that actual companies put load-bearing (lol) work behind 3rd party api's like claude or gpt.. also its really funny because some of these companies fired actual humans to put said 3rd party api to work in place
3.
▲
by
macwhisperer
23d ago
super dope!! the future of musical instruments may be web-based.. exciting times
4.
▲
by
macwhisperer
1mo ago
4-bit quant sits at 15.72GB
5.
▲
by
macwhisperer
1mo ago
yeah this model is chefs kiss.. running an untouched, vanilla 4-bit version (Q4_0) I baked myself today (benched it against Q4_K_M (16gb) and IQ3_M (12gb), Q4_0 (15gb) is king)... this model-- 1: over 60% faster than qwen 3.6 version of the
6.
▲
by
macwhisperer
1mo ago
pro tip: if u are building ur own harness with python (recommended) , use "llama-cpp-python"... I started with "llama-server" and custom stuff around it, which is great for single model setups.. but for multi-model harne
7.
▲
by
macwhisperer
1mo ago
big week for open models... seems like companies are noticing the 26-35b sweet spot... though I think a 12b-a1b-MoE model would be helpful for the 16gb folks
8.
▲
by
macwhisperer
1mo ago
I feel like all the ppl complaining itt don't even run open weight models.. wahh billionaire bad is true, but you are missing the forrest for the trees.. they are literally bending the knee and ur mad about it? it literally means loc
9.
▲
by
macwhisperer
1mo ago
yeah you realize you can turn reasoning off on the newer the models right? also try my version of Gemma-12b https://huggingface.co/macwhisperer/Gemma4-12B-SuperDense or try my qwen 3.5-9b
10.
▲
Show HN: Is the Moon Full?
(isthemoonfull.neocities.org)
2 points
by
macwhisperer
1mo ago
|
0 comments
11.
▲
by
macwhisperer
2mo ago
Basically, he proved that *information is power.* If you don't know which way to go (the subgradient), you're gonna be calculating forever!
12.
▲
by
macwhisperer
3mo ago
also for those with only 16gb-- try this model https://huggingface.co/macwhisperer/Gemma4-12B-SuperDense its exceptional!
13.
▲
by
macwhisperer
3mo ago
hi guys... I run specialized quants on my 24gb air.. (I specialize in 3-bit quants that punch above their weight).. try out my version of 3.6-27b I think you be impressed https://huggingface.co/macwhisperer/Qwen3.6-27B-
14.
▲
by
macwhisperer
3mo ago
qwen3 1.7b- q4_k_m is your best best for that size
15.
▲
by
macwhisperer
3mo ago
literally ask cloud ai like the free gemini or chatgpt.. they could make you an expert on the subject overnight..
16.
▲
by
macwhisperer
3mo ago
I code with like a slew of 20+ custom baked models of all sizes, in various fully custom multi-model harnesses that use different bindings... the harnesses themselves are just as important as the models...different harnesses give different
17.
▲
by
macwhisperer
3mo ago
ai is like the first technology with a conversational service manual inside it.. you should be foaming at the mouth to use claude or codex to make a custom harness, just for your own personal use with local models...
18.
▲
by
macwhisperer
3mo ago
super inspiring! thanks for sharing!
19.
▲
by
macwhisperer
3mo ago
retro-inspired fully custom, swiss army knife style notepad -- https://convert.neocities.org
20.
▲
by
macwhisperer
3mo ago
check out a custom 4-bit quant I made today https://huggingface.co/macwhisperer/Gemma4-12B-SuperDense should run perfect for 12-16gb with maybe 10-20k context seems intelligent enough that I would recommend this as a d
21.
▲
by
macwhisperer
3mo ago
the HITL (human in the loop) is basically the single point...AI is a mirror.. it only "exists" when you talk to it.. much like your reflection in the mirror is only there when you're in view. models can never be self-improvin
22.
▲
by
macwhisperer
4mo ago
this is cool thanks for making it!
23.
▲
by
macwhisperer
4mo ago
cool! what stack are you using for the multiplayer?
24.
▲
by
macwhisperer
4mo ago
can't we just freaking share things? ai models are a crystallization of human effort (available for free on huggingface).. why not use it? AI is like the UBI of intelligence, stop leaving free money on the table by refusing to use it l
25.
▲
by
macwhisperer
4mo ago
this is really cool congrats!
26.
▲
by
macwhisperer
4mo ago
good, the point is that now you have free mental bandwidth to use on building something that truly interests you using AI to help actualize your goal. build something cool for your kiddos idk?
27.
▲
by
macwhisperer
4mo ago
at this point I trust software companies less and less.. being able to build the stack and create bespoke solutions with llms's is incredible.. idk why people get mad about vibe-coding. if ur little brother can make a Spotify clone wit
28.
▲
by
macwhisperer
4mo ago
I run the latest 20b-30b models on a MacBook Air... running inference with an MoE (25 tps) for like 2 hours is like 10% battery.. (look me up on huggingface to download my models) also you gotta realize frontier models have massive "sy
29.
▲
by
macwhisperer
4mo ago
ai is exciting because it shows us what really matters...
30.
▲
by
macwhisperer
4mo ago
can you add in the other quants like IQ3_M? also my personal simple rule of thumb for local ai sizing is: max model size (GB) = ram (GB) / 1.65
More ›