Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
kamranjon
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
1.
▲
by
kamranjon
8d ago
Is nobody using structured outputs? They use constrained decoding at the generation stage to ensure the probability of tokens that would break the format are set to 0. I kinda figured everyone was doing this at this point.
2.
▲
by
kamranjon
8d ago
Hey there! I do the same but I use dwarfstar at a 2-bit quant: https://github.com/antirez/ds4 I'm curious if you've tried dwarfstar and decided to move to llama.cpp and 3 bit quants or what made you go that r
3.
▲
by
kamranjon
12d ago
I follow llama.cpp pretty closely as I use either llama.cpp itself or projects that depend on it all the time, and one thing that I don't think gets talked about is the sheer scale of community involvement. It seems like a logistical n
4.
▲
by
kamranjon
12d ago
Did you post in the wrong thread?
5.
▲
by
kamranjon
12d ago
I think this account should be banned.
6.
▲
by
kamranjon
13d ago
it's funny that the tagline is Radically Open, but you're immediately hit with http login - maybe this was the wrong link?
7.
▲
by
kamranjon
14d ago
They said Opus 5 medium - which does have an intelligence score of 59 (you have to select it manually from the dropdown to see it)
8.
▲
by
kamranjon
14d ago
Yea I am testing through OpenRouter - have you noticed 3.7 flash being significantly faster?
9.
▲
by
kamranjon
14d ago
They've interestingly left out any mention of speed. I have been testing 3.7 flash against 3.5 flash and it seems to lose every time in overall latency. Every benchmark I've seen seems to suggest the opposite[1] - that 3.7 flash i
10.
▲
by
kamranjon
15d ago
I am running 3.8 27b at q6 quant with 160k context on a 32gb video card (arc b70 pro) - I quantized the kv cache at q8 - that is the only trick really - works great.
11.
▲
by
kamranjon
21d ago
Hmm, makes the 2 bit quants actually seem pretty reasonable…
12.
▲
by
kamranjon
21d ago
Trained on roughly 1/9th of the training FLOPS - it’s the pretty incredible that they are making these advances and at the same time sharing their learnings in these papers - I wish we saw more of this from US labs.
13.
▲
by
kamranjon
22d ago
I actually do this with my MBP - it's a LLM server when I'm working - and then when I'm not it's just a really great machine for video editing and other media work.
14.
▲
by
kamranjon
22d ago
Where did they say that? My understanding of this 3.8-Flash-Next release is that it's a MOE (as per the title of the posting here, 125B a6b)
15.
▲
by
kamranjon
22d ago
Have you tried FreeToken yourself? I was hoping to find some benchmarks on their github but took a quick pass at their research paper and it seems they're showing ~2x performance on qwen 3.6 35b when compared to llama.cpp - but llama.c
16.
▲
by
kamranjon
23d ago
I think you misread, it’s 170gb/s for base M6 model and 1.2tb/s for M5 ultra.
17.
▲
by
kamranjon
23d ago
You would want to get the M5 pro version with 307gb/s if you were interested in running local LLMs.
18.
▲
by
kamranjon
23d ago
1tb would likely be ~$20k - given the current >$10k price tag of 256gb. Would you still be considering it at that price?
19.
▲
by
kamranjon
23d ago
You can get the m5 pro in the Mac mini with 307GB/s at 64gb of memory it’s $2899
20.
▲
by
kamranjon
24d ago
Yeah the article feels as though its describing Khan Academy from 10 or more years ago - it is quite odd. "What he has never had is pedagogical knowledge: an understanding of how people learn, what motivates them, what makes the differ
21.
▲
by
kamranjon
26d ago
Is anyone familiar with the laws surrounding police basically operating their own pseudo cell towers? I would assume this would be highly illegal for individuals, what sort of hoops did law enforcement need to jump through to get this type
22.
▲
by
kamranjon
26d ago
Hi there! I actually thought your Dia models were amazing and very natural sounding, I haven’t tried qwen 3 tts yet - has your focus shifted away from building your Dia models and shifted more towards hosting and infrastructure?
23.
▲
by
kamranjon
28d ago
Just wanted to share this, I found it was a really nice resource to understand how diffusion Gemma worked: https://newsletter.maartengrootendorst.com/p/a-visual-guide-... The really interesting thing to me was that the
24.
▲
by
kamranjon
28d ago
I wonder if this could be used by insurance companies to determine premiums?
25.
▲
by
kamranjon
28d ago
Are you using the recommended settings for temperature and such? https://unsloth.ai/docs/models/qwen3.8#recommended-settings Often times I run into issues like this it’s because I am using settings for a different
26.
▲
by
kamranjon
28d ago
I think if a real human doctor completely fabricated a drug usage history for one of their patients, that would also be news.
27.
▲
by
kamranjon
28d ago
What size context are you able to squeeze in with less than 2gb of headroom? I have had some luck using a quantized kv cache but i fear that also decreases overall quality.
28.
▲
by
kamranjon
1mo ago
A no-thinking pelican! I hope to see more, it's surprisingly good for just 2 minutes.
29.
▲
by
kamranjon
1mo ago
I actually disagree that this doesn’t mean anything. I understand the contention that it’s not measuring the quality of the model in general, but I think it is measuring something useful. A good example of this is planning hardware projects
30.
▲
by
kamranjon
1mo ago
Have you thought about running a second tier of the Pelican benchmark where you see which model makes best pelican on lowest or no reasoning settings? I think that'd be pretty interesting and might help highlight which models have a ba
More ›