Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
danielhanchen
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
9 ms
·
1.
▲
by
danielhanchen
28d ago
We made something called Divergence-300 @32 (and later @512) which tests actual inference across 32 tokens on a held out test (Terminal Bench, DeepSWE, Math etc) We do plan to do larger benchmark suites though!
2.
▲
by
danielhanchen
28d ago
Hey we did not remove the MTP for sizes above 8GiB - but yes for small GGUFs under 8 ish GiB, we removed the MTP module (IQ2_XXS and lower), because it's 500MiB to 750MiB in size, and on small 8 GiB machines, even 500MiB is needed. As
3.
▲
by
danielhanchen
1mo ago
Actually we do publish non KLD benchmarks - top-1% is better - for NVFP4 for eg we did MMLU Pro, GPQA, AIME 2025: https://unsloth.ai/docs/models/qwen3.6#nvfp4-benchmarks Sometimes they're just slow and expens
4.
▲
by
danielhanchen
1mo ago
That wasn't our problem right? Gemma officially updated tool calling which we adopted
5.
▲
by
danielhanchen
1mo ago
We also made NVFP4 ones if that helps! https://huggingface.co/unsloth/Qwen3.8-27B-NVFP4
6.
▲
by
danielhanchen
1mo ago
Ok no worries - if there are any future issues - feel free to message / make a HF issue - we'll fix promptly! Also note its best to follow Gemma4's official sampling params since they evaled with it - dry multiplier sometimes
7.
▲
by
danielhanchen
1mo ago
We will investigate Ling!
8.
▲
by
danielhanchen
1mo ago
Thanks for the support and to the community!
9.
▲
by
danielhanchen
1mo ago
Hey :)
10.
▲
by
danielhanchen
1mo ago
Thanks haha
11.
▲
by
danielhanchen
1mo ago
Hey yes - if you could describe what the issues are - we will gladly fix them!
12.
▲
by
danielhanchen
1mo ago
Hey sorry what are the problems that you're experiencing - we're more than happy to help fix them!
13.
▲
by
danielhanchen
1mo ago
Oh thanks for sharing! We launched a Desktop app with fast diffusion, video gen, training support, inference + tool calling, web search, canvas, HTTPS secure remote access via Cloudflared, API / model swapping + more! It works in [Wind
14.
▲
by
danielhanchen
2mo ago
Oh thanks for sharing! The llama.cpp PRs should generally be fine for now - I'm fixing a few small edge cases as well!
15.
▲
by
danielhanchen
3mo ago
Very cool write-up and GitHub repo!
16.
▲
by
danielhanchen
4mo ago
Thank you appreciate the support! It's all thanks to you guys and the community!
17.
▲
by
danielhanchen
4mo ago
Update - Just got rid of the spiced up intro
18.
▲
by
danielhanchen
4mo ago
Thank you!
19.
▲
by
danielhanchen
4mo ago
Oh thanks :) We're also going to add MTP support soon for Qwen3.6! 95% of it is fully human done - the maths, algos, code snippets, screenshots & benchmarks are done / conducted by us and NVIDIA :) We did use AI to fix spellin
20.
▲
Mistral Medium 3.5 YaRN bug fix
(huggingface.co)
1 points
by
danielhanchen
5mo ago
|
0 comments
21.
▲
by
danielhanchen
5mo ago
Sorry on the delay - so it installs https://github.com/Blaizzy/mlx-vlm and other components and sets up the commands - you don't need to use it but we thought it might be easier for folks
22.
▲
by
danielhanchen
5mo ago
Sorry on the delay - oh haha that would be cool :) We did release 2bit dynamic ones, but unsure if they'll be helpful
23.
▲
by
danielhanchen
5mo ago
Yes we do! Sorry on the delay
24.
▲
by
danielhanchen
5mo ago
We use Duck Duck Go - sorry on the delayed response as well
25.
▲
by
danielhanchen
5mo ago
Thank you and appreciate it! Sorry on the delayed reply as well
26.
▲
by
danielhanchen
5mo ago
Oh yes LM Link is cool!
27.
▲
by
danielhanchen
5mo ago
Hey sorry on the delay - we just added API support, so you can access a remote server - it includes optional python, tool call, bash and web search support if you enable them. For SSH - we haven't yet done that - for now we have a SHA2
28.
▲
by
danielhanchen
5mo ago
Hey! Sorry for not replying sooner - yes we'll keep publishing more KLD - sadly some are saying we are "optimizing" for KLD now since we posted so many haha - but the whole purpose of quantization is to match the BF16 logits
29.
▲
by
danielhanchen
5mo ago
Hey so sorry didn't reply sooner - yes the docker used to be I think 4-8GB ish since CUDA sadly itself is 4GB I think, and PyTorch takes the rest. So unfortunately the Unsloth Docker image has ballooned due to this. We tried reducing i
30.
▲
by
danielhanchen
5mo ago
Apologies as well didn't reply sooner - Studio supports AMD out of the box now! We worked with AMD to make it work! One thing that is still missing is pre-compiled AMD ROCM binaries, which we're trying to see if we can integrate t
More ›