7 ms·
nvfp4 mlx, literally barebones pi. edit: on bigger tests, got it to loop pretty easily unfortunately, probably local settings.
by Lwerewolf 2mo ago
nvfp4 mlx, literally barebones pi.
edit: on bigger tests, got it to loop pretty easily unfortunately, probably local settings.
- sosodev 2mo agoWhat inference server are you using? They have a custom branch for llama.cpp, but I wouldn't be surprised at all if it still needs fixing.
- Lwerewolf 2mo agoThis: https://github.com/Blaizzy/mlx-lm/tree/pc/add-lg https://github.com/Blaizzy/mlx-lm/tree/pc/add-lg ...and this is what I should probably wait for (not sure why it's in vlm): https://github.com/Blaizzy/mlx-vlm/tree/pc/laguna-s-nvfp4 https://github.com/Blaizzy/mlx-vlm/tree/pc/laguna-s-nvfp4 ...or perhaps I should've just used the gguf with the provided llama.cpp instead of trying to run the nvfp4-mlx from the get go, but where's the chaos in that :) Running deepseek flash on something locally now, this will have to wait a bit. I still stand by my initial quick assessment - looks capable. Some people on r/localllama also reported loops. We'll see in ~10 hours. Hopefully I haven't terribly mislead people.
- ilc 2mo agoThis is definitely one where wait 2-3 weeks and look again, feels like the right strategy. It could be good, it could not. It is too volatile to say.
- Lwerewolf 2mo agoMight wanna try it now, seems to have been largely fixed. Check huggingface threads and reddit. Works for me, very memory hungry and PP speed drops off a cliff around 200k context (~30tok/s decode and ~40tok/s pp - like... hope it isn't debugging with big logs), but it brings a fresh perspective alongside ds4-flash. The more, the merrier - I prefer keen eyeballs more than speed anyways.
- embedding-shape 2mo ago> edit: on bigger tests, got it to loop pretty easily unfortunately, probably local settings. Been playing around for a few hours with the poolside/Laguna-S-2.1-NVFP4 + poolside/Laguna-S-2.1-DFlash-NVFP4 + vLLM, been seeing the same behaviour. Usually new model releases are plagued with issues at release though, best to wait 1-2 weeks then retry, or better yet, investigate yourself :) Personally I haven't found any obvious issues.
- embedding-shape 2mo agoUpdate: Seems quite literally they have bugs on the hardware I'm trying to run this with: From Poolside CEO Eiso Kant on Twitter: > Learning we have some bugs on the RTX6000. We’re on it. Team has worked non stop last days and it’s getting late for a lot of the inference folks, so might be until tomorrow till we have a solution. - https://x.com/eisokant/status/2079693050796785720 https://x.com/eisokant/status/2079693050796785720 Update2: I'm now running poolside/Laguna-S-2.1-NVFP4 with vLLM 0.23.1rc1.dev1378+gd6dbdb9b0 (FlashInfer 0.6.14) and seeing slightly better results in regards to the looping. I can't see any specific changes that would affect this though, strangely enough.
- embedding-shape 2mo agoUpdate3, summarized from RTX6kPRO Discord: > The BF16 checkpoint doesn't exhibit any of the complaints [...] with our current quants, the models tend to choose the wrong logits sometimes [...] why we'll need a requant [...] We're not aware of any bugs in any runtimes themselves [...] We have two remaining things that we're trying to tackle: some people are reporting thinking being too hard to trigger (^^) , and others are saying it thinks too much. We've seen much more of the latter internally
- Lwerewolf 2mo agoDeleted earlier, didn't see you post, pasting here: /* Just started testing with the gguf (with gpu offload, m5 max 128gb), q4_k_m, running seemingly well. Speed is initially slightly faster than antirez/ds4 - decode tok/s in the 30s, prefill ~400-ish. Expected, given the slightly smaller size. Looks to be working fine, but too early to tell. Definitely likes to "think". */ Anyways, guessing that discord might be focusing on the nvfp4 stuff. I've noticed spelling mistakes in the thinking traces, tool calls have been fine so far.