5 ms·
This: https://github.com/Blaizzy/mlx-lm/tree/pc/add-lg https://github.com/Blaizzy/mlx-lm/tree/pc/add-lg ...and this is what I should probably wait for (not sure
by Lwerewolf 2mo ago
This:
https://github.com/Blaizzy/mlx-lm/tree/pc/add-lg https://github.com/Blaizzy/mlx-lm/tree/pc/add-lg
...and this is what I should probably wait for (not sure why it's in vlm):
https://github.com/Blaizzy/mlx-vlm/tree/pc/laguna-s-nvfp4 https://github.com/Blaizzy/mlx-vlm/tree/pc/laguna-s-nvfp4
...or perhaps I should've just used the gguf with the provided llama.cpp instead of trying to run the nvfp4-mlx from the get go, but where's the chaos in that :)
Running deepseek flash on something locally now, this will have to wait a bit. I still stand by my initial quick assessment - looks capable. Some people on r/localllama also reported loops. We'll see in ~10 hours. Hopefully I haven't terribly mislead people.
- ilc 2mo agoThis is definitely one where wait 2-3 weeks and look again, feels like the right strategy. It could be good, it could not. It is too volatile to say.
- Lwerewolf 2mo agoMight wanna try it now, seems to have been largely fixed. Check huggingface threads and reddit. Works for me, very memory hungry and PP speed drops off a cliff around 200k context (~30tok/s decode and ~40tok/s pp - like... hope it isn't debugging with big logs), but it brings a fresh perspective alongside ds4-flash. The more, the merrier - I prefer keen eyeballs more than speed anyways.