5 ms·
Issue is that llama.cpp is the best way to run models on hardware that isn't nvidias.
by noosphr 21d ago
Issue is that llama.cpp is the best way to run models on hardware that isn't nvidias.
- alightsoul 21d agoExcept when they have less than 16 gb of ram?
- dannyw 20d agoA lot of llama.cpp contributions come from the community and ecosystem, like Unsloth. If something goes awry, I fully expect lots of forks.
- MrDrMcCoy 20d agoThere already are a lot of forks for things they decline to implement. TurboQuant, ROCmFPX, and more. I need to set up an agent that will loop on merging them.