6 ms·
To be honest, running Deepseek v4 flash 0731 is enough for most what I need, and I like its responses way more. It's crazy that I can run this in a Q8 quantizat
by magic_hamster 29d ago
To be honest, running Deepseek v4 flash 0731 is enough for most what I need, and I like its responses way more. It's crazy that I can run this in a Q8 quantization in a home setup. It feels and performs like a frontier model.
The only issue with relying on local models is when you need them to prompt other models, and you might need to offload or switch models constantly which adds significant overhead.
But when it all works, its truly awe inspiring.