6 ms·
That doesn't really track, though. Even tiny models, like Gemma 4, have coherent prose. I know Opus isn't being quantized that small. It seems like it must be s
by SwellJoe 9d ago
That doesn't really track, though. Even tiny models, like Gemma 4, have coherent prose. I know Opus isn't being quantized that small. It seems like it must be some kind of...I dunno. Maybe over-fitting toward some user metric that doesn't track how readable its writing is?
It seems to still be good at code (though I haven't directly compared to earlier Opus versions lately), so it's not a general model collapse type problem, nor quantization errors. If quantization problems, I would expect it to break down on logic before prose, since a 4-bit quantized Gemma 4, even the small versions, have pretty good and, more importantly, coherent written English.
- cyanydeez 8d agoQwen3.8 quant 4 27b works coherently and takes up very little space. If the cloud AIs arnt downsizing the majority of their customers, they will be out of business.
- cyanydeez 8d agoQwen3.8-flash-next can run in ~60gb, offloading 50gb to ssd, and is comparible. The business has to downgrade to be sustainable, and open models prove it.