10 ms·The new IQ4XS has been working pretty well so far on 4090 16gb.by jadbox 28d agoThe new IQ4XS has been working pretty well so far on 4090 16gb.kamranjon 28d agoWhat size context are you able to squeeze in with less than 2gb of headroom? I have had some luck using a quantized kv cache but i fear that also decreases overall quality.beacon294 28d agoTry the llama.cpp fork by thetom. It's called turboquant after the technique