7 ms·
It usually boils down to people trying to convince themselves that keeping their macs hot and with very little ram to spare only to get sub 50 tokens per second
by bel8 15d ago
It usually boils down to people trying to convince themselves that keeping their macs hot and with very little ram to spare only to get sub 50 tokens per second on a subpar lobotomized (quantized) model is worth it.
And I'm not even considering their time spent fiddling, fine tuning configs to adjust for ram, updating/benchmarking models, etc. Which is probably more expensive than the mac so the math is even more wrong.