6 ms·
I am eyeing one of these specifically for this use case, could you please post roughly what kind of tokens per second numbers you get for text generation for th
by busfahrer 1mo ago
I am eyeing one of these specifically for this use case, could you please post roughly what kind of tokens per second numbers you get for text generation for this 27B model?
edit: and which quant you are using, please :-)
- noman-land 1mo agoUsing the 4bit quant on an M1 64GB I'm getting ~65 tps for prompt processing and ~11 tps token generation using oMLX to serve the models and pi as a harness.