13 ms·
My perf sucks compared to yours. Added it to the post - same model averages 325 tok/s in processing prompts, and 34 tok/s in token generation. What am I doing w
by phazonoverload 15d ago
My perf sucks compared to yours. Added it to the post - same model averages 325 tok/s in processing prompts, and 34 tok/s in token generation. What am I doing wrong..?
- argee 14d agoWow, that's just about half the perf. I'm not sure what you're doing differently, though our hardware is a bit different: I am on a Macbook Pro M4 Pro, while you're on a Mac Mini. I would try a different version of the model from HuggingFace while ensuring it's MLX. I'm also using LM Studio, not oMLX, and I've seen some threads like these: https://www.reddit.com/r/LocalLLaMA/comments/1spuwir/omlx_10ts_slowlier_than_lm_studio_qwen36_35ba3_on/ https://www.reddit.com/r/LocalLLaMA/comments/1spuwir/omlx_10...