7 ms·
Because OP is running it on an M3 Ultra with Ollama. He'll get much better perf with llama.cpp or MLX, both of which Ollama wraps, albeit very poorly.
by woadwarrior01 21d ago
Because OP is running it on an M3 Ultra with Ollama. He'll get much better perf with llama.cpp or MLX, both of which Ollama wraps, albeit very poorly.