6 ms·
Open weight models have been getting better/smaller every year. Also, from what I can tell, MLX inference is not as well optimized as CUDA, and the M5 Ultra ha
by srcreigh 14d ago
Open weight models have been getting better/smaller every year.
Also, from what I can tell, MLX inference is not as well optimized as CUDA, and the M5 Ultra has additional kinds of AI compute which is unavailable on other M models. With the massive 1.2 TB/s 512GB Mac studios coming out, I think MLX will get a lot more attention.
In short: Todays models should run faster next year, and next year's models should also be more efficient.