7 ms·
M4 Max is typically better than M5 Pro for inference IIRC.
by fouc 2mo ago
M4 Max is typically better than M5 Pro for inference IIRC.
- harrouet 2mo agoIt depends on what you are looking at. Time to 1st token is faster on the M5 because of HW accelerators helping the prompt interpretation (and it is CPU-bound). Token generation after that is GPU-bound and will profit from the higher bandwidth of the M4 Max.