7 ms·
I'm also trusting `get_peak_memory` + some small buffer for now. Still, it reports accurate peak memory usage for tensors living on GPU, but seems to miss some
by aukejw 1y ago
I'm also trusting `get_peak_memory` + some small buffer for now.
Still, it reports accurate peak memory usage for tensors living on GPU, but seems to miss some of the non-Metal overhead, however small (https://github.com/aukejw/mlx_transformers_benchmark/issues/1 https://github.com/aukejw/mlx_transformers_benchmark/issues/...).