17 ms·
Running a Mac Mini M4 as a home server for a bunch of automation scripts right now. The mmap-based layer streaming is the part I'm most curious about -- how doe
by ryanholtdev 6mo ago
Running a Mac Mini M4 as a home server for a bunch of automation scripts right now. The mmap-based layer streaming is the part I'm most curious about -- how does latency look when you're streaming layers from disk mid-inference? I'd expect throughput to degrade sharply once you exceed unified memory, but maybe the Top-K sparsity masks enough of the weight accesses that it's not as bad as sequential streaming would be. What's the actual tokens/sec at 140B scale on the base Mac Mini config?
- anentropic 6mo agoYeah... https://github.com/opengraviton/graviton?tab=readme-ov-file#benchmarks https://github.com/opengraviton/graviton?tab=readme-ov-file#... the benchmarks don't show any results for using these larger-than-memory models, only the size difference it all smells quite sloppy
- hu3 6mo agoWhat could find in the readme shows: ~19 tok/s for Apple M1 Max (64 GB) with TinyLlama-1.1B-Chat-v1.0