6 ms·
For inference you could use a maxed-out Mac Ultra; the RAM is shared between the CPU and GPU.
by dmbaggett 2y ago
For inference you could use a maxed-out Mac Ultra; the RAM is shared between the CPU and GPU.
- alecco 2y agoFor single user (batch_size = 1), sure. But that is quite expensive in $/tok.