7 ms·
Running large MOE models won't saturate the EPYC memory. You'll never see that 200 Gb/s running inference.
by jaxytee 2mo ago
Running large MOE models won't saturate the EPYC memory.
You'll never see that 200 Gb/s running inference.
- edg5000 2mo agoBecause it'll bottleneck at the CPU? Or because 200 bB/s is overkill when running a model (that fits in memory)?