4 ms·
> you can combine Spark with M3U, the former streaming the compute, lowering TTFT, the latter doing the token generation part Are you doing this with vLLM, or
by echion 9mo ago
> you can combine Spark with M3U, the former streaming the compute, lowering TTFT, the latter doing the token generation part
Are you doing this with vLLM, or some other model-running library/setup?
- coder543 9mo agoThey're probably referencing this article: https://blog.exolabs.net/nvidia-dgx-spark/ https://blog.exolabs.net/nvidia-dgx-spark/