7 ms·
The second sentence of the article: > We have a number of these now and have been running our RDMA network backend for the 8x NVIDIA GB10 cluster for several m
by MarkSweep 15d ago
The second sentence of the article:
> We have a number of these now and have been running our RDMA network backend for the 8x NVIDIA GB10 cluster for several months now, using one.
I have a friend who uses something similar for his 4 NVIDIA GDX Spark cluster.
- avidiax 15d agoI see, so only if you are running an AI cluster. And what is it transferring between the nodes? Partial computations?
- ericd 15d agoEssentially, yes. The DGX Sparks, at 128 gigs for ~$4700 are one of the cheapest ways to get enough high-ish speed memory cobbled together to run one of the more capable open weight models to run at home/for a small biz. This switch lets you combine the memory of 2-4 of them.
- bitbckt 15d ago4-8 of them. 2 Sparks can operate without a switch.
- ericd 15d agoWell, this one only has 4 ports, would you daisy chain them somehow to get to 8? Or bridge three of these switches?
- dgl 15d agoThey are QSFP56-DD which can be split using breakout cables. So it's 4 x 400GbE ports, which could in theory be split into as many as 32 x 50GbE ports. (Although whether that then ends up being an economical thing to do is a different matter.)
- ericd 15d agoAhh yeah, I've seen people using those. I was thinking you'd ideally want to connect them full speed, but it looks like they're 200gbps ports anyway, and I guess even if they were 400gpbs, you don't really need that much for cross-GPU traffic for MoE models doing tensor parallel.
- nubinetwork 15d agoOh I'm sorry, just let me casually spend 30k...
- MarkSweep 15d agoI'm not defending the use of money; I think it's questionable. It's still a home lab by my definition: it's computer stuff that costs less than a sports car in your home.