7 ms·
The way it does tensor splitting without all-reduce cost over PCIe bus wasn't something I thought was possible. What kind of performance are you getting with 4
by karmakaze 9d ago
The way it does tensor splitting without all-reduce cost over PCIe bus wasn't something I thought was possible.
What kind of performance are you getting with 4x R9700s--what do you do with all the VRAM (batching, concurrent requests, etc)?
- intothemild 9d agopersonally? i have 2x gpus.. but i get bursts of ~200tok/s generation, and around 4500-5000tok/s prefil Yeah the R4D Kernel rules imho.
- karmakaze 9d agoSimilar here peak ~250 and down to ~120 as it gets close to 128k (which is where I set DSH compaction) though it can readily do 256k. I just got DeepSeek Harness (DSH) set up with 2x R9700 and it's rather mind blowing that these can do actual work and quickly. Up until now I've always been evaluating and searching for better hardware/model/tweaks. This is much more than I even hoped for and considered getting extra 3090/4090. Now I can stop looking/tweaking and start using it for all the different things I've yet to discover it's good for. I do plan to also try/use Hermes and Pi. DSH is annoying that every plugin install/remove requires a restart--given that "everything's a plugin".