6 ms·
Pool spare GPU capacity to run LLMs at larger scale
- iwinux 6mo agoYou lost me on "spare GPU". I don't have any capable GPUs, let alone spare ones :)
- vagrantJin 6mo agoThis is very promising, definitely looks more user friendly than exo. Can't wait to try it out.
- lostmsu 6mo ago> MoE models via expert sharding with zero cross-node inference traffic This makes the whole project questionable