5 ms·
Very cool. I jumped in here thinking it was gonna be something else though: a packaged service for distributing on-prem model running across multiple GPUs. I'
by singlepaynews 1y ago
Very cool. I jumped in here thinking it was gonna be something else though: a packaged service for distributing on-prem model running across multiple GPUs.
I'm basically imagining a vast.ai type deployment of an on-prem GPT; assuming that most infra is consumer GPUs on consumer devices, the idea of running the "company cluster" as combined compute of the company's machines
- mhamann 1y agoGreat point. I can see how you'd land there. Also a great idea! xD Maybe a better descriptor is "self-sovereign AI?" "Self-hosted AI?"
- olokobayusuf 1y agoWe're building something closer to this at Muna: https://docs.muna.ai https://docs.muna.ai . Check us out and let me know what you think!
- jochalek 1y agoSounds like something that could be implemented with llm-d, though I've not experimented with it. https://llm-d.ai/blog/intelligent-inference-scheduling-with-llm-d https://llm-d.ai/blog/intelligent-inference-scheduling-with-...