6 ms·
Sure, but integrated graphics usually lacks vram for LLM inference.
by om8 1y ago
Sure, but integrated graphics usually lacks vram for LLM inference.
- adastra22 1y agoWhich means that inference would be approximately the same speed (but compute offloaded) as the suggested CPU inference engine.
- deleted 1y ago[deleted]