5 ms·
yeah we tried out popular solutions like exllama and llama.cpp among others that support inference of 4bit quantized models
by junrushao1994 3y ago
yeah we tried out popular solutions like exllama and llama.cpp among others that support inference of 4bit quantized models