5 ms·
Hi, it does not work with llama.cpp right?
by shirman 1y ago
Hi, it does not work with llama.cpp right?
- codelion 1y agoOptillm works with llama.cpp but this approach is implemented as a decoding strategy in PyTorch so at the moment you will need to use the local inference server in optillm to use it.