7 ms·In llama.cpp You can offload some of the layers to gpu with -ngl X. Where x is the number of layersby Dkuku 3y agoIn llama.cpp You can offload some of the layers to gpu with -ngl X. Where x is the number of layers