22 ms·
MKML says that they reduced the size of Llama2-13B model from 26GB to 10.5GB. Similar offering from TheBloke (your first link) is a 10.7GB Q6_K model. Maybe, th
by kapildev 3y ago
MKML says that they reduced the size of Llama2-13B model from 26GB to 10.5GB. Similar offering from TheBloke (your first link) is a 10.7GB Q6_K model. Maybe, they are using GGML and llama.cpp and packaging it in an attractive way while making people believe it is some proprietary tech.
- polishgladiator 3y agoBased on the integration examples, I don't think they are simply repackaging llama.cpp Rather it looks like they are reimplementing their own quantization scheme, in such a way that it is a little easier to integrate for basic python users, at the cost of performance (compared to llama.cpp and others). Given that the bar for integrating something with higher perf like llama.cpp isn't very high (and that's the way the world is heading -- ask any 15 year old interested in this stuff), I can't see anything of value here.
- Lindon4290 3y agoLooks their performance is better than llama.cpp - https://news.ycombinator.com/item?id=37018989 https://news.ycombinator.com/item?id=37018989 - and scales to batches of prompts.
- polishgladiator 3y agoActually no -- that post shows they are not performing measurements and comparisons correctly. These are not serious people.