6 ms·
We're running it on vLLM and are working with others in the community to bring it to other optimized inference frameworks.
by juberti 2y ago
We're running it on vLLM and are working with others in the community to bring it to other optimized inference frameworks.