6 ms·
You can just use Triton which is basically TFserve for Tensorflow, Pytorch, Onnx and more.
by cygn 3y ago
You can just use Triton which is basically TFserve for Tensorflow, Pytorch, Onnx and more.
- albertzeyer 3y agoCan you explain that? My understand of Triton is more that this is an alternative to CUDA, but instead you write it directly in Python, and on a slightly higher-level, and it does a lot of optimizations automatically. So basically: Python -> Triton-IR -> LLVM-IR -> PTX. https://openai.com/research/triton https://openai.com/research/triton
- chillee 3y agoIt's confusing, there's OpenAI Triton (what you're thinking of) and Nvidia Triton server (a different thing).
- jerrygenser 3y agoOriginal comment is referring to Nvidia triton inference server