11 ms·
I suspect the model is fundamentally the same underneath, but that various tricks like quantization are being performed in the deployed model to improve inferen
by ActivePattern 3y ago
I suspect the model is fundamentally the same underneath, but that various tricks like quantization are being performed in the deployed model to improve inference speed/cost at the expense of output quality.