6 ms·
This is a pretty standard technique if you're running the models yourself. e.g. ChatGPT almost certainly does this. There's even work that is more sophisticate
by sshumaker 2y ago
This is a pretty standard technique if you're running the models yourself. e.g. ChatGPT almost certainly does this.
There's even work that is more sophisticated in this domain that allows 'template' style partial caching:
https://arxiv.org/abs/2311.04934 https://arxiv.org/abs/2311.04934