7 ms·
Why not? It's caching the state of the model after the cached prefix, so that inference workload doesn't need to be run again.
by maged 2y ago
Why not? It's caching the state of the model after the cached prefix, so that inference workload doesn't need to be run again.