6 ms·
Haven't read the full report yet, just a quick question. Are your numbers for cold start without pre fill or is it after warmed cache?
by narrationbox 28d ago
Haven't read the full report yet, just a quick question. Are your numbers for cold start without pre fill or is it after warmed cache?
- toebee 28d agoWe do graph capture etc at startup (same as vLLM) but this model variant doesn’t require prefix caching - the prefix is just 10 tokens.