10 ms·
Maybe play around with CPU inference on smaller models too. Frontier models are probably overkill for simple language responses at this point; I suspect you cou
by ramesh31 1mo ago
Maybe play around with CPU inference on smaller models too. Frontier models are probably overkill for simple language responses at this point; I suspect you could get pretty good results with Qwen doing this, and Kokoro TTS is really good now.