6 ms·
I've been hosting Qwen3.8-27B myself. On my endpoint it's $0.30/1M in, $0.10 cache, $2.03 out - so those agent turns that re-send the same prefix get a lot chea
by anotherCodder 16d ago
I've been hosting Qwen3.8-27B myself. On my endpoint it's $0.30/1M in, $0.10 cache, $2.03 out - so those agent turns that re-send the same prefix get a lot cheaper when cache hits. UI at inference.tiyuvta.ai/app if you want to try it. Hosted is up to 210 tok/s and 280ms TTFT with reasoning off.