5 ms·
Which models will this be able to run at an acceptable token/s rate?
by nik736 10mo ago
Which models will this be able to run at an acceptable token/s rate?
- simlevesque 10mo agogpt-oss:120b https://til.simonwillison.net/llms/codex-spark-gpt-oss https://til.simonwillison.net/llms/codex-spark-gpt-oss
- hamdingers 10mo agoAm I missing it or is there no information about performance? Looking for a tokens/sec
- simlevesque 10mo agoHe didn't give that info but the transcript linked at the end shows how much time was spent for each query.
- aseipp 10mo agoRight now I get 59 tok/sec on GPT-OSS 120B using Unsloth's dynamic 4-bit quants, via llama.cpp https://news.ycombinator.com/item?id=45881049 https://news.ycombinator.com/item?id=45881049