4 ms·
The above tests were done with 24k context window. Testing was mostly driven by ChatGPT analyzing oMLX server logs and suggesting changes. Finally, I settled o
by akg_67 15d ago
The above tests were done with 24k context window. Testing was mostly driven by ChatGPT analyzing oMLX server logs and suggesting changes.
Finally, I settled on Qwen3.6-35B-A3B-4bit with 32,768 context window and 16,384 max tokens.
---
Additional results from Qwen3.6-35B-A3B-4bit (Can't edit previous comment)
Qwen3.6-35B-A3B-4bit, 329.7 PP, 41.3 TG