6 ms·
I would test this, might be cheaper per task even costing more per token, probably faster too
by repparw 14d ago
I would test this, might be cheaper per task even costing more per token, probably faster too
- sfink 13d agoI'm still leeching off the free tier, so it's going to be hard to beat the price. But yes, I intend to support several models, to handle the overload situation (automatic failover). And switch to a cheap paid plan, though it seems like that'll mostly improve rate limits, which barely matters for my usage. Faster is always good, though. I do care about latency.