5 ms·
For my application, I'm still happily using gemini-2.5-flash and the only problem is when it reports being overloaded. It's for interpreting a downscaled phone
by sfink 14d ago
For my application, I'm still happily using gemini-2.5-flash and the only problem is when it reports being overloaded. It's for interpreting a downscaled phone camera photo of a hand-written shopping list on a whiteboard, and it works stunningly well. My handwriting sucks, too.
(I guess the only relevance here is that if your problem matches a model's strengths, then you can do fine with a model that is several generations out of date.)
- brap 14d agoI believe the older models are being gradually phased out, newer ones have no availability issues
- repparw 14d agoI would test this, might be cheaper per task even costing more per token, probably faster too
- sfink 13d agoI'm still leeching off the free tier, so it's going to be hard to beat the price. But yes, I intend to support several models, to handle the overload situation (automatic failover). And switch to a cheap paid plan, though it seems like that'll mostly improve rate limits, which barely matters for my usage. Faster is always good, though. I do care about latency.
- w4yai 13d agoplease do yourself a favor and use something far more efficient ! GLM5.3 will make you super happy