7 ms·
Is that something we have credible evidence for? Do they serve a better model when artificialanalysis (the benchmark site) is making the requests, and so on?
by kzrdude 6d ago
Is that something we have credible evidence for? Do they serve a better model when artificialanalysis (the benchmark site) is making the requests, and so on?
- varispeed 6d agoIt is my own anectodal. Few days something I worked on usually got one shotted or got quality result. Today it is very much going nowhere and is stuck in reasoning loops. Probably someone should build nerf tracker, because this is quite common that models get substantially worse once PR hype wears off and they quantise them more or simply route requests to older models with system prompt changed to say it is Astra and not Sol etc.
- EPWN3D 6d agoSo it's your own anecdotal experience, but someone should build a service to track it?
- singingtoday 6d agoThere is a service to track anthropic models. Not sure how accurate it is, but my vibe says somewhat.