4 ms·
You know what. As a big Anthropic fan that pays €120 a month for Claude Max x5 and maybe €30 a month for API access I got extremely annoyed at the extremely var
by Roark66 23d ago
You know what. As a big Anthropic fan that pays €120 a month for Claude Max x5 and maybe €30 a month for API access I got extremely annoyed at the extremely variable quality of service I'm getting.
It is not only that every single turn with opus on high reasoning (high is the middle setting) takes at least 5-7min. It is barely usable interactively. Instead of a chat it feels like you're sending emails to it. Tasks that used to take an hour when it "reasoned" for 45s before it started doing anything now take almost entire day.
At least until few weeks ago it was horribly, mind bogglingly slow (a little better during US nighttime), but the quality was still good. I could not do things interactively, but providing prompts were fine it built stuff fine.
This is no longer the case. It makes stupid errors all the time. So you cannot leave it to complete some work, for example write infrastructure migration scripts a night before then you simply run the scripts and perform the migration during the day. Nope, every single script has stupid issues requiring use of the model to fix them. As they are written in it's own "spaghetti code" fixing them by hand is not an option.
It is clear to me they are doing some shenanigans behind the scenes to try and optimise their compute use. Either they quantized these models dynamically or do other things that affect quality.
In top of that they now do this stupid fingerprinting.
Anyone who knows how output vectors are turned into tokens knows it will eat up a lot of compute or destroy quality.