7 ms·
I'm only waiting for OpenAI to provide an equivalet ~100 USD subscription to entirely ditch Claude. Opus has gone down the hill continously in the last week (a
by andreagrandi 7mo ago
I'm only waiting for OpenAI to provide an equivalet ~100 USD subscription to entirely ditch Claude.
Opus has gone down the hill continously in the last week (and before you start flooding with replies, I've been testing opus/codex in parallel for the last week, I've plenty of examples of Claude going off track, then apologising, then saying "now it's all fixed!" and then only fixing part of it, when codex nailed at the first shot).
I can accept specific model limits, not an up/down in terms of reliability. And don't even let me get started on how bad Claude client has become. Others are finally catching up and gpt-5.3-codex is definitely better than opus-4.6
Everyone else (Codex CLI, Copilot CLI etc...) is going opensource, they are going closed. Others (OpenAI, Copilot etc...) explicitly allow using OpenCode, they explicitly forbid it.
This hostile behaviour is just the last drop.
- dannersy 7mo agoNo offense, but this is the most predicable outcome ever. The software industry at large does this over and over again and somehow we're surprised. Provide thing for free or for cheap, and then slowly draw back availability once you have dominant market share or find yourself needing money (ahem). The providers want to control what AI does to make money or dominate an industry so they don't have to make their money back right away. This was inevitable, I do not understand why we trust these companies, ever.
- NamlchakKhandro 7mo agobecause it's easier than paying $50k for local llm setup that might not last 5 years.
- dannersy 7mo agoWell, yes. They know what they are doing. They know when given the option the consumer makes the affordable choice. I just don't have to like or condone their practices. Maybe instead of taking on billions of dollars of debt they should have thought about a business model that makes sense first? Maybe the collective "we" (consumers and investors, but especially investors) should keep it in our pants until the product is proven and sustainable? It will be real interesting if the haters are right and this technology is not the breakthrough the investors assume it to be AFTER it is already sewn into everyone's work flows. Everyone keeps talking about how jobs will be displaced, yet few are asking what happens when a dependency is swept out from underneath the industry as a whole if/when this massive gamble doesn't pay off. Whatever. I am squawking into the void as we just repeat history.
- newswasboring 7mo agoOr the companies can be transparent about their product roadmap. I can guarantee this enshittification was on the roadmap way before we knew about it. They let us operate under false information, that's just weak behavior.
- player1234 7mo ago[dead]
- andreagrandi 7mo agoNo offense taken here :) First, we are not talking about a cheap service here. We are talking about a monthly subscription which costs 100 USD or 200 USD per month, depending on which plan you choose. Second, it's like selling me a pizza and pretending I only eat it while sitting at your table. I want to eat the pizza at home. I'm not getting 2-3 more pizzas, I'm still getting the same pizza others are getting.
- ifwinterco 7mo agoOpus 4.6 genuinely seems worse than 4.5 was in Q4 2025 for me. I know everyone always says this and anecdote != data but this is the first time I've really felt it with a new model to the point where I still reach for the old one. I'll give GPT 5.3 codex a real try I think
- kilroy123 7mo agoI agree with you. Codex 5.3 is good it's just a bit slower.
- andreagrandi 7mo agoIt is (slower), especially at xhigh setting. But if I have to redo things three times, keep confirming trivial stuff (Claude Code seems to keep changing the commands it uses to read code... once it uses "bash-read", once it uses "tree", once it uses "head" and I have to keep confirming permission), I definitely waste more time than give a command to codex (or in my case OpenCode + codex model) and come back after 10 minutes.
- mosselman 7mo agoI asked Codex 5.3 and Opus 4.6 to write me a macos application with a certain set of requirements. Opus 4.6 wrote me a working macos application. Codex wrote me a html + css mockup of a macos application that didn't even look like a macos application at all. Opus 4.5 was fine, but I feel that 4.6 is more often on the money on its implementations than 4.5 was. It is just slower.
- seu 7mo ago> Opus has gone down the hill continously in the last week Is a week the whole attention timespan of the late 2020s?
- marcus_holmes 7mo agooh shit we're in the late 2020's now
- testdelacc1 7mo agoSorry, I don’t agree. And I won’t be taking questions at this time.
- latexr 7mo agoWe’re still in the mid-late 2020s. Once we really get to the late 2020s, attention spans won’t be long enough to even finish reading your comment. People will be speaking (not typing) to LLMs and getting distracted mid-sentence.
- imafish 7mo agoSeems we're already there. My brain trailed off after "won’t be long enough to even finish"...
- mraart 7mo agoThat's still impressive, given your claim of being a fish...
- Bengalilol 7mo agoThe most impressive thing is that this looks like it is your only comment on this network ^^ ps: imafish may only be a fan of <https://mumband.bandcamp.com/track/if-i-were-a-fish https://mumband.bandcamp.com/track/if-i-were-a-fish>
- eamag 7mo ago
- neya 7mo agoIt's the most overrated model there is. I do Elixir development primarily and the model sucks balls in comparison to Gemini and GPT-5x. But the Claude fanboys will swear by it and will attack you if you ever say even something remotely negative about their "god sent" model. It fails miserably even in basic chat and research contexts and constantly goes off track. I wired it up to fire up some tasks. It kept hallucinating and swearing it did when it didn't even attempt to. It was so unreliable I had to revert to Gemini.
- resiros 7mo agoIt might simply be that it was not trained enough in Elixir RL environments compared to Gemini and gpt. I use it for both ts and python and it's certainly better than Gemini. For Codex, it depends on the task.
- deleted 7mo ago[deleted]
- cactusplant7374 7mo agoNo developer writes the same prompt twice. How can you be sure something has changed?
- kasey_junk 7mo agoI regularly run the same prompts twice and through different models. Particularly, when making changes to agent metadata like agent files or skills. At least weekly I run a set of prompts to compare codex/claude against each other. This is quite easy the prompt sessions are just text files that are saved. The problem is doing it enough for statistical significance and judging the output as better or not.
- baq 7mo agoRalph Wiggum would like a word
- cactusplant7374 7mo agoSame prompt assumes same context state. But I think you get what I mean.
- andreagrandi 7mo agoI suspect you may not be writing code regularly... If I have to ask Claude the same things three times and it keeps saying "You are right, now I've implemented it!" and the code is still missing 1 out of 3 things or worse, then I can definitely say the model has become worse (since this wasn't happening before).
- co_king_5 7mo ago[dead]
- andreagrandi 7mo agoI haven't experiences this with gpt-5.3-codex (xhigh) for example. Opus/Sonnet usually work well when just released, then they degrade quite regularly. I know the prompts are not the same every day or even across the day, but if the type of problems are always the same (at least in my case) and a model starts doing stupid things, then it means something is wrong. Everyone I know who uses Claude regularly, usually have the same esperience whenever I notice they degrade.
- abm53 7mo agoI’m unsure exactly in what way you believe it has gone “down the hill” so this isn’t aimed at you specifically but more a general pattern I see. That pattern is people complaining that a particular model has degraded in quality of its responses over time or that it has been “nerfed” etc. Although the models may evolve, and the tools calling them may change, I suspect a huge amount of this is simply confirmation bias.
- super256 7mo agoOpenAI forces users to verify with their ID + face scan when using Codex 5.3 if any of your conversations was redeemed as high risk. It seems like they currently have a lot of false positives: https://github.com/openai/codex/issues?q=High%20risk https://github.com/openai/codex/issues?q=High%20risk
- andreagrandi 7mo agoThey haven't asked me yet (my subscription is from work with a business/team plan). Probably my conversations as too boring
- stogot 7mo agoTry something not boring and see what happens?
- bbstats 7mo agoall this because of a single week?
- andreagrandi 7mo agoNo, it's not the first time their models degrade for some time.
- GorbachevyChase 7mo agoI was underwhelmed by Opus4.6. I didn’t get a sense of significant improvement, but the token usage was excessive to the point that I dropped the subscription for codex. I am suspect that all the models are so glib that they can create a quagmire for themselves in a project. I have not yet found a satisfying strategy for non-destructive resets when the systems own comments and notes poisons new output. Fortunately, deleting and starting over is cheap.
- trillic 7mo agoThe rate limit for my $20 OpenAI / Codex account feels 10x larger than the $20 claude account.
- choilive 7mo agoYES. I hit the rate limit in about ~15 mins on Claude. But it will take me a few hours with Codex. A/B testing them on the same tasks. Same $20/mo.
- WarmWash 7mo agoMy favorite conspiracy explanation: Claude has gotten a lot of popular media attention in the last few weeks, and the influx of users is constraining compute/memory on an already compute heavy model. So you get all the suspected "tricks" like quantization, shorter thinking, KV cache optimizations. It feels like the same thing that happened to Gemini 3, and what you can even feel throughout the day (the models seem smartest at 12am). Dario in his interview with dwarkesh last week also lamented the same refrain that other lab leaders have: compute is constrained and there are big tradeoffs in how you allocate it. It feels safe to reason then that they will use any trick they can to free up compute.
- thepasch 7mo ago> I’m only waiting for OpenAI to provide an equivalet ~100 USD subscription to entirely ditch Claude. I have a feeling Anthropic might be in for an extremely rude awakening when that happens, and I don’t think it’s a matter of “if” anymore.
- submain 7mo ago> And don't even let me get started on how bad Claude client has become The latest versions of claude code have been freezing and then crashing while waiting on long running commands. It's pretty frustrating.