8 ms·
Ok long time Claude Code user here; lately I've started to realize there's other great models out there I should be trying, but I'm hesitant to leave Claude Cod
by Syntaf 2mo ago
Ok long time Claude Code user here; lately I've started to realize there's other great models out there I should be trying, but I'm hesitant to leave Claude Code behind for something new.
What's the consensus today on codex vs claude code, does it really matter anymore?
- indy 2mo agoYou wouldn't be leaving Claude Code, just trying something new. If you don't like it just resume using Claude.
- Daedren 2mo agoUse a harness that doesn't lock you into a moat, like OpenCode.
- AntonyGarand 2mo agoCan't use a claude code subscription in another harness though
- greenavocado 2mo agoYou absolutely can; they are not banning anymore. The bigger problem is that subscription versions of the models are way crappier than when the "same" model is hit via API (Bedrock/Vertex) You can also make it not count against extra usage. OpenCode docs show it because Anthropic specifically ambushed them with a PR to remove support so simpletons can't use it easily.
- infberg 2mo agoDo you have a source for that?
- deleted 2mo ago[deleted]
- AntonyGarand 2mo agoThe opencode docs[0] still say otherwise, do you have a source? [0] https://opencode.ai/docs/providers/#anthropic https://opencode.ai/docs/providers/#anthropic
- llm_nerd 2mo agoThey aren't banning it anymore, they just make it count as "extra usage". e.g. you're paying for every token in addition to your subscription. Further, the claim that the subscription "version" of the model is worse sounds like bullshit (and the sort of anecdotal nonsense that you see on sites like this). Do you have anything substantiating this?
- ghostpepper 2mo agoI am curious about the claim that the subscription models are different. Has anyone benchmarked this?
- retinaros 2mo agomaybe he meant about server-side features? otherwise its dumb. OAI and Anthropic are using azure/aws/gcp...
- AlexCoventry 2mo agoHow does that work? Doesn't that mean Microsoft/Amazon/Google have full de facto access to OAI's and Anthropic's model weights and operational processes? This is one reason it surprised me that Anthropic decided to run stuff on Musk's hardware. It seems overwhelmingly likely that the new Grok release is informed by what Musk has been able to learn from that relationship.
- zakisaad 2mo agoThis is not at all aligned with my experience. Do you have a source for this?
- deleted 2mo ago[deleted]
- I_am_tiberius 2mo agoDon't use providers that don't allow it.
- kristianc 2mo agoYou can however for now use wrappers which are not harnesses such as T3Code though. They were going to cut under the Programmatic API, but have at least temporarily walked it back.
- thebigspacefuck 2mo agoFWIW Claude Code works with OpenRouter so you can use any model.
- theturtletalks 2mo agoCC system prompt is bloated, use Pi to test Codex instead
- maxloh 2mo agoCodex CLI is open source too. I don't think there is a difference.
- akmarinov 2mo agoCodex is open source and lets you use any model https://learn.chatgpt.com/docs/config-file/config-advanced#oss-mode-local-providers https://learn.chatgpt.com/docs/config-file/config-advanced#o...
- maxloh 2mo agoOnly its CLI is open source. The desktop app is always proprietary.
- magicalhippo 2mo agoYou can use Codex with any endpoint compatible with OpenAI Response API[1], like llama.cpp. [1]: https://unsloth.ai/docs/basics/codex https://unsloth.ai/docs/basics/codex
- moomoo11 2mo ago[dead]
- CuriouslyC 2mo agoIt never really mattered (except when codex was very new). If anything, codex's remote session integration is better, so outside of some "ultracode" orchestration bells/whistles where Claude Code is ahead, I think Codex is a better tool.
- mountainriver 2mo agoAgree, I think there was just a blind study that showed no one could tell the difference even though the users were avid they could
- cmrdporcupine 2mo agoI left Claude for Codex months ago. I was an early Claude Code adopter but I have found Codex consistently better since about the February time frame. And far more reliable. It's more diligent and empirical and results focused, and less creative. It sometimes needs a kick to avoid a Zeno's paradox of incremental steps to get to the goal. But it produces more reliable code with fewer race conditions, unhandled negative cases, etc. It's also better value from a $$ POV, or at least has been. This fluctuates a bit. You're also free to use your Codex subscription with other harnesses, like opencode, etc. Unlike Anthropic. Plays better with others.
- dgritsko 2mo agoI use Conductor pretty much exclusively and it makes it incredibly easy to try different models, even within the same workspace - definitely recommend giving it a shot. Whenever I'm forced to use the Claude Code app directly it just seems woefully inadequate compared to Conductor
- wyre 2mo agoClaude Code is a massively bloated agent harness. Try Pi: https://pi.dev/ https://pi.dev/
- blovescoffee 2mo agoPi is so “unbloated” that it’s extra effort to use. You can decide how much work to put into it. I get the trade off. But this is a big jump from CC. I’d recommend some middle ground like opencode.
- trollbridge 2mo agoIt's worth trying out OpenCode, then oh-my-pi, and also the commercial harnesses like Codex. (I haven't yet bothered to try Antigravity and have no interest in Gemini-cli now that it's not available except on expensive plans.) pi is also worth tinkering with, particularly if you have an eye towards automating some things.
- lubujackson 2mo agoEven simpler, use Cursor with any frontier model. I see others sweat to add enough context to Claude Code while Cursor has a ton of contextual awareness, uses subagents automatically and is significantly faster with no drop off I have found. I'm not sure why devs are so enamored with living in the CLI, but Cursor has one of those too.
- blovescoffee 2mo agoSimilar to running arch Linux. Many people do need to. But many people just like tinkering. Tinkering can lead to positive outcomes but it’s usually not “doing work”.
- ejpir 2mo agowhat are you building that doesn't work with read,write,edit,bash and skills? Genuinely interested.
- AgentMasterRace 2mo ago
- bakies 2mo agoI spent the last couple days switching because Anthropic keeps locking stuff behind API pricing. OpenAI lets you do anything with your sub right now. I'm building headless and web interfaces around Pi.dev. I had this previously with Claude Code but they are going to lock away all those features. I think the Claude does a better job at being proactive to solving things, but I'm going to keep tweaking my harness to nudge gpt to do more in it's turn. Not sure!
- postalcoder 2mo agoThere is so much less drama involved with the Codex world. You don't realize how oppressive CC is until you've escaped it. Outages, weird restrictions, degradation, accelerated usage, etc etc etc.
- timcobb 2mo agoTotally. My experience as well. After some time with codex you're like come on Claude can you just stfu! Haha. I now almost always instruct Claude with specific length requirements when I ask questions. Otherwise, it just blathers and blathers in the most annoying of ways. "Oppressive" is spot on in my opinion
- jakswa 2mo agoI'll agree and expand on "weird restrictions" -- I used to check the claude usage graphs multiple times a day to see where I'm at on my weekly budget. With gpt 5.5 I don't think I'm working differently but haven't felt the need to check anything because I think I've hit my limit... once? on some egregious edge case scenario iirc
- rafaelmn 2mo agoSame here - it's probably that OpenAI needs to buy goodwill with developers to infiltrate corps and Anthropic is trying to squeeze the lead into revenue. The only question is - how much longer can OpenAI burn money before it needs to start showing signs of profitability
- Amir6 2mo agoLet alone getting banned out right with no reason, zero updates after weeks, and not even being able to download your chat history (despite the feature being available (I assume they vibe coded it and it does not work!). My story below; https://news.ycombinator.com/item?id=48597861 https://news.ycombinator.com/item?id=48597861
- Cider9986 2mo agoEven less drama with open models like GLM.
- arikrahman 2mo agoFor me the biggest shift was using Deepseek through an American provider with reasonix as the harness, making cache hits at a rate of practically free.
- tandr 2mo agoWhich provider do you use, if you don't mind sharing? I tried Digital Ocean (looking for Zero retention and no training), but their context limits are rather small for DeepSeek inference.
- arikrahman 2mo agoI used some others but eventually decided to use Deepseek directly. I don't mind giving China my data, I'm more concerned about the American government imprisoning me than a foreign nation an ocean away.
- saberience 2mo agoThey are both excellent but excel in different areas. Fable is super super proactive and great for doing a LOT of work with a single prompt, also for creative work. Codex is more details focused, often catches wonky bugs and correctness issues that Fable misses, feels more terse and less "friendly", more like a stern senior engineer versus a friendly talkative engineer (Claude). Codex is also better if you're already an engineer, Claude is better for non-engineers. I.e. Codex works better if you know exactly what you want and know the right way of explaining it.
- wahnfrieden 2mo agoWith the exception of Fable which is going away anyway, Codex is better especially after the last couple Opus releases. It’s also no longer slower than Claude. You get much more generous usage from the 20x plan. And you get far better uptime. If benchmarks and early tester impressions are accurate, you also get access to Fable level capability at greater speed and lower cost (included in subscription).
- petesergeant 2mo ago> Fable which is going away anyway $2 says nah. You can't take Fable away in a week where GPT-5.6 and Grok 4.5 launch, if you want to hold on to customers.
- mortenjorck 2mo agoThe fact that they already extended subscription Fable once would suggest it won’t be solely locked behind API next week, but at the same time it really does look like they are doing everything they can to avoid serving it continuously at scale. Knowing Anthropic, this unfortunately might end up meaning a quietly quantized Fable on subscription.
- wild_egg 2mo agoCan anyone explain this "quietly quantized" model idea to me from a business perspective? Coca-Cola doesn't "quietly water down" its product to save a few bucks. They know people will take a sip, say "oh that's not what i wanted", and go buy a Pepsi. If they serve me a quantized Fable, I'm just going to think Fable sucks and go get my tokens elsewhere. What's the point?
- petesergeant 2mo agoPepsi may not water down its product, but your local diner may well decide to put a little less ice-cream in its milkshakes if it thinks it can get away with it and most people won't be able to tell.
- kristianc 2mo agoCodex has been good for a long time, more expensive but very focused on efficiency. Working with it feels faster and more to the point than Opus models and I trust it more with long-running jobs. Also regular resets vs being at the whim of Anthropic drama all the time is hella nice.
- anukin 2mo agoCodex is cheaper on average no? I think the models are expensive but the token efficiency of the harness itself solves the problem.
- kristianc 2mo agoYes that's what I meant, the per token cost is higher but as you say the efficiency levels it out / works slightly in Codex's favour.
- timcobb 2mo agoAnyone know what the deal is with the resets?
- kristianc 2mo agoThey've discovered it's a good marketing strategy. Whenever there's an outage, or a new launch, there's often a reset with it, which helps keep people engaged with OAI / Tibo and reduces churn. They've also introduced banked resets, which are really clever. If you have a $200/month plan and three banked resets, you're not churning because you will overweight giving up those resets (loss aversion theory).
- timcobb 2mo agoI ran out of resets :( hehe I had 3 and used them all
- prospector1065 2mo agoTry OpenCode and you can point it at either model
- sidrag22 2mo agoPersonally, I started using openai models to mess with other harnesses. I was pretty oppositional to CC and how they don't let you kinda plug and play freely, or give transparency into -p usage with other harnesses. So i mix and match a bunch of openai and some chinese models im trying out into opencode. I keep hearing codex is great, on the tier of current CC, I've tried it and it just ate my entire 5 hour usage window looping without asking for clarification on something and none of it was usable. that was the only time i tried codex as i could got that same task done with maybe 20% of my window with my existing openai opencode workflow. I had put a decent amount of effort into setting up that initial codex attempt and it went so poorly that i've been entirely uninterested in trying again. This was maybe a month or so ago, and i know stuff moves fast, but for me, i like the models, dont care for the harness.
- thebigspacefuck 2mo agoIMO LMArena is the best benchmark that avoids benchmaxxing https://arena.ai/leaderboard/agent https://arena.ai/leaderboard/agent 5.6 isn’t on there yet but Fable leads by a significant margin atm
- TranquilMarmot 2mo agoThe results here match up to my real-world experience using these models every day at work and switching between them regularly.
- NiekvdMaas 2mo agoIn my experience, for coding Codex is definitely far ahead of Claude Code, even when using Fable 5 as a model.
- postflopclarity 2mo agoyou have a very strange experience
- arcanemachiner 2mo agoPeople work on different things, fail to mention their field of usage in the comments, and then misunderstand the experiences of others who do the same. Repeat ad nauseam.
- AgentMasterRace 2mo agohe's writing novels
- ralusek 2mo agoI use both constantly for different things. You don't need to be a one-model Andy
- qaq 2mo agoI use both especially for checking each others work. Pretty happy with results
- catketch 2mo agoNot sure there's going to be a consensus, but I can tell you that when i have codex review claude-written code, it finds important gaps and fixes. The reverse is also true. Both are powerful, but even better when used in combination
- Jtarii 2mo agoLiterally every top model is identical and anyone saying otherwise is engaging in astrology.
- postflopclarity 2mo agoanybody saying they're identical clearly doesn't use both...
- iugtmkbdfil834 2mo agoHonestly, I would even push it further. People who would claim that don't use either one.
- Supermancho 2mo agoThe outputs, ui, and overall behavior (tokenization) are not identical.
- vessenes 2mo agoMore literal, less fluid verbally, harder time understanding nuance, more correct code, fewer bugs. Less pretty UI. I switch back and forth but find I have less 'clean up' work with codex; more upfront communication though to properly specify. High hopes for 5.6!
- 217 2mo agoConsensus itself does NOT matter, omp is objectively the best harness for power users yet it has 0 hn posts about it, zero. You're fully free to use and try anything and without caring about what others think is right
- efficax 2mo ago"objectively the best"?
- saberience 2mo agoProbably means subjectively according to his own opinion...
- incognito124 2mo agoTo quote a friend: > Well it's objective _to me_
- trollbridge 2mo agoomp is really good. I have one non technical people in my firm using it. One is using it to assist with editing books, basically using it to gather up manuscripts from e-mail / Google Doc etc. submissions, and then switch models between a cheap one and Opus (for actually analysing the manuscript). The other non-technical person has done really surprising things with it AI, like a long-running GPT 5.5 Pro chat session which is basically her expense tracker - it has an .xlsx file "carried" in the chat, and she just tells ChatGPT (or scans a receipt) whenever she has a new expense, and then prompts it in natural language when she needs a report. I'm looking forward to seeing what she can do with omp.
- joe_mamba 2mo ago> omp is objectively the best harness for power users Care to detail this?
- saberience 2mo agoHow can this be "objective"? Surely its subjective. I've tried a fuck load of harnesses but keep coming back to Codex as my harness.
- sk4rekr0w 2mo agoCodex app is a much different experience than CC CLI. I would try it out for a couple days with the new model suite and see what you prefer after that.
- simianwords 2mo agoCodex UI is way way way better than Claude Code - codex UI is much more responsive - i get feedback about the progress easily - the tool calls and results are very legible, I can click them and see the progress - no one talks about this but the tool call and response notification are handled much more elegantly in Codex. In Claude Code, it is handled in a clunky way using loops which always causes some delay - you can steer the conversation midway in Codex - /side is underrated (/btw is the equivalent and is much worse in Claude Code) - I have to admit subagents are handled better in Claude Code
- timcobb 2mo agoIf you can afford it and you have something to justify the expense, I would get both. they're interesting to run side by side, you can hand things off from one to the other. Pretty neat. Unfortunately now I just want to have both :(
- xur17 2mo agoI personally use opencode so I can swap between models and try different options. I'd say I prefer claude (fable and opus 4.8) so far, but curious to see where gpt 5.6 lands. For personal stuff, I've been pretty happy with chatgpt's $20 plan. I believe it has considerably higher limits than claude's $20 plan, and it's enough for the personal stuff I play with (hermes, and some small coding stuff). Also allows me to keep up to date on openai models.
- timcobb 2mo agoThe $20 GPT plan with GPT 5.5 lasted me, somehow, exactly one smallish fixup feature
- xur17 2mo agoWhich limit? Weekly or 5 hour? I've been using it with hermes and some coding (with opencode), and I am getting a LOT more than one feature out of it, but the work is spread throughout the week.
- drschwabe 2mo agoWhen you use opencode to use Claude and Chatgpt models ie- Fable or GPT 5.6 I assume you are getting billed on pure credits? I was under the assumption that you can only use your Claude/ChatGPT paid plans when using Claude Code or Codex. ie- that you would be paying way more via opencode since you would not be getting the extra limits subsidized by your paid plans.
- nilkn 2mo agoCodex has arguably been better than Claude Code for months now, but it's flown under the radar because it just didn't capture the same viral marketing effect and OpenAI in general has had more optics / PR issues than Anthropic amongst the online developer crowd. I use the word "better" not in the sense that the underlying GPT models are fundamentally smarter or more intelligent, but rather that as a product Codex is just simpler, cheaper, and abundantly reliable and low-drama.
- hk__2 2mo agoI’d argue the opposite. I’ve switched back and forth from one to the other and Opus/Fable has been constantly better than any GPT in my daily work. It’s a bit slower but it does the things right, with as little code as possible, some comments where needed. Codex is faster but you always have to correct it because it got something wrong; it writes tons of code ("let me add a small helper") with obvious comments.
- nilkn 2mo agoI'm not sure how meaningful this is. Fable only just recently become more broadly available, and GPT-5.6 is launching broadly today.
- hk__2 2mo agoThe comment I was responding to was talking about Codex usage in the past few months. This is a general feeling about Codex with Claude, not a model-to-model comparison.
- deleted 2mo ago[deleted]
- ljm 2mo agoPurely anecdotally the one persistent issue I have with LLMs writing code is that they are absolutely paranoid and add a load of indirection and defensive crap and even if you prompt to avoid that it will often require manual steering to remove the cruft.
- giancarlostoro 2mo agoLast time I tested Codex on a cheap plan, it barely lasted an hour? I think this was for the $20 plan. I was afraid to try the more expensive plan after that. Not sure, I might just outright rip my Claude Code bandaid if the current usage quotas do die off after the 17th or whatever date they said they would "return on".
- maxloh 2mo agoI wish they open source their desktop app and built-in skills one day. That would be a final blow for me.
- maxloh 2mo agoEdit: Found that their built-in skills are actually open-source: https://github.com/openai/plugins https://github.com/openai/plugins
- wwind123 2mo agoI've been using Claude Code, Codex, Gemini (now Antigravity) at the same time for half year now, ever since I dipped my toe into agentic coding. I'd say in general Claude Code and Codex are equally powerful, Gemini is lagging behind. One thing I appreciate with Codex is, OpenAI nowadays sometimes just gives you quota resets you can bank, so when you use up weekly quota before the week ends, you could just reset the quota, to continue using Codex. I've been much less anxious about Codex quota because of this perk. I just used one reset in the bank yesterday, and still have 3 resets left. Whereas with Claude, when you've used 95% quota 3 days before the week ends, you'd be much more anxious. On the other hand, Claude Code's /remote-control mechanism is extremely helpful when I am running it in the cloud and wants to monitor it or control it on my phone. Codex currently doesn't support this kind of usage. Codex only allows you to use your phone to connect to a session on your desktop, not in the cloud.
- WhitneyLand 2mo agoCodex is supported well on iPhone/iPad, it’s inside the ChatGPT app. It’s amazing how much work you can get done on your phone now, especially if you already have a design mapped out in your head.
- ghostpepper 2mo agoI have used claude and codex extensively but only from their CLI app (heavily sandboxed using rootless podman, network filtering, etc), so I don't really know what I'm missing with the GUI apps. One killer feature that Claude has, and AFAIK Codex still lacks, is the ability to start a session in the terminal and then hand it off (actually just remotely control it), from the iOS app. Last time I tried Codex on iOS it required a ton of set up to link a github project etc. The way claude lets me remote into a session I've already started on my actual machine is much better IMHO.
- akmarinov 2mo agoThey’ve addressed that. Codex in the ChatGPT app on iOS is way better than Claude Code now. You sign in the Codex app on your Mac same on iOS and are able to completely control your sessions - fork, side chats, plugins - everything. It’s really great i often work through it. And you can connect any number of Codex instances on any number of macs and then manage them all through the iOS app.
- petesergeant 2mo agoIn my projects, Claude writes and Codex reviews, and I've had a lot of code I've been very happy with out of that, although as of today, Grok _also_ reviews, and finds interesting new stuff.
- mountainriver 2mo agoThere was just a study showing that when presented blindly no one could tell the difference yet users were avid they could
- hk__2 2mo agoThere _is_ a difference in the way Claude and GPT write. Last Friday I felt Opus was becoming dumb because it was writing like GPT.
- aroman 2mo agoClaude Code fan here... Codex is very good. Sometimes better. The killer feature is price. After 6+ months of exclusive Claude Code usage, I was begrudgingly forced to try Codex once Anthropic rejiggered their limits such that I kept maxing out my $200/mo plan in just a few days. These days I pay both $200/mo plans, and it's just about enough to get me through a week's work (small game studio - infinite code to write!)
- ValentineC 2mo ago> (small game studio - infinite code to write!) Curious: what multiplier do you think your productivity has increased by, from before AI?
- aroman 2mo agoIn terms of ability to ship? Easily tenfold. We literally ship 10 times more than before AI. This does not, however, translate into a tenfold increase in actual business success, of course :)
- winrid 2mo agoYes, because now your competitors do the same. The winners: inference providers
- shwaj 2mo agoInference providers, sure, but wouldn't we expect customers to win also?
- winrid 2mo agooh for sure! I think in some ways LLMs have raised the bar in terms of what you an expect from software (once we get past the hurdle of increased bugs, but that seems to be getting better)
- ValentineC 2mo ago
- rib3ye 2mo agoIt's not clear replies to this thread aren't openAI employees or incentivized influencers, but every benchmark has gpt-5.5 underperforming opus 4.8, sometimes by as much as 10%. Can they all be wrong/paid-off?
- deleted 2mo ago[deleted]
- deleted 2mo ago[deleted]
- novaleaf 2mo agoI sub both codex and claude at 20x. I like opus+fable more than gpt5.5 because it seems gpt tries to finish tasks by leaving any ambiguity unresolved. claude seems better at surfacing open questions. This is using the same AGENTS.md prompts, which were designed firstly for Claude use, so maybe it's something that could be optimized better if I understood gpt as well?
- freely0085 2mo agoIs it you have Fable delegate work to Sol? How do you do that? Do you run it in Codex/Desktop app?
- novaleaf 2mo agono, I didn't get access to sol until a few hours ago. I just have my claude protocol files linked to inside codex. trying Sol this morning for the first time, so I can't really comment on that vs gpt5.5. However you can do what you are asking "fable--> sol" you need to setup a mcp or have fable run a bash tool, just invoke the `codex.exe` cli tool with whatever cmdline args are needed.
- znpy 2mo ago> I'm hesitant to leave Claude Code behind for something new. Codex and Claude Code are not mutually exclusive, you can use both.
- erichocean 2mo agoI use Claude for planning, writing CRs, and code review. Codex writes all of the code, no exceptions. Works great, especially when you ask Claude to break up large CRs into roughly 10 minutes of Codex work each.
- silksowed 2mo agoSame here. I find the design, architecture, system design discussion to be better on Claude, but after Opus 4.6 I switched over to Codex for actual coding and love the results. I use both via the CLI and generally tell Claude to output the result of our decisions as a markdown that will be easy to read and implement by an agentic coding tool. Then I fire up Codex and read said markdown as the input of the session and way to build all the appropriate context needed. I see this as a way to step into letting the agents go run on their own and interact with each other, but I still like to steer so I put these manual steps in the flow. Letting the agents go off on their own and one shot big chunks is not reliable enough yet imo.
- epolanski 2mo agoI do exactly the opposite.
- erichocean 2mo agoI think the key is to get two LLMs looking at the same problem. I use Codex because it's better at the kind of code I need written (math-heavy, 3D geometry code). But if I was doing mainly UI code, I would do the opposite.
- jghn 2mo agoI have found Claude Code to be so much better than other common harnesses that it's kept me solely in the Anthropic ecosystem.
- YuechenLi 2mo agoThe answer is it depends. Claude's generally better at frontend and debugging tasks, while Codex is stronger at backend features and exploratory work. They have very different coding styles and thus very different strengths.
- TranquilMarmot 2mo agoAny actual data backing this up? Or is this just your personal experience?
- YuechenLi 2mo agoJust personal experience, I just find it way easier to do frontend work with Claude than it is with Codex.
- hraxz 2mo agoNo, this is all nonsense. It is so hard to tell at this point between the models to make generalizations like this. Just complete nonsense.
- TranquilMarmot 2mo agoI agree. Some sort of weird placebo effect that people get.
- osigurdson 2mo agoI use both. Not because I am cool, but because it is cost effective for personal projects with two $20 / month plans. It is also nice to be able to see what the state of the art is like for both. Personally, I find it very interchangeable. I open codex --yolo or claude with whatever there yolo flag is (have an alias).
- urams 2mo agoIf you can afford to test it seriously, running both in parallel, it's worth a test to see which you prefer. If you can't, don't bother. You're not likely missing anything since they are close to personal preference with most people I know who have meaningfully tried both preferring Claude
- alexhans 2mo ago> What's the consensus today on codex vs claude code, does it really matter anymore? Consensus is probably the wrong word for the popular opinions reflected in HN that you might get. I would recommend that you have 2 of each at all times when it comes to AI so you don't necessarily become overly locked to quirks of one thing. You'll soon realize that things move so fast that you just start internalizing common patterns instead of depending on one specific vendor. I recommend that you try pi and codex besides claude, to get your own feel for it.
- pkulak 2mo agoI can't tell the difference between Fable and GPT 5.5. I tried Fable while it was in trial $20 mode, used up my whole quota, and it was great, but as soon as I went back to GPT 5.5, everything was the same. But what I love about Openai is that they still let you hook OTHER harnesses up to a subscription. My Pi setup has been built up for a few months now into exactly what I want and moving over to CC or even Codex is really annoying. Caveat: I vibe code in tiny little chunks. I see what I want to do, and exactly how I want it done, then prompt that, refine, what was output, then repeat. I bet Fable is better at building a whole app from a 2-sentence prompt; but that's just not important to me at all.
- akmarinov 2mo agoSame here - gave 5.5 a web design to implement and it sucked. Gave the same to Fable and it still sucked.
- AgentMasterRace 2mo agodid you use Claude design, their tool meant for Web design? because if not then you're the problem .
- edumucelli 2mo agoNot sure about the consensus, but during an entire week I have done every task on my workplace with both Opus 4.8 and GPT 5.5. GPT won hands down. I would even sometimes copy the plans and solutions (using different Git worktrees) from GPT and paste it on Opus and itself would say GPT plans were better. At that point I have migrated. Fable is not enabled in our workspace so I have not tried. Claude lost my trust around February this year when the plan would say nonsensical things as "delete this method" that was clearly a key method on that part of the codebase. For personal projects I am using Codex 20$ plan and when that is over I use DeepSeek which is insanely good for the cost.
- InsideOutSanta 2mo ago> does it really matter anymore? They're different models with different philosophies behind them. This is anecdotal with a user group of 1, but in my experience: Claude has a stronger personality and is more creative. If you give it vague instructions, it's better at filling in the blanks with reasonable ideas. GPT-5.5 is better at following instructions. If you know exactly what you want, it will do it without going off the rails. It's also less likely to imply that you're dumb, but I don't really care about that. Some people do.
- akmarinov 2mo agoI’ve found that Claude is very literal. When I talk to 5.5 it gets what i want it to do, when I talk to Opus 4.8 it does what I say literally and doesn’t get the intent behind it.
- mingqiz 2mo agoThis. Claude is very good at instruction and treat my mistake in prompt as law. Gpt is just smarter at understanding my intent.
- pkulak 2mo agoI wish models called me out more! I can’t count the number of times I’ve had an absolute shit plan, prompted it, and the model built it to a tee. Then I look at the code, and it’s an absolute mess, because it had to make my stupid idea work. Maybe I need a better system prompt.
- athrowaway3z 2mo agoThey blocked Claude from being used in a different harness as well squeezed the usage like crazy. Switched to Codex and haven't cared since. Between the two the biggest difference by far is ... getting your harness / AGENTS.md / skills / tools set up right.
- 3371 2mo agoMy experience is that Codex's auto review is extremely costly, with $20 on both sides, I can run CC with auto mode for longer than with Codex's auto review enabled. Also in my own experience Claude's usage is actually bigger than Codex, but I am not sure if that's due to I stick to 5.5 with Codex while keep Sonnet as the default to orchestrate other models in CC.
- Razengan 2mo agoI've subscribed to ChatGPT/Codex for over a year and tried a Claude sub twice 1 month each, with a gap of several months in between. I tried them both side by side, mostly for reviewing existing Godot/GDScript code, or sometimes generating Swift Mac apps, including converting ancient relics I wrote eons ago in Visual Basic on Windows Codex was consistently better than Claude: https://i.imgur.com/jYawPDY.png https://i.imgur.com/jYawPDY.png Besides the useless "This is good" findings while reviewing and the excessive "oops you're right" backtracking, Claude's atrocious UX and borderline "spyware" make me never want to try an Anthropic product again for a long long while.
- firemelt 2mo agojust try it you will back to codex because gpt is trash, I ask for refund under 7 hours
- linsomniac 2mo agoI'm also a long-time Claude Code user here, though the last 3 weeks I've been doing loops having claude use codex to review until they reach consensus; uses tons of tokens but the result is really good. I'm trying Codex as my primary the last day or so, because I'm at 98% use and reset in 3 days on Claude. I'm worried about a lot of our skills and CLAUDE.mds and the like getting lost unless I migrate them, but otherwise codex seems to be working great.
- nvarsj 2mo agoThe harness is so much better than cc which is a buggy mess. Gpt is also way faster than Claude. I’ve been using gpt for a while now and I know a lot of people that swapped away from Anthropic for multiple reasons. However - fable still seems to be the best coding agent, it’s just slow and the harness sucks. So I still use it in some rare cases like to review codex. I’m hoping 5.6 lets me drop it entirely.
- killix 2mo ago[flagged]
- ThunderBee 2mo agoIME it entirely depends on your work. I find myself using both daily for different things. Codex with GPT 5.5 is much better at general SWE tasks but Claude Code with Opus is far better at complex reasoning tasks like reading and summarizing research papers, replicating experiments, identifying research gaps and proposing interesting follow ups.
- yokoprime 2mo agoI prefer codex for most tasks, but stil use Claude if i need to make something "nice but generic", i.e. a html artefact or touch up of front end code.
- davidhs 2mo agoI recommend trying Codex too. In fact, I recommend running them side-by-side if you have the budget, e.g. have both independently plan the same feature or implement in a different worktree, or have them critique each other's work. I personally find GPT-5.5 to be a better programmer than Opus 4.8, it is extremely thorough, but I don't like the code it generates ("austere"), and find Opus 4.8 to write more "human friendly" code. The programming comments GPT-5.5 makes is pretty awful where-as Opus 4.8 is good. I feel like Opus 4.8 is better at grasping my intention than GPT-5.5, and honestly find GPT-5.5 to be kind of "autistic". I do prefer the language (not the writing) of GPT-5.5, as I find the philosophical flowery language of Opus 4.8 kind of annoying. I have only managed to try Fable 5 a little bit, which feels like a much more generally smarter version of Opus 4.8, that is much better a programming and grasping your intention, and I think even the intention of your code, and is _really_ good at spotting bugs or problems with logic in your code. It feels wicked smart but is extemely expensive. It feels smart in the sense like it has a "bigger brain" and is much more sensitive to subtleties/details. These are different "brains", have different "personalities", etc. I think the best thing is to develop a feeling for it yourself.
- novaRom 2mo agoI haven't tried Codex yet, but I for my tasks GPT-5.5 may correctly point to a proper direction but its code feels a bit weird. Opus 4.8 is way better in coding, and actually it's the only one who could catch very very sophisticated bug in a large codebase (I tried different models including GPT-5.5 and DeepSeek). Interestingly Gemma 4 under opencode running locally performs not bad at all, it's far yet from DeepSeek level, but it manages to understand tools quite well, and code quality is pretty good. So, for simple coding projects I can say local models already won. It's amazing how smart open models of desktop size have become today. I mean it's quite plausible to manage small codebase today relying on only open tools and local models, you don't need any subscription to produce high quality code, but yes I assume you already experienced and know what you're doing :)
- noisy_boy 2mo agoI did the side by side between Claude code (effort medium) and Cursor (auto). Asked Claude to prepare a plan and asked Cursor to review it and it found tons of gaps in the plan. Cost-wise, it came out better too. I have been using Cursor daily (along with Claude) and the former has been 20-30% cheaper despite me spending more time on it.
- teki_one 2mo agoI had great results combining the two. If you (or your employer) can afford then you can ping-pong the models in the plan phase (not really ping-pong as humans should get a say too) and then let one implement and the other review. I got better results working this way than just to stick to a single model.
- setnone 2mo agolike others said in the thread: much less drama and i'll add much less attitude from the company and the models, overall i'm having much calmer experience with codex, hope it stays that way
- purpleidea 2mo agoThe codex software is garbage compared to Claude, but open source is the future, so you should at least switch.
- marknutter 2mo agoIt honestly baffles me how people can ask a question like this and get such a wide spectrum of answers in response. It's all so much based on vibes and anecdotal evidence. I've not really noticed much of a difference in capability since Opus 4.6 and I've used a ton of different models. They all work pretty damn well for me.
- corford 2mo agoPersonally I use Open Code with a copilot sub. Then all models are available in my session with just a /model and /variants command combo. Makes it super low friction to try different models & combos (my favourite right now is DeepSeek V4 Flash for initial PRD then Fable 5 high for implementation).
- siva7 2mo agoHave been long time clauder but honestly codex feels much more liberating. Something you can't buy..
- lizardking 2mo agoI consistently have better results with Codex for the work that I do. People have been saying that for six months, but until 5.4 the experience was sufficiently slower that it wasn't worth the switch. Making the switch was frictionless. Give it a try
- Skidaddle 2mo agoA few less obvious niceties of Codex: - built-in image generation using your subscription, which can be super handy - can actually edit Google Docs and Google Sheets (Claude can only create new or sometimes append) - I get a surprising amount of mileage out of the $20 plan They both have their places for sure.
- steve-atx-7600 2mo agoSet yourself up to be able to try / switch between models easily. I was a claude only user and just have my user level AGENTS.md for codex and others simply point at my user CLAUDE.md. Have a script that syncs my skills (just directories) between all models. Also, if you want to use /simplify or similar from claude in another model, you can ask claude for the prompt and put that in a skill for the other models.
- killix 2mo ago[flagged]
- SatvikBeri 2mo agoIt's trivial to try another agent. You can spend $20 for a monthly subscription and ask it to import all your settings from Claude Code.
- Kerbonut 2mo agoClaude Code is not the model, it's the harness. You can use any model you want with Claude Code to varying degrees of success. I use Qwen3.6-27b daily with Claude Code as an example.
- small_model 2mo agoNow we have various Opus+ level models (Opus/Fable, Grok 4.5, GPT 5.6) I prefer to focus on price/speed and harness as models are all generally good enough for coding. (Fable is overkill for 90% of work but is still level above). So I use Grok Build with 4.5 as its VERY fast and cheap, Codex is next best for me with sol/lunar 5.6. and Claude Code Fable for the 10% of tasks that need that level of reasoning. However I find Claude Code harness responsiveness much less than other two (all TUI versions) I wish they would fix this.
- lavela 2mo agoAm I missing something or isn't sol/lunar 5.6 only out for like 3 hours? How did you evaluate?
- small_model 2mo agohow do you mean, I always use the latest models so evaluating all the time.
- forshaper 2mo agoDon't know about consensus, but I personally still find Opus to be better for sniffing codebase intent and checking things as a whole, while Codex seems more detail-oriented for individual files.
- bmurphy1976 2mo agoI use both. Both are great. But in terms of Desktop Apps I think Codex has the better UI. It's more straightforward, just works, and has small conveniences like the open in editor icon. Claude's very bloated and convoluted by comparison. Maybe you need the bloat (Claude Design), but I prefer the more razor's edge efficiency of Codex. Model wise, I can't really tell. They all do what I want them to do most of the time and go off the rails occasionally. The question is increasingly becoming who's faster and cheaper and gives me more tokens, not who's better.
- gregwebs 2mo agoI run my AI agent as a different user (in addition to using the sandbox functionality provided by cc/codex). It does not seem possible to run the Codex GUI as a different user. I can run the TUI (/Applications/Codex.app/Contents/Resources/codex) but it has the shortcoming that remote control is only available in the GUI. I installed the Claude Code Codex skill provided by Anthropic and I am having Claude invoke it automatically to review all plans and changes. The nice thing about this is that for an additional $20/month pro plan I can extend the runway for Claude rate limiting and compare frontier model responses. I am looking for more ways now to work in Codex as a subagent that gets used automatically from Claude Code.
- Lucasoato 2mo agoI had to switch to Opencode from Claude code because the latter wasn’t supporting GitHub Copilot as model provider. I didn’t think I could have found a better solution, spawning multiple subagents with different models is such a great thing. I built in the past very small cli wrappers to call other models; Claude Code often refuses to do that, lies and does the job itself instead of delegating to another provider’s llms.
- ra0x3 2mo agoMy final answer on this is that we just can't say anything affirmative because all of our projects/codebases are completely different. I've gone back and forth on the "codex vs claude" being better, and while I'm currently of the believe that Claude is superior, I understand that might be the case for _my_ particular set of projects and _my_ personal way of interacting with the model.
- OkWing99 2mo agoI used to have the CC $200 plan, and moved to Codex 6 months ago. I have the anthropic $20 plan + API billing for rare use. Use Codex daily. Not having to deal with Anthropics constantly changing policies, token-gating, and carrot-and-stick marketing helps me to focus on work, rather than dealing with their company problems.
- jeffybefffy519 2mo agoIMO Codex has been the same rollercoaster ride as Claude. GPT 5.3-codex was incredible for backend/system tasks, GPT5.5 is better all rounder but weaker in some spots. There has also been many weeks when Codex's models were dumb AF. Same rollercoaster ride as anthropic between Opus 4.5 to 4.8... IMO the two biggest problems not really being answered by both OpenAI and Anthropic are: 1. Why not make specific models good at specific tasks for Codex/Claude Code. Theres a handful of types of work here whereby small good quality models would do better than these generalised all purpose models whereby someone discovers Fable is bad at biology.... 2. Why cant they consistently run these models and keep them performing? Performance of the models seems to directly correlate with amount of compute available, but they dont talk about it...
- jatora 2mo agoCodex has been comparable for a while. 5.1-5.5 have competed closely with 4.5-4.8. Fable blew them all out, now Sol comparable to Fable again. Some slight tooling differences with skills and hooks but for the most part I think if people are so engineered into one CLI that swapping to another inhibits them, then that is an error in usage habits. Codex historically will follow tasks more closely with less creativity, whereas Opus will do more than you specify. I wouldnt consider either one better due to this fact, just makes them useful for different situations. Generally they'll perform similarly for most tasks. Opus and Fable dominate 5.5 in artistic design (pixel art, ascii art), and edge out 5.5 slightly in general UI design taste. Have not tested Sol in that regard yet. So far in my usage Sol has been superior to Fable at graphics rendering engine optimization. Codex will work longer, and in single sessions without as much subagent usage. Codex only has 256k context but its compaction is absolutely next level. You will not notice compactions and they will happen multiple times during a complex task or set of tasks without you ever having to notice or care. Claude code on the other hand still has fairly poor compaction. Codex has more generous usage limits, and they also give you usage resets (weekly+5h resets) that you can bank for a month or so. Not sure how often they give these out. Codex also seemingly never has outages or weird delays like Claude code does. OpenAI randomly resets usage just like Anthropic does I would use both if you code often
- snthpy 2mo agoOT but how are y'all sharing your skills and agents across harnesses? I have a bunch of Claude Code Plugins and yesterday asked Codex to make them accessible to itself. It wanted to rewrite most of it. I was hoping i could get by with some symlinks or something to avoid drift.
- alchemism 2mo agoIt’s almost as complex as compiling software for different runtime environments; it requires a knowledge artifact repository and publishing system, e.g. https://github.com/thinkingsage/context-bazaar https://github.com/thinkingsage/context-bazaar
- hagen8 2mo agoIn my opinion Opus is waaayy better in agentic orchestration. It feels like it can natively deal with multiple subagents whereas gpt needs to be taught extensively.
- davedx 2mo agoI switch between both as my daily drivers. I do almost all my regular coding tasks with Codex 5.5 on medium. Sometimes for niche edge cases, or when I run out of tokens on my Codex sub, I'll switch to Claude. Some recent examples where Claude was able to solve things Codex couldn't: - 3D gamedev layout: I asked Codex to render a solar system in a certain camera positioning, saying it needed to fit the planets of the system to the viewport. Codex just couldn't do it, even on high reasoning: Claude Opus did it first attempt. - Tricky Tiptap image drag-n-drop layout implementation: Codex failed this after numerous iterations. Claude Opus also struggled mightily to get it to work, but I think around 3 attempts it nailed it. Both of them ended up grepping the Tiptap code from node_modules - that's the kind of task it was. But these are really isolated examples. Across all my projects (I have many; mostly TypeScript, but also things like C#), Codex "Just Works" (tm), with minimal prompting effort from me.
- algoth1 2mo agoA great thing about codex is that, even if run out of usage, it finishes the task. Claude code will abruptly break the work and leave it there unfinished as soon as it runs out of tokens. Also, antrophic randomly resets the token usage which is annoying when I’m trying to ration them. While openai gives you extra resets that you can apply when you want to
- fny 2mo agoI have them talk to each other via tmux to great effect on complex changes. Its great for auditing changes as work is done.
- sdoering 2mo agoTo me, the question of switching (and any recommendation) depends highly on the type of work you do as well as your setup in terms of harness and memory, context management, and so on. I have my context managed in a structured (but nowadays way too big) Obsidian Vault. I also built myself a vector based "vault search" capability and have my harness use this as a tool to find thematically similar things across the different contexts, when needed. I also build a few custom skills and extensions for my harness to be able to do my work. Talking about harness: I use pi.dev and have taken care of, that i set it up in a way as to easily be able to switch the intelligence layer without loosing context. Yes, there are differences in how well models perform, but if a model refuses a task - like gpt-5.5 not willing to build a downloading tool for Annas Archive - I switch the model to something less finicky. Thus I was able to switch to gpt based models after about a year with Claude (and having had a Claude Max since the early days it was available). I played a lot with other models recently, to see how stabl my setup is for switching, should something like Fable happen on a broader scale with the US government. As said, minor changes in tonality, minor issues ith the quality of long text being written by the model, but most of it is actually managed in by the tonality docs, guard rails, coding standards and the likes, I set up over the last 9+ months of intensive work with it (first in Claude Code, then Codex and now as said pi.dev). So YMMV and it heavily depends on your setup. But I more and more treat those models as interchangable.
- deleted 2mo ago[deleted]