7 ms·
SWE-1.7 Reach Near GPT 5.5 and Opus Intelligence
- harmonic18374 2mo agoA company whose first demo was completely fraudulent announces that its model beats GPT-5.5, on its own benchmark? I’m gonna wait a little before I trust this. This whole company seems to optimize for raising money and impressing VCs. Lying about their products, ignoring consumer market to target enterprise, bragging about how they work their employees like slaves, and writing these posts full of intimidating technical jargon...
- giancarlostoro 2mo agoWould love to see these companies use benchmarks done by third parties.
- anthonypasq 2mo agothey are right there? it shows swe-bench multilingual and terminal bench
- w4yai 2mo agoWhat happened ?
- SubiculumCode 2mo agoLink for this?
- jeffnv 2mo agohttps://news.ycombinator.com/item?id=40008109 https://news.ycombinator.com/item?id=40008109
- andy99 2mo agoIs it just me or does all that* seem pretty tame by today’s standards? Not saying it’s right, but it barely raises eyebrows. Sounds like a pretty typical startup demo. * Based on the first comment in the link that claims to summarize the video.
- oa335 2mo ago> "A company whose first demo was completely fraudulent" Could you expand on this?
- deleted 2mo ago[deleted]
- achandra03 2mo agoTo be fair it does seem like most AI startups are now like this (particularly when it comes to constantly mentioning how hard they work and ignoring consumer markets).
- parineum 2mo ago> it does seem like most AI startups are now like this Remember when AGI was going to replace all jobs in 6 months? It's always been like that.
- sigbottle 2mo agoI highly respect many people at cognition but yeah that's put a sour taste in my mouth. I want to work in the AI space on actual AI research, at any part of the stack. Even if I'm developing training infra - as long as people are advancing knowledge of what intelligence could be. But it seems like either it's big labs or grifters, that's it, and even the big labs, at least publicly, seem very grifty at times. Not like I have the technical chops probably, but still.
- alansaber 2mo agoThis is inevitable when the primary incentive is to raise aggressively. Overall I dont find cognition blogs that jargony, there are definitely worse offenders
- throwaw12 2mo agoOpen source for the win! Imagine how far community might have pushed if 2 past versions of 'morally superior' Anthropic and 'completely Open AI' open sourced their models for the community to build on top of them
- spott 2mo agoIs this open source? I can't find a link to download the weights.
- UncleOxidant 2mo agoIt's based on an open weight model (Kimi 2.7) so shouldn't it also be open weight?
- andy99 2mo agoThere is no obligation to do that. I think the landscape would be very different now if one of the big labs had released an earlier “frontier” model under copyleft that requires sharing fine tunes. I hope it still happens.
- cmrdporcupine 2mo agoDario is convinced that will create SkyNet, and so no, it will never happen. Only the blessed members of the True Church Of Effective Altruism can approach the Ark of the Covenant. The unwashed cannot be trusted.
- api 2mo agoRationalism and EA is Scientology for the Bay Area.
- cmrdporcupine 2mo agoIt's just the 2020s version of Ayn Rand's "Objectivism." Distillation of exploitative personality traits covered over in sophistry and philosophical excuses. Lets people be dickheads and be smugly superior about it at the same time.
- llmslave 2mo agoThese models are never as good, the benchmarks dont tell the full story
- SubiculumCode 2mo agoFunny, the cheerleading at HN for leading Chinese models, but a non Chinese lab (building on top of a Chinese model) gets dissed here.
- llmslave 2mo agoall the open source models are a waste of time relative to the bleeding edge from openai/anthropic
- pixel_popping 2mo agoNot true since a few months, genuinely try GLM 5.2 and Minimax M3, especially in adversarial/gating... as a general model, I can agree, but as a coding model, they are not bad, comparable to maybe Opus 4.5 in real usage which is quite impressive.
- llmslave 2mo agoyeah but why waste your time on these models, just use the one that gets the better results
- somenameforme 2mo agoI was going to respond until I saw your account name lol.
- llmslave 2mo agohaha i outsource my thinking to the smartest model
- 2mo ago
- achierius 2mo agoI've always had mixed feelings about Cognition. Obviously they have some very, very smart people working there (I even know a few), and they do make real products. But at the same time, they've made suspicious marketing claims more than once and even been caught making outright fabricated ones; and while they certainly seem to have shaped up from that, I still find their claims to be in a sort of grey area where they seem to avoid unfavorable comparisons and lean on their own benchmarks. Certainly when I've tried their models they have not been nearly as useful as comparable versions of Claude, GLM, etc. -- though I haven't had a chance to try SWE-1.7 yet.
- nibbleyou 2mo agoUnrelated: what's the point of "*equal contribution"? Why would someone specify this
- edot 2mo agoBecause papers are often referred to by the first author’s name, and often the first author is the primary researcher and therefore deserves the extra credit. When two or more primary authors are equally involved, they’ll often do a random ordering but annotate this so that no one thinks one did more than the others.
- nibbleyou 2mo agoInteresting. Thank you
- tancop 2mo agosome journals and colleges actually have a policy to always use random order to help fight the "senior researcher gets all the credit" culture in academia. theres a lot of cases where a prof forced their students to put them first even if they had an advisor role, or even credit someone for zero real work because they threatened to block submission and prevent the students from getting their degree.
- yousif_123123 2mo agoWe need more models that optimize for coding and that can be cheaper than frontier models, like what SWE 1.7 and composer 2.5 are trying to do. I don't think there's an effort to make something GLM-5.2 level but focused only on coding.
- UncleOxidant 2mo agoQwen was doing something like this with their coder models. But alas, they seem not to be releasing those anymore. Last one was Qwen3-coder-next.
- yousif_123123 2mo agoIts crazy that OpenAI and Anthropic themselves aren't doing that. No attempts at reducing inference cost for code as far as I know from them.
- andai 2mo agoOpenAI do have codex models, which are half the price. I haven't used them enough to comment on the quality though. I remember them saying a few years ago that, they didn't think it was worth specializing models for code, because their general purpose models kept beating them. I guess they changed their mind? Since they did start making codex models again.
- arcanemachiner 2mo agoThey also stopped making them again.
- andriy_koval 2mo agoMy speculation is that frontier models are MoE, and they just have some number of experts for coding.
- gsibble 2mo agoI use this model. It's pretty good but not Opus 4.8 or Fable levels obviously. I'm really hoping we get more models like it (and better) soon. I run it locally and it's great that way.
- hedgehog 2mo agoHeads up to anyone else curious, I installed the Devin CLI and SWE-1.7 is not currently available there.
- taf2 2mo agoNot finding anything about this while searching huggingface: https://huggingface.co/search/full-text?q=SWE-1.7 https://huggingface.co/search/full-text?q=SWE-1.7 i assume this is another closed source model?
- mirekrusin 2mo agoOpen weight models should have GPL-like license where it says if you train model on it, it needs to be open weight as well.
- joecot 2mo agoYes, and not only that but you can't even access it via API, you can only use it in Devin (formerly Windsurf). I'm an OpenCode user, but I'll fall back to Claude Code if I want to use Opus end to end for something, given my company has a subscription. But I'm not using yet another tool and subscription for a model that isn't even winning.
- pants2 2mo agoKinda funny that their "cost-vs-performance" chart looks the same as the one for Composer 2.5[1], except that it includes Composer 2.5 at a completely different spot. What are the chances that CursorBench ranks Cursor's model highest, and Cognition's bench ranks Cognition's model highest? Both are to be RL'd from Kimi as a base model, BTW. I'd posit that it's not deliberate deception, but for both companies their training data and benchmarks come from the same dataset (Devin/Cursor interaction logs) so they naturally overfit. 1. https://cursor.com/blog/composer-2-5 https://cursor.com/blog/composer-2-5
- culi 2mo agoI think it's also telling that they left out the usual hallmarks of the Pareto distribution: GLM 5.2, Qwen 3.7, Minimax M3, and Mimo 2.5 https://arena.ai/leaderboard/code/webdev/pareto https://arena.ai/leaderboard/code/webdev/pareto
- petesergeant 2mo ago> they left out ... GLM 5.2 They did not.
- bluelightning2k 2mo agoGood observation. I actually started typing the same point that the chances are actually high because of train/eval overlap then realised you answered your own question with that same observation. It is interesting though! Perhaps in some way this means we should decide which eval set aligns best with our taste? Back to the blog post. This is an excellent write up of an excellent technical achievement. I have a lot of respect for the Cognition/Devin (always "Windsurf" to me) and Cursor teams. I found it interesting - but justified - that they referred to themselves as a foundation lab rather than a dev tools company.
- swyx 2mo agoagent lab, not foundation lab
- 2mo ago
- fallinditch 2mo agoI'm looking forward to trying this out. I've been using SWE 1.6 quite a lot for grunt work alongside Opus for higher level planning and tricky stuff - a good combo. As a (former) Windsurf user I'm pretty happy with the progress of the Cognition/Devin ecosystem after they took over Windsurf, now known as Devin Desktop.
- ryandvm 2mo agoOkay, let's give software engineers a break for a bit and focus on obsoleting other high-linguistic context occupations.
- silvertaza 2mo agoLike diplomacy? heh
- cromka 2mo agoSome of these models could be particularly small, depending on the market they'd target...
- FergusArgyll 2mo agoI know you didn't mean this but have you ever seen Meta's model which plays (the game) Diplomacy? really cool https://ai.meta.com/research/cicero/ https://ai.meta.com/research/cicero/
- znsnjanwnwwn 2mo agoAnd to do that you’ll need development so until we’re all out of a job they’ll keep pushing. Once automating is automated it’s done.
- blauditore 2mo agoAny day now... Just a bit more space in the context window, trust me bro.
- Sabinus 2mo agoWe had decades of "just more and faster bits" in computing and that produced a lot of new capabilities.
- yieldcrv 2mo agosoftware and agentic workflows will obsolete those things RL environments building on top of each other will get these models there needs people doing software development lifecycles to figure it out and implement
- reflectix 2mo ago[dead]
- spate141 2mo agoFeels like they discovers that if you build your own benchmark, you can win it
- londons_explore 2mo agoPretty sure most benchmarks are being gamed by people training on the test set deliberately or accidentally anyway.
- ahk-dev 2mo agoThe benchmark debate is fair, but I think the more interesting signal is how quickly coding models are becoming a category of their own rather than just smaller frontier models. More specialization, more competition on cost, and probably a lot more benchmark gaming along the way :)
- petesergeant 2mo agoI think it's a bit odd to show the API prices for competitors when that's not how most people pay for them. I do like that it's provisioned by Cerebras though. I think I'd have leant towards focusing on the TPS.
- akshaydeshraj 2mo agoWould have been worth a consideration if it could have been used beyond it's own harness. Unfortunately, doesn't seem to be the case. https://x.com/theodormarcu/status/2074896486047834380 https://x.com/theodormarcu/status/2074896486047834380
- esafak 2mo agoUgh, that changes everything. If I wanted an arranged marriage I could go back to Claude Code.
- 2001zhaozhao 2mo agoHarness-wrapper tools that support multiple harnesses and allow sharing workspace features (skills, slash commands, etc.) between them will be meta.
- RussianCow 2mo agoIronically, Devin Desktop is one of those tools. It supports any harness that supports ACP (which is most of them)—you can use Claude Code, Codex, OpenCode, etc from the Devin Desktop UI. I'm currently experimenting with OpenSpec[0] as the "framework" and using different subscriptions for different parts of the spec-driven process: Opus via Claude Code for exploration, Devin SWE for building, and GLM 5.2 via the Z.ai Coding Plan for verification. I don't love having to mix and match harnesses, but in practice it's barely more effort than switching models. [0]: https://openspec.dev/ https://openspec.dev/
- thereitgoes456 2mo agoIs that surprising? It is standard "embrace, extend, extinguish" from a company not in a strong enough position to do the third one.
- mrinterweb 2mo agoI really don't want harness lock-in. I am trying to decouple myself from Claude Code now. I love the model of OpenRouter and being able to switch models at will let's your harness focus on your personal tooling and you can easily switch to the flavor of the month LLM with a single slash command instead of rewiring your entire workflow to use a harness to use a model. I like Cerabras, but I really wish they would make more of their hosted models generally available.
- Mitchem 2mo agoWhile I am skeptical of the results here, I am very excited for this new trend of making models faster. Running capable models at 1k TPS is more valuable for me than running better models at 30 TPS. I can only imagine the trend continues to move from "let's only make models smarter" to just incremental intelligence gains but with step improvements in speed.
- lnenad 2mo agoWhy? I'm personally on the opposite end. Less babysitting/higher quality means more time goes back to me/the user. 1000tps of bad code means you have to keep validating the output and circling back.
- anthonypasq 2mo agoid rather iterate multiple times than wait 15 minutes to notice it made a mistake.
- lnenad 2mo agoAgain, my point is exactly the opposite. Higher quality implies a mistake isn't made in a significant % of cases.
- unshavedyak 2mo agoIt's a lossy conversion though. "Mistake" is relative to the stated goals and specifications which are often heavily lacking. So unless you write with a high degree of architectural and implementation specificity then it might make very high quality code that is still not what you wanted.
- lnenad 2mo agoYou can ask for a complete feature/app/business. Or you can split up the work into verifiable/testable pieces and rely on a high quality AI to deliver. As time goes by the pieces will get larger as capability grows. I still trust myself and my experience when arch is involved, but AI has been great at tackling lower level stuff. And with Fable I don't really care it takes a while for it to complete, as I know I can trust it a lot more (which is what I personally prefer). Yes, with a 10k tps model you can iterate quickly. But that's not me personally.
- anentropic 2mo agohttps://devin.ai/pricing https://devin.ai/pricing Apparently 'free' on the $20/mo Devin plan (presumably within some quota still) and that is "via Cerebras at 1000 TPS" according to the announcement I live on Opus 4.8 High and their benchmark scores SWE-1.7 slightly higher ... if at all realistic that sounds like a great deal ... too good to be true?
- yousif_123123 2mo agoThe 1000 TPS shows for me as "SWE 1.7 Lightning" and took 14% of my daily quota in one prompt on the $20 plan. But the normal speed one seems to be free or with very generous limits.
- RussianCow 2mo agoThe "Lightning" (Cerebras) variant isn't free, only the regular one, which runs closer to 50 TPS in my experience with SWE 1.6.
- anentropic 2mo agoOh, the free plan said "Slow" so I thought maybe the others had the fast version :)
- settled 2mo agoI used SWE-1.5 and 1.6 when it was Windsurf (before Devin Desktop), it's not that bad (grunt work, tests, can actually plan and implement some medium level stuff) but you get a much much better value and better models (GPT-5.4^) going with a Codex subscription (plus you get resets). That company truly subsidized its user base to the extreme before, the $15/mo subscription was the best value on Earth paired with weekly deals reducing credits for premium models. Now it's barely any messages for paid models, completely watered down.
- chris_st 2mo agoFWIW, Cognition has all the Sonnet/Opus/Fable models, and all the GPT ones, as well as GLM, Kimi, and Gemini.
- 2001zhaozhao 2mo agoWait Devin has a CLI? Time to support it in my agent IDE just like Cursor's...
- hudo 2mo agoHow do i use it from opencode/openrouter?!
- RussianCow 2mo agoYou cannot; you must use either their Devin Desktop app or the Devin CLI.
- kgeist 2mo agoOn artificialanalysis.ai, Kimi 2.7 Code is way worse than GLM 5.2 at everything (general intelligence, coding, agentic tasks). But here, both Kimi 2.7 and its derivative SWE-1.7 are ahead of GLM 5.2. This tells me the benchmarks they use are cherry-picked.
- petesergeant 2mo ago> This tells me the benchmarks they use are cherry-picked. Which benchmarks would you have chosen instead, and why?
- cogman10 2mo agoIt honestly seems like there's not a great way to currently benchmark AI. The ideal way to run these benchmarks would be to give a 3rd party the model to run in an isolated environment so the prompts don't make their way back to the AI engineers. That seems doable for open weight models, but not for private models.
- p1necone 2mo agoIf you've got money to burn on tokens, the way that seems best to me is to set up a repeatable harness - docker container with a specific past commit from your own project, set of known issues/features that you've already fixed/completed of varying levels of complexity. Set up a script that launches the harness for each model, prompts them to implement one of the tasks, let it churn until either tests pass or it hits some budget limit. Then, most importantly, read the transcript and output and judge subjectively - I don't think this actually can be narrowed down to a score, although tokens burned to fix, whether it actually got the tests green etc are all good signals. (I've done this, but so far only on a codebase that was too complicated with models that were too weak because I didn't want to spend more than a few dollars - results were inconclusive, planning on iterating on my personal benchmark in future)
- bearjaws 2mo agoLikely benchmaxxed. You see it in Qwen and other smaller models all the time.
- ai_slop_hater 2mo ago[dead]
- godzillabrennus 2mo agoCognition... oh what a ride... We were customers when they acquired Windsurf, stopped offering customer support, raised prices, dismantled the brand, and raised prices again. We are not customers anymore. Benchmarks are not the only thing to worry about when you are using models.
- Bnjoroge 2mo agoI’ve unfortunately had to temper my excitement with Cognition’s models/products given the amount of unwarranted hype they created with Devin on first release, but hopefully this is good.
- deleted 2mo ago[deleted]
- amarant 2mo agoVery thankful someone is doing this work! I suspect it will be thankless work for a while: we're not far enough into diminishing gains territory for anything other than the absolute best being worth considering for most people, but I reckon we will be pretty soon. If a year from now swe 2.0 or whatever reaches fable 5 parity for a fraction of the cost, that'll be very attractive indeed!
- messia 2mo agoAnd yet when I use swe it feels like massive shit
- deleted 2mo ago[deleted]
- messia 2mo agoif you are still using these products in 2026 you are really a shit engineer
- modeless 2mo agoWhat is the actual per token price? The benchmarks look similar to Grok 4.5 also released today and priced at $2/M input tokens and $6/M output tokens.
- tonyhart7 2mo agoand what is actual intelligence per dollar benchmark ???? its useless comparing token/dollar while some model inherently generate more thinking output and cost more despite lower cost
- RussianCow 2mo agoThe regular one (not the fast variant) is free but slow. The "Lightning" variant (which uses Cerebras and gets supposedly 1000 TPS) costs $12.50/M output, $2.5/M input, $1/M cached input. So it's quite a bit more expensive than SWE 1.6.
- deleted 2mo ago[deleted]
- bobtheborg 2mo agoI like to use SWE-1.6 for quick help with git. For instance: review the top stash and tell me what's in it (grouped appropriately) 1.6 does this fine nearly instantly. 1.7 tried for 17s before I killed it
- luciana1u 2mo ago[flagged]
- retinaros 2mo agoit is obvious at this stage that most of the gains are in distillation post training and having good RL simulations. the moat of the private labs are just their capability of stealing open research and data and locking the few bit they innovate on the top of it. It will work until they IPO.