6 ms·
in my experience of 1 month daily use, Qwen 3.7 Pro is just unusable. wastes too much time, goes off track, useless stuck loops, cannot debug at all. Deepseek V
by 5701652400 2mo ago
in my experience of 1 month daily use, Qwen 3.7 Pro is just unusable. wastes too much time, goes off track, useless stuck loops, cannot debug at all. Deepseek V4 Pro is night-and-day compare to Qwen. actually Qwen models seems the worst SWE experience so far. and it is super expensive compare to Deepseek. cannot delegate anything to it, cannot use it real-time low-level tasks either. totally unusable.
- ph4rsikal 2mo ago> Qwen 3.7 Pro is just unusable. wastes too much time, goes off track, useless stuck loops, cannot debug at all. D Anthropic should not have bugged their knowledge distillation attacks.
- chewz 2mo ago> Anthropic should not have bugged their knowledge distillation attacks. It is like one of Pizzaro's men crying that someone have stolen his precious golden dublons As Lenin have said - "Loot the looters" (Russian: Грабь награбленное)
- RazorBucksICO 2mo agoAppealing to the Belsheviks for moral authority is, well I will just say an interesting approach. I do not have that much sympathy for Anthropic, but I do not have much sympathy for publishing companies either whose rights to a revenue stream they violated either. Are Chinese AI companies the Robin Hood in this story? Would they be so magnanimous if they had the upper hand? I don’t think so.
- trollbridge 2mo agoConsidering the results from Kimi K3, it appears most the accusations of them “stealing” via distillation are unfounded accusations.
- zobzu 2mo agohow many hn posts do you believe arent propaganda these days? its billions, trillions were talking about. imo hn should display posters origin, such as country, bon, datacenter registered ips, and the discourse will change dramatically.
- neonstatic 2mo ago[flagged]
- chewz 2mo agoFrom my experience Qwen-3.7-Max is above the Opus level but delivers results much faster. Slightly worse then Fable. Way ahead of Deepseek 4 Pro (in speed and overall comprehension) - which is a workhorse on its own. I am using them all with Claude Code mostly. Qwen-3.7-Plus is quite OK, good for subagent use. Way better then Sonnet. Qwen-3.8-Max-Preview seems working just fine for me at the moment - I am playing with is right now but too early to say anything. At 10% of regular price it is a steal so far.
- nullbio 2mo agoIf by Opus you mean Opus 4 and not Opus 4.8, then sure.
- chewz 2mo ago> If by Opus you mean Opus 4 and not Opus 4.8, then sure I meant Opus 4.8 which is rather dumb and ineffective in coding harness, especially with higher thinking levels.
- deleted 2mo ago[deleted]
- porksoda 2mo agoMy experience was so much different to this, that I have the unfortunate impression that you're shilling. It really was not a capable model, it felt like the old oai models back when we were all excited but couldn't actually trust them even in the littlest ways. What harness were you using, did you do any work to make it better? What was I doing wrong? I just pointed opencode at it, with a pretty simple (large-ish) data cleaning project.
- chewz 2mo agoHave you actually used Opus 4.8 in Claude Code? It takes way too long to do any practical task on higher thinking levels due to over-engineering. And I am not the only one complaining. Lots of people downgrade to Opus 4.6 exactly for this reason. Opus 4.8 training works well for agentic work. Not for code harness. EDIT: ``` stronger on coding and raw capability but can be more argumentative, verbose, and costly. Reliability and instruction-following Many users say 4.6 felt more reliable and followed instructions better. "With 4.6, when I tell it something, it actually remembers the spirit of what I asked for and keeps applying it." Others report 4.8 drifts from preferences and can be frustrating to control. "I still find myself getting frustrated when it ignores preferences and drifts from instructions" Some people find 4.7/4.8 push back more and act more adversarial than 4.6. "The biggest complaint against 4.8 is that it is argumentative and "pushes back" constantly" Coding quality and capability Several users praise 4.8’s coding strength and thoroughness. "4.8 is technically impressive, especially for coding" Other reports say 4.6 could be better for certain coding workflows and breaks less. "4.6 still >> 4.8 for anyone else as well? Maybe I'm in the minority, but for my use cases Opus 4.6 is still better than" Some recommend mixing models: use 4.8 for key tasks and 4.6 for general work to save tokens. "What I do is... use 4.8 for key moments, and for everything else 4.6" Cost, speed and token behavior Users note 4.8 often uses more tokens and can feel slower because it “thinks” more. "4.8 is much more cautious, and as a result - slower. It checks everything, thinks for a long time etc." ``` [https://www.reddit.com/answers/601770d4-4059-478d-aa52-b445c60f1cb7/?q=opus+4.8+vs+opus+4.6&source=SERP&upstreamCID=9de6493e-379f-4e87-8435-94de4b569d64&upstreamIID=20b927c0-701c-4c36-8a8b-6edbdb3f4a20&upstreamQ=opus+4.8+vs+opus+4.6&upstreamQID=8ef17d1d-e13f-4378-bb4c-d78a7f69a8ca&tl=en https://www.reddit.com/answers/601770d4-4059-478d-aa52-b445c...]
- big-chungus4 2mo agoQwen3.7 pro is meh, but 3.7 max is a very good model
- Demiurge 2mo agoAre these different models or different efforts for thinking (internal back and forth review) using the same model?
- 2Gkashmiri 2mo agoCan you tell me more about deepseek? I paid $2 for deepseek api, put the key in void editor and made a crypto tool in html. It turned out to be around 67kb. I used sample files in CSV that were a few hundred lines. It spent around $1.8 in the hour or two or light coding and follow up bugs. Is it really really this much? I can't imagine spending a month using it for a day job, it would cost more than the salary so what gives? I understand the local ai and all that but do cloud providers cost this much? Earlier I thought "billion tokens" but now not sure
- aduwah 2mo agoA local AI is not about cost. In fact you will likely pay more for it than with most providers. Just look up the advantages of having access to a technology like this that can be self hosted
- 5701652400 2mo agoso Deepseek 4 Pro cannot go on own sessions for too long. I delegate small-medium tasks: refactors, summaries, research, writing tests + have very good codebase already + extensive history / architecture / docs / linters. so it picks up and does decent small-medium scope work. it is fast, accurate, cheap. does exactly what I want directly and does not waste time nor tokens. definitely not "implement me complex greenfield project".
- k__ 2mo agoMy 2 weeks with DeepSeek V4: Pro is ~50% more expensive than Flash. Both need babysitting. Plan, split in small tasks, give it docs, types, tests, linter, best practice examples, etc. Always start a new session when starting a task. Do regular manual sanity checks, and tell it to find issues in the codebase. I pay like $1,50 per day for Pro.
- 5701652400 2mo agovery simlar experience. I would also add that I run it this way ~12hour a day non-stop. 300M / tokens per day (99.7% cache hit).
- 2mo ago
- 3abiton 2mo ago> in my experience of 1 month daily use, Qwen 3.7 Pro is just unusable. wastes too much time, goes off track, useless stuck loops, cannot debug at all. Deepseek V4 Pro is night-and-day compare to Qwen. actually Qwen models seems the worst SWE experience so far. I have used both Qwen3.6-35B and Qwen3.6-27B locally (both Q8 quantized with llama.cpp). I have also used antirez's quant of DS4-flash. They all performed within the same tier, DS4 being a bit more efficient, but they all gave really good results, mainly used for bash scripting, debugging, python and some C++. I am curious what type of applications/langauges failed with Qwen? One thing to note, the chat templates were "broken" for qwen models and had to debug it, there are already effort on this. Tbh, the same with gemma.
- Grimblewald 2mo agofor usable local layperson applications, I think qwen models are king. However, at hosted/frontier scales I'd agree for code. However for visual comprehension etc. for me qwen is the top of the line, and it isnt even close. Qwen VL models nail tasks frontier models dont even get close to acceptable on.