6 ms·
When it comes to quality of outcome, since at least Feburary, the harness has almost equal, if not more weight than the model itself. It's no longer "which mode
by hadlock 1mo ago
When it comes to quality of outcome, since at least Feburary, the harness has almost equal, if not more weight than the model itself. It's no longer "which model is the best?" it's "which model + harness is the best?"
I get drastically different tool call failure rates using Claude SDK vs OpenCode using Qwen 3.6 models
- azinman2 1mo agoWhich works better for you?
- KronisLV 1mo ago> the harness has almost equal, if not more weight than the model itself This feels like a horrible failing of the models to generalize, then - both basic and intermediate tasks should be possible to do with Claude Code, OpenCode, Pi, ZCode, Kimi Code, Dirac and tbh any other mainstream or even slightly niche harness. Not doubting the claim itself, there's a reason why good benchmarks include the harness.
- dominotw 1mo agoi think thats BS that harness has equal weight. most of intellegice is still coming from training data not from RL. so how is 'coevolved harness' equal weight.
- disgruntledphd2 1mo ago> most of intellegice is still coming from training data not from RL. For coding specifically, I'm not sure this is still true. Given the heavy use of RL to improve coding performance, I'd expect the harness to be important as it defines what tools the model is rewarded for using.
- dominotw 1mo agoi belive they looked at the traces and they were still legible ( RL traces should be gibberish)
- davidlt 1mo agoI just wanted to emphasize this. Harness is a big part of how things perform thus usually it's harness + model co-design that's important.
- JLO64 1mo agoIt's worth nothing that recent Claude models seem to have gotten worse at tool calling outside of Claude Code and the SDK: https://lucumr.pocoo.org/2026/7/4/better-models-worse-tools/ https://lucumr.pocoo.org/2026/7/4/better-models-worse-tools/
- HDBaseT 1mo agoYeah this is a complete lie. You can use effectively any harness and get good results. Harnesses are mostly placebo.
- scrlk 1mo agoWith Opus 4.7, there's a 10 point improvement in the AA coding agent index when you swap out Claude Code for OpenCode: https://artificialanalysis.ai/agents/coding-agents#harness-comparison https://artificialanalysis.ai/agents/coding-agents#harness-c... To put that in perspective, the difference between GPT-5.6 Sol Max and 5.6 Luna Max is 8 points. That's a lot of extra performance that you can get for free just by using the best harness.
- RideOnTime22 1mo agoEvery other week it's a new "X didn't matter, until Y date" without any hard quantitative claims. It's crazy how over the past years a field originating from math ends up succumbing to subjective feels.
- hadlock 1mo agoIt's a forum, not an engineering conference, but here is a demonstrated 10 point difference between two top harnesses anyways: https://artificialanalysis.ai/agents/coding-agents#harness-comparison https://artificialanalysis.ai/agents/coding-agents#harness-c...