5 ms·
It's really interesting how much the AI harness seems to matter. Going from 48% via Google's official results to 65% is a huge jump. I feel like I'm constantly
by mdasen 5mo ago
It's really interesting how much the AI harness seems to matter. Going from 48% via Google's official results to 65% is a huge jump. I feel like I'm constantly seeing results that compare models and rarely seeing results that compare harnesses.
Is there a leaderboard out there comparing harness results using the same models?
- alfiedotwtf 5mo agoFor my local tests the past few months on the same local model, I’ve found Claude Code to be way better than OpenCode, and OpenCode to be better than Codex.
- GodelNumbering 5mo agoI really wish there was! I thought of even creating one but it would be conflict of interest
- manx 5mo agoWe probably want to compare the cartesian product of model+harness.
- culi 5mo agoMaybe the future isn't a human-like centralized intelligence but an octopus-like decentralized intelligence where more focus is placed on making the harness itself "smart"
- dominotw 5mo agoThat would be counter to AI company goals. They want harness to be dumb and models to be smart so they can sell models.
- SwellJoe 5mo agohttps://en.wikipedia.org/wiki/Bitter_lesson https://en.wikipedia.org/wiki/Bitter_lesson History indicates you can't tool and harness your way to effectively competing against a smarter model with more compute.
- satvikpendem 5mo agoNot really. Anthropic for example sells both the harness and the models as a unified kit via Claude Code, it is in their best interest to make sure both parts work as well as possible, via reinforcement learning of previous usage as well for new model performance increases.
- dominotw 5mo agobut harness are not a moat. They wouldnt have to subsidize their own harness massively if that was the case. Anyone can write a good harness .
- satvikpendem 5mo agoThat's not true that anyone can write a good harness because the LLM providers have information like prompts that they can RL train off of that someone writing their own harness would not have. Therefore a good and proprietary harness is a moat.
- dominotw 5mo agothat doesnt answer why claude subsidizes their own harness and bans ppl from using subsidized inference on openclaw ect
- satvikpendem 5mo agoYes it does? They want people to be locked into the Claude Code product.
- dominotw 5mo agowhy do they have "lock" them if its clearly superior to alternatives that merely u se their api.
- satvikpendem 5mo ago
- isege 5mo agoIsn't that what terminal-bench does?
- nikcub 5mo agothe most cited is terminal bench 2.0, but its also plagued by cheating accusations and benchmaxxing. somewhat remarkably, claude code ranks last for Opus 4.6 - which may say something about cc, or say something about the benchmark [0] https://www.tbench.ai/leaderboard/terminal-bench/2.0 https://www.tbench.ai/leaderboard/terminal-bench/2.0