9 ms·
Looking at intelligence vs cost: - Opus 5 is 10% smarter than Grok 4.5 for 10x the cost. - Opus 5 is a bit smarter than Gpt 5.6 Sol for 2.75x the cost ref: ht
by guybedo 2mo ago
Looking at intelligence vs cost:
- Opus 5 is 10% smarter than Grok 4.5 for 10x the cost.
- Opus 5 is a bit smarter than Gpt 5.6 Sol for 2.75x the cost
ref: https://artificialanalysis.ai/?cost=intelligence-vs-cost-per-task https://artificialanalysis.ai/?cost=intelligence-vs-cost-per...
- I_am_tiberius 2mo agoWith Grok you can be sure that you're data ends up in the next model (derived or anonymized, but still).
- adamtaylor_13 2mo agoYou can opt out of training. If you don't believe checking the opt-out box actually opts you out, then this sentence could be said about literally any provider.
- JacobAsmuth 2mo agoAll providers are equally trustworthy :)
- I_am_tiberius 2mo agoMaybe I am biased, but the person in control of that specific company is by far the most untrustworthy.
- jofzar 2mo agoYes but grok has the literal track record of our of the box uploading your whole codebase, secrets included to a remote box.
- adamtaylor_13 2mo agoIf I recall correctly that was a bug, not malicious intent. Hanlon's razor makes me presume that's likely.
- conradkay 2mo agoI don't think can use the AA index to say something is 10% smarter I assume 100 is the max, meaning it's impossible to be 2x as smart as Muse Spark 1.1
- alphabettsy 2mo agoAs always, it requires evaluation with your work because I’m often finding grok to be much more expensive than the price would lead you to believe. There’s also the frustration of it not quite being enough sometimes. It’s extremely capable, but I still find that it needs more concrete guidance and boundaries than other models.
- adverbly 2mo agoThe "current top dog" smartest model available will probably always have a premium to go after use cases where a little more intelligence is worth a lot more value. It did far better at some tasks compared to Sol (e.g. the ARC 3 benchmark). And at those tasks, it's not just "a bit smarter": It got 30% vs less than 8% - so you're talking 2.75x more for almost 4x the coverage.
- dist-epoch 2mo agoThat's not how intelligence works - "IQ 130 is just 7% smarter than IQ 120"
- JacobAsmuth 2mo agoAA isn't the best way to measure relative cost in real world use because some of those benchmark questions are extremely hard for the models. Some models give up quickly on hard questions, other models spin their wheels for a long time before declaring defeat (or getting the answer on token 200k!). A useful measure of real world cost (complementary with total cost like they already report, of course) would be "cost for correct answers". You could look at the ratio between the two costs to get a measure of laziness which many would find quite useful.