6 ms·
I keep seeing these Grok 4 intelligence claims, so I tried something very simple: "Animate a round robin tournament for 10 people." Results: Claude: ~10s, perf
by Rperry2174 1y ago
I keep seeing these Grok 4 intelligence claims, so I tried something very simple: "Animate a round robin tournament for 10 people."
Results:
Claude: ~10s, perfect working demo
ChatGPT: ~20s, solid solution
Grok 4: ~1000s, failed completely, gave me a truncated base64 blob
This wasn't some obscure edge case... it was basic data visualization that any decent model should handle. Yet somehow Grok 4 is "competing with humans" and has "99% tool accuracy"...
I don't buy it..
links:
Claude: https://claude.ai/share/7a413a6a-5c01-44a1-aaed-8b237e5e9e94 https://claude.ai/share/7a413a6a-5c01-44a1-aaed-8b237e5e9e94
Chatgpt: https://chatgpt.com/canvas/shared/687a9f9d4304819187ac7d98d30f8aad https://chatgpt.com/canvas/shared/687a9f9d4304819187ac7d98d3...
Grok 4: https://grok.com/share/c2hhcmQtMw%3D%3D_20b61291-e1bb-45e5-ade4-f6d916209b6e https://grok.com/share/c2hhcmQtMw%3D%3D_20b61291-e1bb-45e5-a...
These benchmarks are either just wrong or measuring something completely divorced from practical utility imo...