7 ms·
I recognize the sarcasm. The data I can find says it's performing at baseline however? https://marginlab.ai/trackers/claude-code/ https://marginlab.ai/trackers
by MattSayar 5mo ago
I recognize the sarcasm. The data I can find says it's performing at baseline however?
https://marginlab.ai/trackers/claude-code/ https://marginlab.ai/trackers/claude-code/
- ACCount37 5mo agoYeah, that's my point. Humans are not reliable LLM evaluators. "Secret model nerfs" happen in "vibes" far more often than they do in any reality.