13 ms·
Frontier models have seen more Mathematica and Powershell code than I ever have in their training, yet they really struggle to produce working output. They seem
by Fredkin 6d ago
Frontier models have seen more Mathematica and Powershell code than I ever have in their training, yet they really struggle to produce working output. They seem heavily tuned to Linux too. Despite adding skills to rectify this, they still fail to realize they're running on Windows and waste tokens. A human with this much training wouldn't have this problem. There are evidently still some pretty big holes still.
- jsenn 6d agoThis is probably a harness problem rather than a model problem. GitHub Copilot will happily and effectively use Powershell while Claude Code struggles in my experience.