7 ms·
That's the best answer I've seen to that question so far, my honest respect. I am much more skeptical than you, for me I have seen a sea change in _tool capabil
by igorkraw 1mo ago
That's the best answer I've seen to that question so far, my honest respect. I am much more skeptical than you, for me I have seen a sea change in _tool capabilities_ (mainly pre opus 4.5 to post 4.5) and harness engineering but no strong change in the type of errors made and the pattern of harness engineering (the pattern of "set things up for the LLM to see when it fucks up and let it flail till the verifier tells it to stop").
I would actually expect the sea changes as you describe it in your first criteria to continue with 1) vision, audio and video natively integrated 2) continued scaling of e2e rlvf for workflows with large scale labeling efforts 3) ASICs and widescale deployment of diffusion models leading to speed ups
But as of right now, I still expect these models to need humans to prune the output to the gold and set up the harness right for both the novel bits, and for the boilerplate to be cohesive with the global intent.
Which is of course an amazing potential boost in productivity, but still a sigmoid flattening.
As for your second criteria that includes cost, I think we might every well see this coming soon, but it's difficult to estimate with the efficiency gains still possible.
Thanks for engaging:-)