5 ms·
One way to gauge how close to viable local inference is the quality of the arguments against it are. $100/mo for ten years is $12K. $200/mo is $24K. When you l
by jmull 21d ago
One way to gauge how close to viable local inference is the quality of the arguments against it are.
$100/mo for ten years is $12K. $200/mo is $24K. When you look at the valuations of the companies providing these tokens, you don't necessarily see these rates going down.
Your argument against privacy is so obviously bad I don't think it needs a response.
Today, most companies and individuals (including me) are still going to be buying tokens from the cloud. But for many it makes sense to go local today, and it seems pretty clear the day where we turn the corner isn't that far away. E.g., suppose the RAMpocalyse eases in 3 years and Apple releases the M7 ultra, which has actually been designed for local inference (and other players are making their moves as well). I suspect it will become something of a no-brainer to do most inference locally, perhaps supplemented by lighter usage of per-token plans from the cloud for access to specialized models.