7 ms·
I wonder when/if we will start to see Apple include a silicon encoded model into their chips. Similar to Taalas build Llama 3.1 silicon with 17,000 tok/s infere
by Humphrey 22d ago
I wonder when/if we will start to see Apple include a silicon encoded model into their chips. Similar to Taalas build Llama 3.1 silicon with 17,000 tok/s inference.
So could the M7 actually include an AFM 3B model, alongside a generic neural engine?
- wmf 22d agoThat model would be larger and more expensive than the entire M7 chip.
- tyre 22d agoWhy would a local model for a consumer device need 17k tok/s? Apple is better off building chips with generalizable TPUs (or equivalent) so they can upgrade/patch models.
- MaxikCZ 22d agoI cant shake the feeling of "640KB is enough for everybody". Imagine not one AI answering over 1 minute but a team of 100+ agents in hieararchical structure taking care of your request in seconds, checking each other.
- jeffybefffy519 22d agoIt would open heaps of use cases, you could almost pass it over frames of images the camera sees in real time for example...
- deleted 22d ago[deleted]
- dyauspitr 22d agoIt would make zero sense. These things are improving by leaps and bounds every week, we are not at the point where you can burn weights into silicon and put it on one of the largest consumer devices on the planet yet.
- aqfamnzc 22d agoOn the other hand, models these days are getting to the point where even if all development halted permanently, they would continue to be useful long into the future. (At least until their knowledge base or linguistics become too outdated.)
- deleted 21d ago[deleted]