6 ms·
Apple's Neural Engine is equivalent to the "NPU" hardware you would find on many cheap ARM SBCs. It's a dedicated coprocessor with a very low performance target
by bigyabai 4d ago
Apple's Neural Engine is equivalent to the "NPU" hardware you would find on many cheap ARM SBCs. It's a dedicated coprocessor with a very low performance target, not intended for giant transformers or LLM acceleration. These NPUs are not particularly hard to design, and their critical flaw is that they don't scale very well. CUDA "won" because a bigger GPU meant having an utterly massive amount of CUDA cores to delegate ALU work to. NPUs/Neural Engine has the opposite problem, where spending $10,000 on an M1 Ultra only gets you ~2x better NPU performance versus the baseline $600 M1 chip. NPUs and Neural Engines are essentially dark silicon on the majority of devices with them, their uses are few and far between.
Considering Apple's refusal to sign Nvidia's ARM/CUDA drivers, it is pretty clear how Apple missed the boat here. Apple Silicon could have dominated the datacenter rollout if macOS supported CUDA properly. The Mac Pro would probably not have been cancelled if the PCI lanes could be used for normal datacenter GPUs and CUDA workloads. The excellent Thunderbolt bandwidth present on so many Macs is wasted supporting RDNA but not eGPU enclosures. There are several hardware features that Apple holds back for no good reason, handing Nvidia the lead in certain markets.
The only thing that stopped Apple from riding AI to the top was their own petty grudge towards Nvidia. Simple changes to macOS would have destroyed Nvidia's Grace CPU sales and made Apple Silicon the crown prince of the AI boom, at zero risk to themselves. The only thing Apple really needed was their own Mellanox equivalent.