5 ms·
Anyone have any idea what the architecture/vendors they are using for inference/compute? Getting the compute to run inference for multi-trillion parameter mode
by choilive 2mo ago
Anyone have any idea what the architecture/vendors they are using for inference/compute?
Getting the compute to run inference for multi-trillion parameter models at any sort of scale and performance is daunting. There are a handful of vendors that have systems that can do this (~ Nvidia NVl-72 class) that pretty much only the frontier labs and hyperscalers effectively have access to.