7 ms·
I can actually see this being useful for fairly narrow workloads in dedicated devices where the model doesn't need to change very often and low-latency inferenc
by Transformanshen 1mo ago
I can actually see this being useful for fairly narrow workloads in dedicated devices where the model doesn't need to change very often and low-latency inference matters more than flexibility
I don't see it replacing general-purpose GPUs but it seems like a reasonable option for that kind of workload