6 ms·
Your maths is incorrect, dividing by the number of devices doesn't make sense in this case, you should be multiplying them. E.g. - 466 minutes of fourth-gener
by codefined 5y ago
Your maths is incorrect, dividing by the number of devices doesn't make sense in this case, you should be multiplying them. E.g.
- 466 minutes of fourth-generation TPU time.
- 1,600 minutes of third-generation TPU time.
Using this logic, fourth generation TPUs are 3.4x better. But, comparing different cluster sizes is pointless. These things don't scale linearly.
- skyde 5y agoyou are most likely correct! I read ´Training took 1.82 minutes with 256 fourth-gen TPUs’ to mean Using all 256 TPU it took 1.82 wall clock time to do the training not 466 minutes. This mean each TPU contributed to (1/256) assuming it scale linearly. Even if it doesn’t the 256 TPU ran in parallel not sequentially right ?