4 ms·
Instrumentation checklist for running large GPU clusters
- deleted 2y ago[deleted]
- amaldavid 2y agoJust stumbled upon this blog which details out testing and validating large GPU clusters before running training workloads. Any other similar blogs which adds more nuance in terms of debugging the issues once we identify them as well?
- roanakb 2y agothis is a good one for debugging rdma: https://docs.redhat.com/en/documentation/red_hat_enterprise_linux/7/html/networking_guide/sec-testing_early_infiniband_rdma_operation https://docs.redhat.com/en/documentation/red_hat_enterprise_...