DSN AI HA Hao Elastic Training: Dynamic Scaling for Fault-Tolerant and Cost-Efficient Large Model Training Abstract: Large-scale GPU clusters are inherently dynamic environments — hardware failures remove nodes, maintenance windows shrink…
DSN SCI AI HA Hao Intelligent Training Job Orchestration: How HyperPod Bridges the Gap Between Hardware Failures and… Abstract: When a GPU node fails during large-scale distributed training, detecting the failure is only the first step. The training job —…