viable/strict/1790933554: Add NCCL2 debug server health checks (#199092)
- PyTorch: 1794 events in the last 90 days
- PyTorch: 1772th Release in the last 90 days
- Previous: earlier the same day · trunk/77700655e1ddd21002e444ea9d9a501cfe259da9: [torchtitan hash update] update the pinned torchtitan hash (#199225)
What happened
NCCL2's flight recorder exposed a dump endpoint but did not publish watchdog health to the debug server, so the server could not trigger an all-rank dump when one rank detected a timeout or communicator failure. Port the TorchComms health-check model: have the debug frontend request NCCL2 flight-recorder dumps from every rank when any rank is unhealthy. Give a fatal abort the same configurable grace period and route…
Summary assembled by rule from the sources below