trunk/a476c633b36df133775c886b3c10b12ad0087d9e: [ROCm] Enable native AsyncTP (#177961)
- PyTorch 近 90 天出现 1941 次
- PyTorch 近 90 天第 1915 次发版
- 上一次:同一天稍早 · ciflow/trunk/199870
发生了什么
Summary This PR enables the native AsyncTP path of fused_all_gather_matmul on ROCm (gfx942 / MI300X and gfx950 / MI355X). The native path overlaps the all-gather of the activation shard A with A @ B , using a persistent GEMM that waits on a per-chunk signal before reading each gathered chunk of A . On CUDA this GEMM is a CUTLASS kernel; this PR adds the ROCm equivalent using the Composable Kernel (CK) tile Persisten…
摘要按规则整理自下方来源原文