← 返回事件
持续讨论AI发版

viable/strict/1791262552: [ROCm] Enable native AsyncTP (#177961)

发生了什么

Summary This PR enables the native AsyncTP path of fused_all_gather_matmul on ROCm (gfx942 / MI300X and gfx950 / MI355X). The native path overlaps the all-gather of the activation shard A with A @ B , using a persistent GEMM that waits on a per-chunk signal before reading each gathered chunk of A . On CUDA this GEMM is a CUTLASS kernel; this PR adds the ROCm equivalent using the Composable Kernel (CK) tile Persisten…

摘要按规则整理自下方来源原文

为什么在扩散

来源