← 返回事件
持续讨论AI发版

viable/strict/1791014076: [MPS] Reduce strided inputs in place for full reductions (#198646)

发生了什么

Since #198494 dropped the Strided pass-1 kernel, a full reduction ( dim=None ) over a non-contiguous input materialises a contiguous copy and then runs Flat over it. The copy moves more bytes than the reduction reads, so e.g. x[:, ::2].sum() in fp32 became 2x slower than before the refactor, and max / min / all / any have paid that copy since they moved off MPSGraph. This adds a FlatStrided plan: at::collapse_dims r…

摘要按规则整理自下方来源原文

为什么在扩散

来源