viable/strict/1791014076: [MPS] Reduce strided inputs in place for full reductions (#198646)
- PyTorch 近 90 天出现 1848 次
- PyTorch 近 90 天第 1824 次发版
- 上一次:同一天稍早 · trunk/66f0744d7ce600ca807699a2a3c19f3cd1337886: Mark unused parameters in inductor and nativert (#199467)
发生了什么
Since #198494 dropped the Strided pass-1 kernel, a full reduction ( dim=None ) over a non-contiguous input materialises a contiguous copy and then runs Flat over it. The copy moves more bytes than the reduction reads, so e.g. x[:, ::2].sum() in fp32 became 2x slower than before the refactor, and max / min / all / any have paid that copy since they moved off MPSGraph. This adds a FlatStrided plan: at::collapse_dims r…
摘要按规则整理自下方来源原文