trunk/326909e0a95f0fbd10fc85dffe0bf18858e12998: [MPS] Enable gemv kernels for linear (#198721)
- PyTorch 近 90 天出现 1789 次
- PyTorch 近 90 天第 1767 次发版
- 上一次:同一天稍早 · trunk/77700655e1ddd21002e444ea9d9a501cfe259da9: [torchtitan hash update] update the pinned torchtitan hash (#199225)
发生了什么
We have gemv kernels which only dispatched with torch.mm before 🤦♂️. This enables it for F.linear , i.e. what models actually run when decoding. Summary of speedups: dtype shapes geomean speedup min max slower than old bf16 96 1.23x 0.98x 5.05x 1 fp32 96 1.14x 0.92x 4.91x 7 Per shape perf, bf16 input weight bias calls per token old us new us speedup old GB/s new GB/s 1x1x896 128x896 True Qwen/Qwen2.5-0.5B: 48 29.2…
摘要按规则整理自下方来源原文