trunk/326909e0a95f0fbd10fc85dffe0bf18858e12998: [MPS] Enable gemv kernels for linear (#198721)
- PyTorch: 1788 events in the last 90 days
- PyTorch: 1766th Release in the last 90 days
- Previous: earlier the same day · trunk/77700655e1ddd21002e444ea9d9a501cfe259da9: [torchtitan hash update] update the pinned torchtitan hash (#199225)
What happened
We have gemv kernels which only dispatched with torch.mm before 🤦♂️. This enables it for F.linear , i.e. what models actually run when decoding. Summary of speedups: dtype shapes geomean speedup min max slower than old bf16 96 1.23x 0.98x 5.05x 1 fp32 96 1.14x 0.92x 4.91x 7 Per shape perf, bf16 input weight bias calls per token old us new us speedup old GB/s new GB/s 1x1x896 128x896 True Qwen/Qwen2.5-0.5B: 48 29.2…
Summary assembled by rule from the sources below