← 返回事件
持续讨论AI发版29.2

trunk/326909e0a95f0fbd10fc85dffe0bf18858e12998: [MPS] Enable gemv kernels for linear (#198721)

发生了什么

We have gemv kernels which only dispatched with torch.mm before 🤦‍♂️. This enables it for F.linear , i.e. what models actually run when decoding. Summary of speedups: dtype shapes geomean speedup min max slower than old bf16 96 1.23x 0.98x 5.05x 1 fp32 96 1.14x 0.92x 4.91x 7 Per shape perf, bf16 input weight bias calls per token old us new us speedup old GB/s new GB/s 1x1x896 128x896 True Qwen/Qwen2.5-0.5B: 48 29.2…

摘要按规则整理自下方来源原文

为什么在扩散

来源