trunk/3f83700b397fdb04beb392d6b55fde076fe75083: Fix CUDA SDPA grid overflow for large batches and heads (#199594)
- PyTorch 近 90 天出现 1899 次
- PyTorch 近 90 天第 1875 次发版
- 上一次:同一天稍早 · ciflow/trunk/199719
发生了什么
Human Note For batch and head dims that exceed the grid cap, we just launch multiple kernels used some features found in #177985 Agent Report Agent details Prepared with AI assistance. CUDA grid.y/grid.z cannot exceed 65,535 blocks. Chunk the memory-efficient forward/backward launch grids while retaining full tensor dimensions and logical batch/head offsets. Kernels write directly into final outputs, preserving GQA…
摘要按规则整理自下方来源原文