- viable/strict/1791434203: [CUDA] Add fork/join execution on CUDA streams (#199128)
PyTorch Releases
- trunk/5ba7f2bf58f6177091cdf890429ffd5f5dc2139a: [CUDA] Add fork/join execution on CUDA streams (#199128)
PyTorch Releases
- viable/strict/1791320088: [CI] Stop double-running Inductor tests in CUDA trunk (#199291)
PyTorch Releases
- viable/strict/1790773504: Extend `native_layer_norm` param dtype check from CUDA to XPU (#198647)
PyTorch Releases
- viable/strict/1790725852: [inductor] Partition cross-device fallbacks from CUDA graphs (#190555)
PyTorch Releases
- v0.31.0rc1: [CI/Build] Skip the snapshot runtime on CUDA 12.x images (#59118)
vLLM Releases
- trunk/79f81b7884ec64cdba0ed2320394cb4b345922df: Update B200 smoke test workflow to CUDA 13.4 sm100 (#198912)
PyTorch Releases
- viable/strict/1790375212: Fix int32 overflow in cdist CUDA backward kernel (#198452)
PyTorch Releases
- trunk/c262fc64bd3a40834366246d22e9e7575f219a16: [CUDA] Capture Python launch stacks for CUDA graphs (#198419)
PyTorch Releases
- trunk/769b13f46a24d108e74a807bc3bdb5cc8b478758: Add tiled CUDA kernel for dense 2D transpose copies (#194310)
PyTorch Releases
- trunk/1426d759da6cd18784700a459d4050b240146ab9: Revert "Add tiled CUDA kernel for dense 2D transpose copies (#194310)"
PyTorch Releases
- trunk/4a81d5787ea55c5a1cb73afec4cd30a171f8cb5b: [inductor] Gate remaining CUDA TMA tests on CUDA devices (#197141)
PyTorch Releases
- viable/strict/1789252151: [inductor] Gate CUDA TMA tests on CUDA devices (#196636)
PyTorch Releases
- viable/strict/1788560676: Build Windows MAGMA for CUDA 13.4, drop CUDA 12.x builds (#196014)
PyTorch Releases
- trunk/cf8577d42b4750c22ad369fd8cef2b8e84bce74e: [CI] Add sm80 to the clang20 CUDA pull builds' arch list (#194685)
PyTorch Releases
- trunk/68d20d4ee3956ceb5fecbc1676112aa32d2ad7d9: Disable USE_FLASH_ATTENTION when no CUDA arch supports it (#194502)
PyTorch Releases