trunk/38d3f6fd0d6e7ea5b68ee3854e1033ff08a5c92c: [FSDP2] Separate sharded and unsharded gradient dtypes (#194434)
- PyTorch: 1780 events in the last 90 days
- PyTorch: 1758th Release in the last 90 days
- Previous: earlier the same day · ciflow/torchtitan/199419: [DCP] Read checkpoint items in parallel in FileSystemReader
What happened
Fixes #170648 feature 0 : when mp_policy.reduce_dtype is None, it follows orig model.parameters() .grad_dtype. .grad_dtype were silently ignored before this PR if orig grad_dtype is set, mp_policy.reduce_dtype is resolved to orig grad_dtype (including explicit None) if not set, mp_policy.reduce_dtype is resolved to orig param dtype this is BC breaking : mp_policy.reduce_dtype=None no longer defaults to mp_policy.par…
Summary assembled by rule from the sources below