v0.26.0
- vLLM: 11 events in the last 90 days
- vLLM: 11th Release in the last 90 days
- Previous: 3 days earlier · v0.26.0rc1
What happened
vLLM v0.26.0 Release Notes Highlights This release features 411 commits from 212 contributors (61 new)! New Inkling model family with a full support stack: base modeling ( #48799 ), piecewise CUDA graph support ( #48822 ), Hopper FA4 relative attention ( #48858 ), MTP=1 speculative decoding ( #48869 ), LoRA ( #48884 ), and standard ModelOpt NVFP4 quantization ( #48990 ). DeepSeek-V4 performance push across vendors:…
Summary assembled by rule from the sources below