← Back to events
ResolvedAIRelease0.26.0

v0.26.0

What happened

vLLM v0.26.0 Release Notes Highlights This release features 411 commits from 212 contributors (61 new)! New Inkling model family with a full support stack: base modeling ( #48799 ), piecewise CUDA graph support ( #48822 ), Hopper FA4 relative attention ( #48858 ), MTP=1 speculative decoding ( #48869 ), LoRA ( #48884 ), and standard ModelOpt NVFP4 quantization ( #48990 ). DeepSeek-V4 performance push across vendors:…

Summary assembled by rule from the sources below

Why it's spreading

Sources

Release