v0.34.4-rc0: mlx: speed up Qwen 3.8 prompt processing (#18550)
- Ollama: 45 events in the last 90 days
- Ollama: 43th Release in the last 90 days
- Previous: 4 days earlier · v0.34.3-rc1
What happened
mlx: speed up Qwen 3.8 prompt processing Use MLX's gated-delta kernel for long scans and fold dense MLP global scales into SwiGLU. address comments
Summary assembled by rule from the sources below