← Back to events
ActiveAI

New Deepseek model V4.1-Flash cuts memory needs for AI agents

Photo: The Decoder

What happened

Deepseek releases V4.1-Flash, a multimodal model with 552 billion parameters that cuts KV cache memory to a quarter of its predecessor. On the DeepSWE coding benchmark, it narrowly beats Opus 5 and GPT-5.6 Sol, even though only 16 billion parameters are active per token. The model ships under the MIT license and targets much cheaper AI agents.

Summary assembled by rule from the sources below

Sources