open-models

552B MoE opens DeepSeek's new architecture family

Promtime

open-models

DeepSeek has published V4.1 Flash on Hugging Face, a mixture-of-experts model with 552 billion parameters built on a new encoder-decoder structure and presented on Threads as the smallest model in the lab's new architecture family, with native visual understanding.

At a glance

  • Alongside the structural change, DeepSeek says the model went through a new pre-training method and a larger-scale reinforcement learning stage in post-training, with visual understanding native to the model.
  • At 552 billion parameters, V4.1 Flash is the smallest entry in the new family, far below the 1.6 trillion total parameters of the V4-Pro flagship that went live in August.
  • DeepSeek opened a limited beta of V4.1 Flash on September 8, 2026, with a two-day testing window and an automatic shutdown set for September 10, giving outside testers a fixed short trial.

The parameter count is the least novel part of the release. An encoder-decoder stack with vision trained in from the start reads as a rebuilt recipe rather than another checkpoint on an existing line, and calling V4.1 Flash the smallest member of a family implies larger models on the same structure. For practitioners tracking open weights, the release likely matters more as a preview of DeepSeek's next generation than as a model to deploy today.

V4.1 Flash runs 552 billion parameters on an encoder-decoder mixture of experts

V4.1 Flash runs 552 billion total parameters in a mixture-of-experts layout. DeepSeek describes the model as built on a new encoder-decoder structure and calls it the smallest model in its new architecture family. It is published on Hugging Face under the deepseek-ai account as DeepSeek-V4.1-Flash.

In a mixture-of-experts design only part of the parameter pool is active for any single token, so total size and compute per token diverge. Encoder-decoder stacks separate the processing of input from the generation of output, a structure long used in translation and multimodal systems. Decoder-only transformers, the dominant shape of recent large language models, fold both roles into a single stack.

DeepSeek changed both the pre-training method and the post-training reinforcement learning

DeepSeek says V4.1 Flash adopted a new pre-training method and underwent larger-scale reinforcement learning in its post-training stage. The same description from DeepSeek groups those training changes with the encoder-decoder structure and the native visual understanding as elements of the new architecture family.

Reinforcement learning after pre-training is the stage where labs shape reasoning behaviour and response style, and scaling it is one of the standard levers available once a base model exists. Pre-training changes sit earlier in the pipeline, before any of that tuning begins.

Native multimodality means image handling is trained into the model rather than added afterwards through a separate adapter at inference time. In serving terms one set of weights covers text and images, without a second vision model in the request path.

The limited beta ran from September 8 to an automatic shutdown on September 10

DeepSeek opened a limited beta for V4.1 Flash on September 8, 2026. The trial carried a two-day testing window and was set to shut down automatically on September 10. For scale, V4-Pro, the DeepSeek flagship released in August, totals 1.6 trillion parameters.

By total parameter count, V4-Pro is close to three times the size of V4.1 Flash. DeepSeek presents the two at different points of its lineup: V4-Pro as the flagship that shipped in August, V4.1 Flash as the smallest model of the new architecture family with native visual understanding.

Time-boxed betas with a hard cutoff are a familiar pattern in the industry, and DeepSeek has built its public profile on releasing open weights through Hugging Face rather than keeping models behind an API alone. The V4.1 Flash page on Hugging Face fits that pattern.

Timing for the larger models

DeepSeek has not said when the larger models implied by the description of V4.1 Flash as the smallest of the family will arrive, and it has not said whether V4-Pro will move to the encoder-decoder structure. The announcement also carries no active-parameter figure, no context length, no benchmark results and no throughput numbers for V4.1 Flash.

Comments

No comments yet. Be the first.

Join the conversation

Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.

We only use your name and avatar from Google. We never store your email address.

552B MoE opens DeepSeek's new architecture family · News