DeepSeek V4.1-Flash Ships With New Architecture
๐Ÿš€ News

DeepSeek V4.1-Flash Ships With New Architecture

DeepSeek added V4.1-Flash to its API on the date recorded in the September 10 changelog entry with 552B MoE parameters, native vision, and lower API rates.

The AI Dude ยท September 10, 2026 ยท 4 min read

Changelog Sequence

The September 10 entry in the change log at api-docs.deepseek.com states that DeepSeek-V4.1-Flash is the smallest model in a new architecture family with native multimodal visual understanding. The same entry records benchmark scores for the instruct model under maximum reasoning effort: GPQA Diamond at 90.9, HLE at 36.8 on the text subset, Codeforces rating at 3471, MathArena Apex at 65.6, Terminal-Bench 2.1 at 90.6, Terminal-Bench 3.0 at 30.0, Terminal-Bench 4.0 at 31.2, DeepSWE v1.1 at 74.2, ProgramBench at 20.3, NL2Repo-Bench at 65.4, CyberGym at 88.1, SEC-Bench Pro at 62.8, ExploitGym at 15.3, HLE with tools at 63.9, Automation-Bench at 54.8, Agents' Last Exam at 31.8, Chartography with tools at 78.9, BabyVision with tools at 89.6, and ZeroBench-main with tools at 49.0.

API changes listed on that date include the requirement to set the model name to deepseek-flash. The legacy names deepseek-v4-flash and deepseek-v4-flash-vision-exp remain accepted for compatibility but now route to the new model and bill at the Flash price. The entry notes that V4-Flash and V4-Flash-Vision-Exp are retired.

The September 10 changelog states: "After extensive testing, V4.1 Flash has comprehensively surpassed V4 Pro in performance, cost, speed, and total time, so we plan to retire V4 Pro in an orderly manner."

The pricing table on api-docs.deepseek.com shows the new rates in units per million tokens. Input cache-hit off-peak costs $0.003 and peak $0.006. Input cache-miss off-peak costs $0.15 and peak $0.3. Output off-peak costs $0.6 and peak $1.2. The concurrency limit stands at 2500. Off-peak hours are defined as all times outside 01:00-04:00 and 06:00-10:00 UTC Monday through Friday. The page states that these prices took effect at 04:00 UTC on September 10, 2026.

Metricdeepseek-flashdeepseek-v4-pro
Input cache hit off-peak$0.003$0.022
Input cache hit peak$0.006$0.044
Input cache miss off-peak$0.15$0.66
Input cache miss peak$0.3$1.32
Output off-peak$0.6$1.98
Output peak$1.2$3.96
Concurrency limit2500500

Earlier entries in the same changelog trace the path. The August 21 post introduced the experimental vision variant DeepSeek-V4-Flash-Vision-Exp with scores such as Terminal-Bench 2.1 at 83.9 and DeepSWE at 59.3. The August 13 post rolled out the GA version of V4-Pro with HLE scores of 42.7 without tools and 60.0 with tools, plus Terminal-Bench 2.1 at 87.9. The July 31 post released the prior V4-Flash version with Terminal-Bench 2.1 at 82.7 and DeepSWE at 54.4. The September 10 post records that V4.1-Flash outperforms V4-Pro across performance, cost, speed, and total time.

The Hugging Face page for deepseek-ai/DeepSeek-V4.1-Flash lists a 1M context length, MIT license, support for both non-thinking and thinking modes, and a vision encoder trained from scratch. The technical report hosted there states the model was pretrained on a 45T-token multimodal corpus and adopts a Causal Encoder-Decoder architecture with 552B backbone parameters, 8B active parameters during prefill, and 16B during decode.

Cost Compression Pattern

The sequence of releases shows DeepSeek repeatedly reducing KV cache size and active parameters while adding multimodal support and raising concurrency limits. The September 10 announcement states that V4.1-Flash KV cache needs one-quarter the HBM and one-eighth the SSD storage of the previous generation. The technical report adds that SWA Bounded Replay and CSA2 attention reduce the global KV cache footprint to 890 bytes per token. The same document projects that the architecture can scale to larger models without proportional cost increases.

Official partners WorkBuddy including CodeBuddy and OpenCode now support the model. The pricing page shows off-peak rates remain half of peak rates, continuing the pattern that began with the August 13 V4-Pro release. The September 10 changelog entry states that all deepseek-v4-pro requests will route to V4.1-Flash at the new rates starting 04:00 UTC on September 14, 2026, until V4.1-Pro launches. This routing decision follows the documented outperformance of V4.1-Flash on Terminal-Bench 2.1 at 90.6 versus 87.9 for V4-Pro, DeepSWE v1.1 at 74.2 versus 62.7, and CyberGym at 88.1 versus 83.3 under identical evaluation settings.

The pattern points to sustained focus on serving higher throughput at lower per-token cost through the Causal Encoder-Decoder design and continued cache compression. The Hugging Face technical report notes plans to work with the open-source community on inference support and to explore additional deployment options, including large-scale clusters of 2,000 GPUs plus storage.

Against GPT-5.6 Sol the new model records Terminal-Bench 2.1 at 90.6 versus 88.8 and DeepSWE v1.1 at 74.2 versus 73.0 on the same scaffolds and settings listed in the technical report comparison table.

Migration Reversal Conditions

From 12:00 Beijing Time on September 14, 2026, and until the future release of V4.1 Pro, requests to deepseek-v4-pro will all be routed to V4.1 Flash and billed at the V4.1 Flash price. The changelog entry on api-docs.deepseek.com records this schedule explicitly. The pattern would end if that routing date slipped or if V4.1-Pro launched before the migration completed.

A second break would occur if measured concurrency or cache-hit rates on the new model fell below the 2500 limit and pricing tiers published on September 10. The announcement leaves open the exact date for V4.1-Pro and the final open-weights release timeline beyond the current Hugging Face upload. Those two milestones remain the clearest points that could alter the observed direction.

DeepSeek V4.1-Flashdeepseek-flash APIDeepSeek pricingV4.1 vs V4-Pro

Keep reading

Weekly issue

The 5 AI tools that mattered this week.

One email, Fridays. No spam, unsubscribe anytime.