The Lab Notes

DeepSeek

DeepSeek scraps plan to retire V4 Pro, citing user demand

DeepSeek's documentation showed Sept. 15, 2026, that it reversed a plan to retire V4 Pro days after announcing the shutdown alongside its V4.1-Flash launch.

An orange arrow looping backward over a rack of server hardware.

DeepSeek will keep selling access to its V4-Pro model unchanged, reversing a plan to shut it down that the company itself announced days earlier.

The reversal surfaced only inside DeepSeek’s live technical documentation, which stated, as confirmed Sept. 15, 2026, that “in response to user demand,” the company decided to continue providing API service for DeepSeek V4 Pro after Sept. 14, 2026, with the billing method unchanged, and that it would give further notice of any changes. That text sits inside a changelog entry dated Sept. 10, 2026, and on DeepSeek’s pricing page. The company has not issued a separate, dated statement announcing the reversal.

The retirement plan it reverses traces to DeepSeek’s Sept. 10 announcement of a new model, DeepSeek-V4.1-Flash, released that same day. In that post, DeepSeek said it was “phasing out V4-Pro”: starting at 04:00 UTC on Sept. 14, all requests to the deepseek-v4-pro API would be rerouted to V4.1-Flash and billed at V4.1-Flash’s lower rates, a plan the company said would hold “until V4.1-Pro launches.” That plan is now reversed.

V4.1-Flash is a 552-billion-parameter mixture-of-experts model built on what DeepSeek calls a “causal encoder-decoder” architecture, using eight billion active parameters to process input and 16 billion to generate output, according to the announcement. It adds native visual understanding and is generally available now on the DeepSeek API under the model name deepseek-flash, DeepSeek said. The company’s API changelog, also dated Sept. 10, 2026, records the same launch. DeepSeek retired its prior V4-Flash and V4-Flash-Vision-Exp models in the process, rerouting their old model names to V4.1-Flash at the Flash price.

DeepSeek said V4.1-Flash’s key-value cache needs one-quarter the high-bandwidth memory and one-eighth the solid-state storage of the prior generation, which it said matters because cache-hit charges make up a large share of agent costs.

DeepSeek’s pricing page, confirmed live Sept. 15, 2026, lists deepseek-flash at $0.15 per million input tokens and $0.60 per million output tokens off-peak, doubling to $0.30 and $1.20 at peak; cache-hit input runs $0.003 off-peak and $0.006 at peak. deepseek-v4-pro, version DeepSeek-V4-Pro-0813, costs $0.66 input and $1.98 output off-peak, $1.32 and $3.96 at peak, with cache-hit input at $0.022 off-peak and $0.044 at peak. Peak hours run 01:00-04:00 and 06:00-10:00 UTC Monday through Friday; every other hour is off-peak.

On its own reported testing, not an independent benchmark, DeepSeek’s changelog credits V4.1-Flash with a GPQA Diamond score of 90.9, a Codeforces rating of 3,471 and a Humanity’s Last Exam score of 36.8, or 39.1 on the text-only subset.

DeepSeek published V4.1-Flash’s weights on Hugging Face and said it would work with the open-source community on inference support and deployment options.

Analysis

Backing off a flagship retirement four days after announcing it, and doing so inside documentation rather than a dated statement, is an unusual way to handle a product reversal. DeepSeek has not said what changed between Sept. 10 and Sept. 14, or how much user pushback it took to keep V4 Pro running.