DeepSeek formally released its V4 Pro model on August 13, 2026. According to the company's announcement, the model is available through its app, website and API. DeepSeek positions V4 Pro as an upgrade focused on AI-agent tasks, including tool use and multi-step workflows. The same release also introduces a new API pricing structure for V4 Pro and V4 Flash that separates peak and off-peak periods.
V4 Pro is available through three channels
DeepSeek's official release note says V4 Pro is now offered through the app, web product and API. Reuters independently confirmed the formal launch and reported the company's emphasis on improved agent capabilities. The news agency listed V4-Pro-0813 at $1.32 per million input tokens and $3.96 per million output tokens. Those amounts match the peak-period levels shown in DeepSeek's new official pricing table.
The new API schedule starts on August 16
DeepSeek's official pricing page says the new rates take effect at 16:00 UTC on August 16, 2026. For V4 Pro, the off-peak table lists $0.022 per million tokens for cache hits, $0.66 for uncached input and $1.98 for output. During peak periods, the same items are priced at $0.044, $1.32 and $3.96. The displayed off-peak amounts are therefore half of the corresponding peak rates.
For V4 Flash, the off-peak prices are $0.007 for cache hits, $0.22 for uncached input and $0.66 for output. Peak-period prices are $0.014, $0.44 and $1.32 respectively. Every figure is quoted per million tokens. The official table also defines the periods used for peak and off-peak billing, so teams operating API workloads will need to account for timing as well as model selection and token volume.
How to read the table
The schedule separates cache-hit input, uncached input and model output. A request can contain different quantities in all three categories, so looking only at the input or output headline does not show the full cost. The off-peak figure listed for each item is half of its corresponding peak rate. Workload estimates therefore need to combine token mix, cache behavior and call timing.
The gap between Pro and Flash is wider
Reuters reported that V4 Pro's quoted $1.32 input price is about nine times the $0.14 comparison figure for V4 Flash, while the $3.96 output price is about fourteen times the $0.28 Flash figure. DeepSeek's new official schedule, however, contains different prices depending on time and cache status. A single ratio should not be applied to every request. The actual bill will depend on the model, whether the input is served from cache, the amount of output and the billing period in which the call occurs.
What independent testing indicates
Reuters cited an Artificial Analysis evaluation in which the reasoning version of V4 Pro scored 53 on the Intelligence Index, compared with 40 for V4 Flash. The index combines nine areas, including tool use, coding, scientific reasoning, long-context performance and agentic work. These are results from an independent benchmarking organization, not a performance guarantee from DeepSeek. Results in production can vary with the task, prompt design, connected tools and evaluation method.
What changes for users
The launch brings two distinct changes. V4 Pro is broadly available as the company's higher-tier option for agent-focused work, while API billing moves to a time-sensitive structure. App and web users mainly receive the new model option. Developers and organizations face the operational impact of the revised pricing table. DeepSeek has published the effective time and the rates for both models; the effect on a particular workload must be calculated from its token volume, output length, cache behavior and call schedule.
The most reliable way to compare costs is to use the official table rather than a single headline price. Reuters provides independent confirmation of the launch and useful context on the Pro-to-Flash gap, but DeepSeek's documentation is the primary source for the exact billing amounts and start time. No claim about overall value or suitability can be made from price and benchmark figures alone; those judgments require testing against a specific workload.
