DeepSeek Makes V4-Pro Price Cut Permanent
DeepSeek locked in a 75% price cut on its flagship V4-Pro model. Here's what it means for AI pricing and the global compute race.
DeepSeek made its 75% price cut on the flagship V4-Pro model permanent, effective May 23, 2026. What started as a promotional discount is now the sticker price, so API costs for V4-Pro sit at a quarter of their original levels with no expiration date attached. Reuters and Engadget confirmed the 75% cut, which landed the same day across DeepSeek's official pricing page.
The claim underneath the announcement is that DeepSeek can sustain rock-bottom pricing on a 1M-context, frontier-class model indefinitely. The timing sharpens it, arriving as OpenAI, Google and Anthropic all push their own agentic models into production.
Permanence removes the budgeting risk
DeepSeek had been running the 75% discount on V4-Pro as a time-limited promotion. Developers building on the API knew the price could snap back at any point, and that uncertainty made it risky to architect cost-sensitive applications around DeepSeek's rates. You would budget for the discount and then potentially absorb a 4x increase overnight.
Making the cut permanent, per DeepSeek's own documentation at api-docs.deepseek.com, removes that risk. Teams can plan long-term around V4-Pro's reduced rates without hedging for a reversion. For startups running inference-heavy workloads, whether RAG pipelines, multi-turn agents or document processing across the 1M context window, the budgeting math just became predictable.
The promotional period reads in hindsight less like a promotion and more like a market test. DeepSeek watched usage patterns at the lower price, confirmed the unit economics worked, and locked it in. It is a more sophisticated pricing move than the headline suggests.
The pricing floor this sets for everyone else
The AI API pricing war has been intensifying all year, mostly in increments: a new tier here, a batch discount there. Permanently cutting a flagship model's price by 75% is a different order of signal. It forces every other provider to answer whether they can match it, and if not, why developers should pay more.
The competitive landscape right now looks like this. OpenAI has been pushing GPT-5 and its variants, with pricing that remains significantly higher for comparable context windows and capability tiers. Google just launched Gemini 3.5 Flash with aggressive agentic pricing, though its flagship Pro models still sit at a premium. Anthropic offers Claude with strong coding and reasoning performance and has not engaged in the same kind of aggressive price cutting. Mistral released Medium 3.5 as an open-weight model, competing on openness rather than pure API price.
DeepSeek's move compresses the pricing floor for the entire market. Even where Western labs decline to match the per-token rates, they face growing pressure to justify premiums with measurable capability gaps rather than brand trust or ecosystem lock-in.
Export controls and the cost basis underneath the cut
There is a strategic dimension beyond competitive tactics. DeepSeek built its infrastructure under US export controls that restrict access to NVIDIA's most advanced chips. Offering a frontier-class model at 75% below its own original pricing, and sustaining that permanently, says something about the efficiency of its training and inference stack.
The V4-Pro announcement comes from a company that has consistently done more with less. DeepSeek's earlier models drew attention for achieving competitive benchmark scores with reportedly fewer GPUs and less compute than Western counterparts. Whether the credit belongs to architectural innovation, distillation techniques or different cost structures in Chinese data centers, the outcome is the same: DeepSeek can price aggressively because its cost basis is genuinely lower, not because it is burning venture money on subsidized pricing.
For Western labs, the permanence is the uncomfortable part. A temporary discount can be a growth hack. A permanent 75% cut points at structural cost advantages that will not disappear when a promotional budget runs out.
The considerations that outlast the price tag
Before anyone migrates production workloads to V4-Pro, several things weigh against the savings.
Data residency and compliance come first. DeepSeek's infrastructure is China-based, and for enterprises in regulated industries such as healthcare, finance and government, that may be disqualifying regardless of price. Data sovereignty rules in the EU, US and other jurisdictions can make Chinese-hosted inference a compliance headache on its own.
API reliability matters just as much, because price means nothing if uptime does not meet production requirements. DeepSeek's API has had availability issues in the past, and the company does not publish the kind of SLA guarantees OpenAI or Google offer enterprise customers.
Capability against price is the third consideration. V4-Pro is competitive on benchmarks, and competitive is not the same as best. For complex agentic workflows, nuanced instruction following or safety-critical applications, the cheapest model is not automatically the right one, so evaluate on your actual use case rather than published leaderboard scores.
Long-term platform risk closes the list. Geopolitical tension between the US and China has not eased, and building critical infrastructure on a Chinese AI provider carries risks that have nothing to do with the technology itself.
For a side project, a prototype or a cost-sensitive application where data sensitivity is not a concern, V4-Pro at these prices is genuinely compelling. For enterprise software that has to pass a security review, the price advantage may never get a chance to matter.
Inference has been commoditizing on a schedule
Zoom out and the permanent cut fits a pattern that has been accelerating through 2026. AI inference is commoditizing faster than most people expected, and the progression is legible.
Through 2023 and 2024, API pricing was a moat, and OpenAI could charge premium rates because the alternatives were limited and significantly worse. In 2025 the open-weight models, Llama, Mistral and Qwen, closed the capability gap, self-hosted inference became viable for many use cases, and API prices started falling. By 2026 multiple frontier-class models compete on price, with DeepSeek, Mistral and others willing to price at or near cost, turning the intelligence layer into a commodity input rather than a premium product.
This does not mean the AI labs are in trouble. Enormous value remains in ecosystems, tooling, safety infrastructure and enterprise relationships. But the raw model API is increasingly a commodity, and DeepSeek just made that harder to deny.
Numbers DeepSeek has not published
A few open questions are worth flagging. Exact per-token pricing is one: the 75% reduction from original V4-Pro rates is confirmed by Reuters and DeepSeek's pricing page, but the specific input and output token costs vary by whether you are hitting cache, using the 1M context, or running batch inference, so check DeepSeek's current pricing page directly for the latest figures.
Rate limits at the new price are another, since permanent cuts sometimes arrive alongside tighter limits or reduced throughput guarantees, and DeepSeek has not publicly addressed whether access terms change. The third is whether this triggers responses. OpenAI and Google have the margins to cut prices if they choose to compete on cost, and whether they will, or whether they double down on premium positioning, is unresolved.
Where the V4-Pro cut leaves the market
A small announcement carries large implications here. It confirms that frontier-class AI inference can be delivered at a fraction of what Western labs currently charge. It puts pressure on every API provider to justify pricing with clear capability advantages. And it shows that China's AI ecosystem, despite chip export controls, is competing on economics rather than benchmarks alone.
For developers the takeaway is straightforward: if V4-Pro fits your use case and your compliance requirements, there is no longer a reason to wait for the discount to expire, because it will not. For everyone else, the number that matters is not the price but what the price implies about where AI costs are heading across the whole market.
Keep reading
News
AI21 Labs Cuts 60% of Staff, Bets on Maestro
AI21 Labs slashes over 60% of staff, drops foundation models, and pivots to its Maestro agent optimization platform after Nebius acquisition talks collapse.
News
Alibaba Bans Claude Code Over Security Concerns
Alibaba told staff to remove Anthropic's Claude Code by July 10 over security concerns. Here's what triggered the ban and what it signals.
News
Anthropic Acquires Stainless: What It Means for AI
Anthropic bought Stainless, the SDK generator behind OpenAI and Cloudflare's client libraries. Here's the strategic play for AI agents.