DeepSeek, the Chinese lab whose ultra-low pricing shook the AI industry in early 2025, announced that its V4 system will transition from preview to an official release in mid-July. The update arrives with an unprecedented twist: for the first time, the company will charge developers twice as much to use the model during peak daytime hours.

According to a report by TechNode, the higher rates will apply on weekdays from 9 a.m. to noon and from 2 p.m. to 6 p.m. Beijing time. Outside of these windows, pricing remains unchanged. The company noted that users will receive an email notification 24 hours prior to any billing adjustments.

The pricing breakdown illustrates how the new system operates. For the flagship model, deepseek-v4-pro, output tokens normally cost 6 yuan per million (approximately $0.85). During peak hours, that rate doubles to 12 yuan (around $1.70). Similarly, the lighter deepseek-v4-flash model will scale from 2 yuan to 4 yuan per million tokens during the same periods. Tokens represent the small segments of text a model processes and generates, serving as the industry standard for API billing.

V4 originally debuted as an open-weight preview on April 24. The official production version retains the preview's standout feature: a 1-million-token context window across the entire lineup, allowing the model to process massive amounts of text simultaneously. DeepSeek highlights that the update also delivers significant improvements in coding, mathematics, and multi-step agentic tasks. The firm's legacy models, deepseek-chat and deepseek-reasoner, are scheduled for retirement on July 24.

This move is particularly striking because DeepSeek single-handedly ignited China's AI price war with its aggressively discounted tokens, recently making a 75% price cut on V4 permanent. Implementing a peak-hour surcharge marks a sharp pivot in strategy. The company attributes this to hardware constraints: during high-traffic intervals, compute demand outpaces available chip capacity. The surge pricing aims to incentivize shifting non-urgent workloads to off-peak hours, which DeepSeek claims will stabilize service reliability under heavy loads.

Beyond capacity management, the pricing structure reflects genuine architectural efficiency rather than a simple loss-leader strategy. Reviewing the preview in April, independent developer Simon Willison noted that V4-Pro was the most cost-effective among major frontier models, crediting architectural optimizations that reduce the memory footprint per request. This efficient baseline ensures DeepSeek remains remarkably affordable even with the surcharge; competitors like OpenAI and Anthropic continue to charge several times more per token.

The geographic timing works to the advantage of international developers. Beijing's peak windows align with overnight hours in the US (roughly 9 p.m. to 6 a.m. Eastern Time), meaning American teams operating during their standard business day will automatically benefit from off-peak rates. For others, however, executing the exact same API call will cost double at 1 p.m. compared to the evening, turning job scheduling into a direct financial decision. It is worth noting that DeepSeek's public pricing page still displays flat rates, so the surcharges should be verified once officially updated. Ultimately, this shift signals that the era of strictly flat, ever-decreasing AI pricing may be ending—and the industry's lowest-cost leader is the one leading the shift.