DeepSeek-V4-Pro-0813 Goes GA With 87.9% on TerminalBench and Peak/Off-Peak Pricing That Halves API Costs


When DeepSeek-V4-Flash-0731 shipped on July 31 and posted an 82.7% score on TerminalBench, it was the fast, agentic half of a two-part V4 family release. On August 13, 2026, the second half arrived: DeepSeek-V4-Pro-0813, the generally available version of the full-scale model that was in preview since the Flash launch. The GA release closes the preview window and brings the Pro model's thinking effort controls, DSpark speculative decoding, and a new pricing structure built around peak and off-peak hours to production deployments.
Vitals
- Architecture: 1.6T total parameters, ~49B active per token (Sparse MoE with DSpark speculative decoding)
- Context window: 1,000,000 tokens input; 384,000 tokens maximum output
- Thinking effort levels: Non-thinking / High / Max
- API ID:
deepseek-v4-pro(no API change required from preview) - License: Open weights (Hugging Face + GGUF); API-accessible
- Peak pricing (01:00–04:00, 06:00–10:00 UTC): $0.044 input (cache hit) / $1.32 input (cache miss) / $3.96 output per 1M tokens
- Off-peak pricing (all other hours): $0.022 / $0.66 / $1.98 per 1M tokens
- Effective date: August 16, 2026
- Release date: August 13, 2026
Benchmark breakdown
Autonomous coding: On Terminal-Bench 2.1, the model scored 87.9% — placing it in a cluster at the frontier alongside Grok 4.6 (88.4%) and Claude Opus 5. On SWE-bench Verified, vendor-reported results show approximately 80.6% in standard mode, rising under max thinking effort. DeepSeek notes that the Pro model significantly outperforms V4-Flash on complex multi-step reasoning tasks while V4-Flash maintains its edge on raw throughput. (Sources: DeepSeek blog, vendor-reported.)
General coding quality: The model scored 3,206 on Codeforces rating and 93.5% on LiveCodeBench (vendor-reported). On GPQA Diamond, which tests graduate-level science reasoning, it scored 90.1% — consistent with the broader pattern that V4-Pro targets tasks requiring sustained reasoning chains more than speed-optimized completion.
Thinking effort sensitivity: DeepSeek's own analysis shows the gap between High and Max thinking effort is smaller than the gap between Non-thinking and High for most tasks. On SWE-bench, moving from Non-thinking to High adds roughly 10–12 points; High to Max adds ~1.2 points. For cost-sensitive production deployments, High effort appears to be the practical default — Max is reserved for tasks where marginal accuracy gains justify the additional token spend.
DSpark speculative decoding
The DSpark module is DeepSeek's answer to the latency problem that comes with large MoE models. Speculative decoding uses a smaller "draft" model to propose multiple candidate tokens, which the large model then verifies in a single forward pass — effectively increasing throughput without changing the model's outputs. DeepSeek has not fully disclosed the architecture of DSpark's draft model, but third-party benchmarkers have noted measurably lower time-to-first-token versus the preview build under equivalent load. For production deployments where latency matters, this is the most significant infrastructure change in the GA release.
The peak/off-peak pricing model
The most structurally interesting aspect of the GA release is not the benchmark numbers — it's the pricing design. Effective August 16, off-peak API calls to V4-Pro are 50% cheaper than peak calls. At peak hours (01:00–04:00 UTC and 06:00–10:00 UTC, a deliberately narrow window that targets DeepSeek's own peak infrastructure demand), input cache miss costs $1.32/M tokens and output costs $3.96/M tokens. Off-peak, those drop to $0.66/M and $1.98/M respectively.
For batch processing workloads — document analysis, test generation, codebase indexing — that can be scheduled outside peak hours, this is a meaningful cost lever. At $1.98/M output tokens off-peak, V4-Pro is priced competitively against DeepSeek-V4-Flash-0731's standard pricing, despite being the full-scale Pro model. Developers using the deepseek-v4-pro API endpoint receive the GA model without any code changes required.
Try it yourself
Illustrative prompts for testing V4-Pro's reported strengths.
- For thinking effort comparison: run the same complex debugging task with
thinking: "high"andthinking: "max"and compare output quality versus token cost. - For off-peak batch processing: schedule large-scale code analysis or test generation jobs to run between 04:00–06:00 UTC for 50% cost reduction.
- On Terminal-Bench-class tasks:
"You have shell access. The test suite in /tests/ is failing with 47 errors after the latest refactor. Diagnose the root cause, fix the underlying issue in /src/, and verify the full test suite passes."
What this means for the V4 family
The V4 family is now complete: V4-Flash (speed, throughput, agentic loops) and V4-Pro (maximum reasoning, complex tasks, high output token budgets). The off-peak pricing model is a differentiated move — no other frontier API provider has introduced time-of-day pricing at this scale, and it creates a new optimization dimension for production AI infrastructure teams. Combined with the open weights on Hugging Face and GGUF support, DeepSeek continues the strategy of maximum accessibility across both self-hosted and API deployment models. The key watch point from here is independent evaluation of the TerminalBench and SWE-bench results against third-party harnesses — DeepSeek's self-reported numbers have historically held up well, but verification is pending.


