DeepSeek

Executive Summary

What it is: DeepSeek is an API-only supplier (no subscriptions) whose V4 family went fully GA in August: V4-Pro-0813 (1.6T/49B, 1M context, 384K output) and V4-Flash-0731, plus an experimental Flash vision model, all via OpenAI/Anthropic/Responses-compatible endpoints with opencode and Codex integration guides. Source: https://api-docs.deepseek.com/quick_start/pricing

What to watch out for: Peak/off-peak pricing went live with an effective increase: Pro output is $1.98/MTok off-peak and $3.96 peak versus July's flat $0.87, with peak defined as Beijing business hours (01:00-04:00 and 06:00-10:00 UTC weekdays). deepseek-chat/reasoner aliases deprecate Oct 24. Source: https://api-docs.deepseek.com/quick_start/pricing

Bottom line: The cheapest-frontier-API brand quietly got 2-4x more expensive depending on when you run, though off-peak Flash ($0.22/$0.66) remains the budget benchmark; budget by timezone, and note the Anthropic-format endpoint makes DeepSeek a drop-in Claude Code backend. Source: https://api-docs.deepseek.com/quick_start/pricing

Key Terms

  • Peak/off-peak pricing - now live: peak hours are 01:00-04:00 and 06:00-10:00 UTC Monday-Friday (09:00-12:00 and 14:00-18:00 Beijing); all other hours, including weekends, are off-peak at exactly half the peak rate. Applies to all billing items. Source: DeepSeek – Pricing
  • DeepSeek-V4-Pro-0813 - the general-availability Pro model (August 12-13), 1.6T total / 49B active MoE, 1M context, 384K max output, agent-focused post-training. Source: DeepSeek – Pricing , Aliyun – Newly Released Models
  • deepseek-v4-flash-vision-exp - experimental vision model on the Flash base; images are tokenized by dimensions and billed as input tokens. DeepSeek's first API vision capability. Source: DeepSeek – Pricing , DeepSeek – Vision
  • Anthropic-format API - Deepseek – Anthropic endpoint enabling Claude Code and Anthropic-SDK compatibility alongside OpenAI and Responses formats. Source: DeepSeek – Pricing

Latest Changes

Changes since the 2026-07 report.

July watch items verified:

New August changes:

  • Price change (increase, plus time-of-day multiplier): V4-Pro moved from flat $0.435/$0.87 per MTok (input miss/output) to $0.66/$1.98 off-peak and $1.32/$3.96 peak; V4-Flash moved from $0.14/$0.28 to $0.22/$0.66 off-peak and $0.44/$1.32 peak. Cache-hit: Pro $0.022 off-peak / $0.044 peak; Flash $0.007 / $0.014. Source: DeepSeek – Pricing
  • New capability: deepseek-v4-flash-vision-exp experimental vision API (August). Source: DeepSeek – Vision
  • Spec increase: max output raised to 384K tokens across models (1M context). Source: DeepSeek – Pricing
  • Deprecation scheduled: deepseek-chat and deepseek-reasoner API aliases will be deprecated October 24, 2026 (migration to V4 models). Source: Aitoolsrecap – Ai News August 31 2026
  • Distribution: V4-Pro-0813 and V4-Flash-0731 listed on Alibaba Bailian (Aug 14 / Aug 1). Source: Aliyun – Newly Released Models

Plans

DeepSeek is API-only: no subscription plans, top-up balance billing. Concurrency limits: 2,500 (Flash), 500 (Pro). Source: DeepSeek – Pricing

API Pricing

Model Input (cache miss) Input (cache hit) Output Notes
deepseek-v4-pro (peak) $1.32/MTok $0.044 $3.96/MTok Peak: Mon-Fri 01:00-04:00, 06:00-10:00 UTC
deepseek-v4-pro (off-peak) $0.66/MTok $0.022 $1.98/MTok All other hours incl. weekends
deepseek-v4-flash (peak) $0.44/MTok $0.014 $1.32/MTok
deepseek-v4-flash (off-peak) $0.22/MTok $0.007 $0.66/MTok
deepseek-v4-flash-vision-exp Same as flash; images billed as input tokens - - Experimental

Context 1M; max output 384K; thinking default (non-thinking optional); FIM and prefix completion in non-thinking mode. Source: DeepSeek – Pricing

Model Performance / Benchmarks

No new DeepSeek-published benchmark tables in August beyond the GA notes. Standing July figures: V4-Flash-0731: Terminal-Bench 2.1 82.7, DeepSWE 54.4, CyberGym 76.7, "far exceeds V4-Pro-Preview" on agentic tasks. Aggregator coverage of the GA claims V4-Pro-0813 posts the strongest agent-benchmark scores in the family. Sources: DeepSeek – Updates , Local-Ai-Zone – Ai Updates August 2026

Latest News

V4-Pro GA and the End of Flat-Rate DeepSeek (August 2026)

V4-Pro-0813 reached GA on August 12-13, and with it DeepSeek activated time-based pricing and raised effective rates: even off-peak, Pro output costs more than double July's flat $0.87. The supplier that built its brand on being the cheapest frontier API is now roughly 3x more expensive at Beijing business hours, though still far below Western frontier pricing (Opus 5 is $5/$25; Fable 5.1 $10/$50). Sources: DeepSeek – Pricing , Local-Ai-Zone – Ai Updates August 2026

Vision Arrives (Experimentally) (August 2026)

deepseek-v4-flash-vision-exp brings image input to the API for the first time, closing a long-standing gap for Claude Code-style agentic workflows that need screenshots. Source: DeepSeek – Vision

Community Signals

No major new HN threads on DeepSeek pricing were found in the August window; the July leaks (Liang Wenfeng's investor meeting, the gigawatt Inner Mongolia data center) remain the standing context for DeepSeek's compute constraints, which the new peak/off-peak pricing indirectly addresses (demand shaping). Sources: News – Item , News – Item

Enterprise Readiness

Feature Available? Details
SSO (SAML/OIDC) No API keys only.
SCIM No Not applicable.
Audit logs No Not documented.
IP indemnity No Not documented.
Data residency No Single-region inference; no residency options.
HIPAA No Not advertised.
Air-gapped / on-prem Partial V4 open weights (GGUF ecosystem) enable self-hosting. Source: GitHub – Deepseek Ai
SLA No Not published.
Admin controls (RBAC) No Balance/budget controls only.

Transparency Gaps

Gap Details Severity
Price-change communication The increase from flat to peak/off-peak rates (Pro output +127% off-peak, +355% peak vs July) was applied without a published migration note on the pricing page; the changelog framing is "policy" rather than a price-hike announcement. High
Vision model maturity "Exp" vision ships without published benchmark scores or a roadmap to Pro-class vision. Medium
Rate-limit tiers Concurrency caps (500/2,500) are published, but per-token TPM limits remain undocumented. Medium
GA benchmark delta V4-Pro-0813's "strongest agent scores in the family" claim has no published score table. Medium