DeepSeek V4 Flash API Pricing and Cost Estimates
DeepSeek DeepSeek V4 Flash API pricing is $0.435 per 1M input tokens and $1.3 per 1M output tokens. Cached input is billed at $0.0145 per 1M tokens. The context window is 1M tokens. No batch tier is offered. Prices verified 2026-09-04 against the official pricing page.
Input / 1M tokens
$0.435
Output / 1M tokens
$1.3
Cached input / 1M
$0.0145
Context window
1M
Monthly cost at three workloads
| Workload | Tokens / month | Standard | 50% cache hit |
|---|---|---|---|
| Light (prototype / side project) | 5M in + 1M out | $3.47 | $2.42 |
| Medium (small production app) | 50M in + 10M out | $34.75 | $24.24 |
| Heavy (production at scale) | 500M in + 100M out | $348 | $242 |
Note: Peak-hour price (Beijing weekdays 09:00–12:00, 14:00–18:00) effective Aug 17, 2026; all other hours bill at 50% of peak.
Closest-priced alternatives
| Model | Input / 1M | Output / 1M | Medium workload / mo |
|---|---|---|---|
| DeepSeek DeepSeek V4 Flash | $0.435 | $1.3 | $34.75 |
| Google Gemini 3.1 Flash-Lite | $0.25 | $1.5 | $27.50 |
| OpenAI GPT-5.6 luna | $0.2 | $1.2 | $22.00 |
| Google Gemini 3.6 Flash | $0.75 | $3.75 | $75.00 |
| Zhipu (z.ai) GLM-5 | $1 | $3.2 | $82.00 |
Medium workload = 50M input + 10M output tokens per month at standard price, no caching. Prices verified 2026-09-04.
DeepSeek V4 Flash vs the rest of the DeepSeek lineup
| Model | Input / 1M | Output / 1M | Cached / 1M | Context | Medium workload / mo |
|---|---|---|---|---|---|
| DeepSeek V4 Flash | $0.435 | $1.3 | $0.0145 | 1M | $34.75 |
| DeepSeek V4 Pro | $1.3 | $3.91 | $0.043 | 1M | $104 |
Calculate with your own volume
LLM API Pricing Calculator
Compare current API prices for Claude, GPT, Gemini, DeepSeek, Kimi, Grok and other LLMs, and estimate your monthly bill from token volume. Includes prompt-caching prices and batch discounts that most comparison tables miss. Pick your current model to see exactly how much migrating to each alternative saves. Prices last verified 2026-09-04. Runs entirely in your browser — no upload, no signup.
Share of input tokens served from prompt cache
Adds a column showing what switching to each model saves (or costs) per month at this usage.
| Model | Input $/M | Output $/M | Cached $/M | Context | Est. monthly | vs current |
|---|---|---|---|---|---|---|
GPT-5.6 luna OpenAI | $0.2 | $1.2 | $0.02 | 1.05M | $22.00 | −$12.75 (−37%) |
Gemini 3.1 Flash-Lite | $0.25 | $1.5 | $0.025 | 1M | $27.50 | −$7.25 (−21%) |
DeepSeek V4 Flashcurrent DeepSeek · open-weight | $0.435 | $1.3 | $0.0145 | 1M | $34.75 | — |
Gemini 3.6 Flash | $0.75 | $3.75 | $0.075 | 1M | $75.00 | +$40.25 (+116%) |
GLM-5 Zhipu (z.ai) · open-weight | $1 | $3.2 | $0.2 | 200K | $82.00 | +$47.25 (+136%) |
Grok 4.20 xAI | $1.25 | $2.5 | $0.2 | 1M | $87.50 | +$52.75 (+152%) |
Kimi K2.7 Code Moonshot · open-weight | $0.95 | $4 | $0.19 | 262K | $87.50 | +$52.75 (+152%) |
Claude Haiku 4.5 Anthropic | $1 | $5 | $0.1 | 200K | $100 | +$65.25 (+188%) |
DeepSeek V4 Pro DeepSeek · open-weight | $1.3 | $3.91 | $0.043 | 1M | $104 | +$69.35 (+200%) |
GLM-5.2 Zhipu (z.ai) · open-weight | $1.4 | $4.4 | $0.26 | 200K | $114 | +$79.25 (+228%) |
Grok 4.6 xAI | $2 | $6 | $0.5 | 500K | $160 | +$125 (+360%) |
Grok 4.5 xAI | $2 | $6 | $0.3 | 500K | $160 | +$125 (+360%) |
Claude Sonnet 5 Anthropic | $2 | $10 | $0.2 | 1M | $200 | +$165 (+476%) |
GPT-5.6 terra OpenAI | $2 | $12 | $0.2 | 1.05M | $220 | +$185 (+533%) |
Gemini 3.1 Pro Preview | $2 | $12 | $0.2 | 1M | $220 | +$185 (+533%) |
Kimi K3 Moonshot · open-weight | $3 | $15 | $0.3 | 1.048576M | $300 | +$265 (+763%) |
GPT-5.6 sol OpenAI | $4 | $20 | $0.4 | 1.05M | $400 | +$365 (+1051%) |
Claude Opus 5 Anthropic | $5 | $25 | $0.5 | 1M | $500 | +$465 (+1339%) |
Claude Fable 5.1 Anthropic | $10 | $50 | $0.25 | 1M | $1000 | +$965 (+2778%) |
Claude Fable 5 Anthropic | $10 | $50 | $1 | 1M | $1000 | +$965 (+2778%) |
Prices are per 1 million tokens in USD, standard (non-batch) tier, verified 2026-09-04. Claude Fable 5.1: Released Sep 1, 2026. Anthropic's Mythos-class flagship tier above Opus 5 — same $10/$50 base rate as Fable 5, but cache reads drop to 2.5% of the input price ($0.25/MTok vs the 10% every other Claude model charges). Shares its underlying model with Claude Mythos 5.1, offered only to approved organizations. Claude Fable 5: Superseded by Claude Fable 5.1 (Sep 1, 2026), which keeps the same base rate but cuts cache reads to $0.25/MTok. Fable 5 remains available as a pinned snapshot with cache reads at the standard 10% ($1.00/MTok). Shares its underlying model with Claude Mythos 5. Claude Sonnet 5: The $2 / $10 launch price is now permanent; the planned Sep 2026 increase to $3 / $15 was cancelled. GPT-5.6 sol: Promotional price, available at least through Nov 21, 2026 (previously $5 / $30). Requests beyond the long-context threshold bill at 2× input / 1.5× output. GPT-5.6 terra: Requests beyond the long-context threshold bill at 2× input / 1.5× output. GPT-5.6 luna: Requests beyond the long-context threshold bill at 2× input / 1.5× output. Gemini 3.1 Pro Preview: Prompts over 200K tokens bill at $4 / $18. Gemini 3.6 Flash: Intro price through Dec 31, 2026; $1.50 / $7.50 from Jan 1, 2027. Grok 4.6: Requests over 200K tokens bill at 2×. Grok 4.5: Requests over 200K tokens bill at 2×. Kimi K3: Always-on reasoning; output includes thinking tokens. DeepSeek V4 Pro: Peak-hour price (Beijing weekdays 09:00–12:00, 14:00–18:00) effective Aug 17, 2026; all other hours bill at 50% of peak. DeepSeek V4 Flash: Peak-hour price (Beijing weekdays 09:00–12:00, 14:00–18:00) effective Aug 17, 2026; all other hours bill at 50% of peak.
Batch discount is applied only to models whose provider offers an async batch tier. Cached-input pricing models the read price; cache-write surcharges (Anthropic) are not included. Long-context surcharges apply above provider thresholds and are noted per model.
FAQ
How much does the DeepSeek V4 Flash API cost?
$0.435 per 1M input tokens and $1.3 per 1M output tokens, with cached input at $0.0145 per 1M. Note: Peak-hour price (Beijing weekdays 09:00–12:00, 14:00–18:00) effective Aug 17, 2026; all other hours bill at 50% of peak. (Verified 2026-09-04.)
What does a month of DeepSeek V4 Flash cost in practice?
A medium workload of 50M input + 10M output tokens per month runs about $34.75/mo, or about $24.24/mo if half the input hits the prompt cache.
How does prompt caching change DeepSeek V4 Flash pricing?
Input tokens served from the prompt cache bill at $0.0145 per 1M — about 3% of the standard input price. Repeated prefixes like system prompts and few-shot examples benefit most.
Is DeepSeek V4 Flash cheaper than Gemini 3.1 Flash-Lite?
At the same medium workload, DeepSeek V4 Flash costs about $34.75/mo versus about $27.50/mo for Google Gemini 3.1 Flash-Lite. Also weigh quality, latency, and context window (1M vs 1M tokens).
How does DeepSeek V4 Flash pricing compare with other DeepSeek models?
By input price: DeepSeek V4 Pro is $1.3/$3.91 per 1M tokens in/out (about 299% of DeepSeek V4 Flash's input price). See the full-lineup table above.