MagicTools
LLM API Pricing Calculator (all models)

GPT-5 mini API Pricing and Cost Estimates

OpenAI GPT-5 mini API pricing is $0.25 per 1M input tokens and $2 per 1M output tokens. Cached input is billed at $0.025 per 1M tokens. The context window is 400K tokens. The Batch API bills at 50% of standard. Prices verified 2026-07-27 against the official pricing page.

Input / 1M tokens

$0.25

Output / 1M tokens

$2

Cached input / 1M

$0.025

Context window

400K

Monthly cost at three workloads

WorkloadTokens / monthStandard50% cache hitBatch API
Light (prototype / side project)5M in + 1M out$3.25$2.69$1.63
Medium (small production app)50M in + 10M out$32.50$26.88$16.25
Heavy (production at scale)500M in + 100M out$325$269$163

Closest-priced alternatives

ModelInput / 1MOutput / 1MMedium workload / mo
OpenAI GPT-5 mini$0.25$2$32.50
DeepSeek DeepSeek V4 Pro$0.435$0.87$30.45
Google Gemini 3.1 Flash-Lite$0.25$1.5$27.50
DeepSeek DeepSeek V4 Flash$0.14$0.28$9.80
Zhipu (z.ai) GLM-5$1$3.2$82.00

Medium workload = 50M input + 10M output tokens per month at standard price, no caching. Prices verified 2026-07-27.

Calculate with your own volume

Back

LLM API Pricing Calculator

Compare current API prices for Claude, GPT, Gemini, DeepSeek, Kimi, Grok and other LLMs, and estimate your monthly bill from token volume. Includes prompt-caching prices and batch discounts that most comparison tables miss. Prices last verified 2026-07-27. Runs entirely in your browser — no upload, no signup.

M tokens
M tokens

Share of input tokens served from prompt cache

ModelInput $/MOutput $/MCached $/MContextEst. monthly

DeepSeek V4 Flash

DeepSeek · open-weight

$0.14$0.28$0.00281M$9.80

Gemini 3.1 Flash-Lite

Google

$0.25$1.5$0.0251M$27.50

DeepSeek V4 Pro

DeepSeek · open-weight

$0.435$0.87$0.0036251M$30.45

GPT-5 mini

OpenAI

$0.25$2$0.025400K$32.50

GLM-5

Zhipu (z.ai) · open-weight

$1$3.2$0.2200K$82.00

Grok 4.20

xAI

$1.25$2.5$0.21M$87.50

Kimi K2.7 Code

Moonshot · open-weight

$0.95$4$0.19262K$87.50

Claude Haiku 4.5

Anthropic

$1$5$0.1200K$100

GPT-5.6 luna

OpenAI

$1$6$0.11.05M$110

GLM-5.2

Zhipu (z.ai) · open-weight

$1.4$4.4$0.26200K$114

Gemini 3.6 Flash

Google

$1.5$7.5$0.151M$150

Grok 4.5

xAI

$2$6$0.3500K$160

Claude Sonnet 5

Anthropic

$2$10$0.21M$200

Gemini 3.1 Pro Preview

Google

$2$12$0.21M$220

GPT-5.6 terra

OpenAI

$2.5$15$0.251.05M$275

Kimi K3

Moonshot · open-weight

$3$15$0.31.048576M$300

Claude Opus 5

Anthropic

$5$25$0.51M$500

GPT-5.6 sol

OpenAI

$5$30$0.51.05M$550

Claude Fable 5

Anthropic

$10$50$11M$1000

Prices are per 1 million tokens in USD, standard (non-batch) tier, verified 2026-07-27 against official pricing pages. Claude Sonnet 5: Intro price through Aug 31, 2026; $3 / $15 after. GPT-5.6 sol: Requests beyond the long-context threshold bill at 2× input / 1.5× output. Gemini 3.1 Pro Preview: Prompts over 200K tokens bill at $4 / $18. Grok 4.5: Requests over 200K tokens bill at 2×. Kimi K3: Always-on reasoning; output includes thinking tokens.

Batch discount is applied only to models whose provider offers an async batch tier. Cached-input pricing models the read price; cache-write surcharges (Anthropic) are not included. Long-context surcharges apply above provider thresholds and are noted per model.

FAQ

How much does the GPT-5 mini API cost?

$0.25 per 1M input tokens and $2 per 1M output tokens, with cached input at $0.025 per 1M. (Verified 2026-07-27.)

What does a month of GPT-5 mini cost in practice?

A medium workload of 50M input + 10M output tokens per month runs about $32.50/mo, or about $26.88/mo if half the input hits the prompt cache. Offline jobs can use the Batch API at 50% of standard.

How does prompt caching change GPT-5 mini pricing?

Input tokens served from the prompt cache bill at $0.025 per 1M — about 10% of the standard input price. Repeated prefixes like system prompts and few-shot examples benefit most.

Is GPT-5 mini cheaper than DeepSeek V4 Pro?

At the same medium workload, GPT-5 mini costs about $32.50/mo versus about $30.45/mo for DeepSeek DeepSeek V4 Pro. Also weigh quality, latency, and context window (400K vs 1M tokens).

API pricing for other models