free tool · 23+ models · updated monthly
AI Token & Cost Calculator
Compare real API costs across 23+ AI models. Enter your token counts and request volume — see exact monthly and annual costs side by side, instantly.
Don't know your token count?
Estimate from word count or paste your actual prompt text.
1 word
~1.3 tokens
100 words
~133 tokens
1,000 words
~1,330 tokens
10,000 words
~13,300 tokens
Your usage
Quick presets
3,000
Total reqs/mo
3.0M
Input tokens/mo
1.5M
Output tokens/mo
Prompt caching
Reduces cost when the same prompt prefix is reused
Calculating for
- 📥 1,000 input tokens / req
- 📤 500 output tokens / req
- 🔁 100 req/day × 30 days
- 📦 3,000 total req/month
Cost comparison
Click a row for full breakdown
| Tier | Cache $/1M | Per req | Annual | Caps | |||||
|---|---|---|---|---|---|---|---|---|---|
🏆 🔵Gemini 1.5 Flash Extremely cheap. Good for batch processing. | Economy | $0.075 | $0.3 | $0.01875 | $0.000225 | $0.675 | $8.10 | 1.0M | visiontools |
🟣Mistral Small 3 Ultra-cheap. Great for classification and routing. | Economy | $0.1 | $0.3 | — | $0.000250 | $0.750 | $9.00 | 32K | tools |
⚪Llama 3.3 70B Instruct Best open-weight model. Via Groq/Together/Fireworks. | Economy | $0.12 | $0.3 | — | $0.000270 | $0.810 | $9.72 | 128K | tools |
⚪Llama 3.2 11B Vision Open-weight multimodal. Good vision at low cost. | Economy | $0.18 | $0.18 | — | $0.000270 | $0.810 | $9.72 | 128K | vision |
🟢GPT-4o mini Best price/performance for simple tasks. | Economy | $0.15 | $0.6 | $0.075 | $0.000450 | $1.35 | $16.20 | 128K | visiontoolsbatch |
🔶Command R Affordable. Good for RAG pipelines. | Economy | $0.15 | $0.6 | — | $0.000450 | $1.35 | $16.20 | 128K | tools |
🟣Codestral Specialised for code. 256K context. | Economy | $0.3 | $0.9 | — | $0.000750 | $2.25 | $27.00 | 256K | tools |
🔷DeepSeek V3 Incredibly cheap. GPT-4 level quality at fraction of cost. | Economy | $0.27 | $1.1 | $0.07 | $0.000820 | $2.46 | $29.52 | 64K | tools |
🔵Gemini 2.5 Flashreasoning Best value with 1M context. Very fast. | Economy | $0.3 | $2.5 | $0.075 | $0.001550 | $4.65 | $55.80 | 1.0M | visiontools |
🔷DeepSeek R1reasoning Open-source o1 alternative. Strong math and coding. | Standard | $0.55 | $2.19 | $0.14 | $0.001645 | $4.94 | $59.22 | 64K | |
🟠Claude Haiku 3.5 Fastest and cheapest Claude. Great for high-volume tasks. | Economy | $0.8 | $4 | $0.08 | $0.002800 | $8.40 | $100.80 | 200K | visiontoolsbatch |
🟢o4-minireasoning Fast, affordable reasoning. Great for coding and math. | Standard | $1.1 | $4.4 | $0.275 | $0.003300 | $9.90 | $118.80 | 200K | visiontoolsbatch |
🟢o3-minireasoning Cost-efficient reasoning model. | Standard | $1.1 | $4.4 | $0.55 | $0.003300 | $9.90 | $118.80 | 200K | toolsbatch |
🔵Gemini 1.5 Pro 2M token context window. | Standard | $1.25 | $5 | $0.3125 | $0.003750 | $11.25 | $135.00 | 2.0M | visiontools |
⚪Llama 3.1 405B Instruct Largest open-weights Llama. GPT-4 class. | Premium | $3 | $3 | — | $0.004500 | $13.50 | $162.00 | 128K | tools |
🟣Mistral Large 2 Strong EU-based model. Good for code and reasoning. | Standard | $2 | $6 | — | $0.005000 | $15.00 | $180.00 | 128K | tools |
🔵Gemini 2.5 Proreasoning 1M token context. Strong coding and reasoning. | Premium | $1.25 | $10 | $0.31 | $0.006250 | $18.75 | $225.00 | 1.0M | visiontools |
🔶Command R+ Strong RAG and enterprise search use cases. | Standard | $2.5 | $10 | — | $0.007500 | $22.50 | $270.00 | 128K | tools |
🟠Claude Sonnet 4.5 Best balance of intelligence and cost. | Standard | $3 | $15 | $0.3 | $0.010500 | $31.50 | $378.00 | 200K | visiontoolsbatch |
🟢GPT-4o Flagship GPT-4 class model. Strong all-rounder. | Standard | $5 | $15 | $2.5 | $0.012500 | $37.50 | $450.00 | 128K | visiontoolsbatch |
🟢GPT-4 Turbo Previous flagship. Superseded by GPT-4o. | Premium | $10 | $30 | — | $0.025000 | $75.00 | $900.00 | 128K | visiontoolsbatch |
🟢o3reasoning OpenAI's most capable reasoning model. | Ultra | $10 | $40 | $2.5 | $0.030000 | $90.00 | $1.1K | 200K | visiontoolsbatch |
🟠Claude Opus 4 Most intelligent Claude. Extended thinking available. | Ultra | $15 | $75 | $1.5 | $0.052500 | $157.50 | $1.9K | 200K | visiontoolsbatch |
How AI token pricing works
What is a token?
A token is roughly 0.75 English words, or about 4 characters. "Hello world" is 2 tokens. Most LLM APIs charge separately for input tokens (your prompt) and output tokens (the AI's response).
Why input and output are priced differently
Generating text (output) requires much more GPU compute than reading text (input). That's why output tokens cost 3-5× more per million tokens than input. Heavy output models like Claude Opus cost $75/M out vs $15/M in.
Prompt caching explained
When you send the same system prompt repeatedly (e.g. in a chatbot), caching stores it on the server. Cached reads cost 60-80% less than fresh input. Anthropic and Google both support this — it can save thousands per month at scale.
Batch API discounts
OpenAI and Anthropic offer ~50% discounts for async batch processing — tasks that don't need a real-time response. If you're doing data processing, evaluation, or bulk generation, batch APIs are the cheapest option.
Reasoning models cost more
Models like o3, o4-mini, and DeepSeek R1 'think' before answering by generating internal reasoning tokens. These tokens are typically charged at full output price, which means a complex query can use 5-20× more tokens than a standard model.
How to reduce your AI bill
Use cheaper models for simple tasks (GPT-4o-mini for classification, Haiku for routing). Enable prompt caching if your system prompt is long. Use batch API for non-real-time jobs. Compress your prompts — every 1K tokens saved = real money at scale.
AI model pricing reference (2026)
| Model | Provider | Input $/1M | Output $/1M | Context |
|---|---|---|---|---|
| 🟠 Claude Opus 4 | anthropic | $15 | $75 | 200K |
| 🟠 Claude Sonnet 4.5 | anthropic | $3 | $15 | 200K |
| 🟠 Claude Haiku 3.5 | anthropic | $0.8 | $4 | 200K |
| 🟢 GPT-4o | openai | $5 | $15 | 128K |
| 🟢 GPT-4o mini | openai | $0.15 | $0.6 | 128K |
| 🟢 o3 | openai | $10 | $40 | 200K |
| 🟢 o4-mini | openai | $1.1 | $4.4 | 200K |
| 🟢 o3-mini | openai | $1.1 | $4.4 | 200K |
| 🟢 GPT-4 Turbo | openai | $10 | $30 | 128K |
| 🔵 Gemini 2.5 Pro | $1.25 | $10 | 1.0M | |
| 🔵 Gemini 2.5 Flash | $0.3 | $2.5 | 1.0M | |
| 🔵 Gemini 1.5 Pro | $1.25 | $5 | 2.0M | |
| 🔵 Gemini 1.5 Flash | $0.075 | $0.3 | 1.0M | |
| ⚪ Llama 3.3 70B Instruct | meta | $0.12 | $0.3 | 128K |
| ⚪ Llama 3.1 405B Instruct | meta | $3 | $3 | 128K |
| ⚪ Llama 3.2 11B Vision | meta | $0.18 | $0.18 | 128K |
| 🟣 Mistral Large 2 | mistral | $2 | $6 | 128K |
| 🟣 Mistral Small 3 | mistral | $0.1 | $0.3 | 32K |
| 🟣 Codestral | mistral | $0.3 | $0.9 | 256K |
| 🔷 DeepSeek V3 | deepseek | $0.27 | $1.1 | 64K |
| 🔷 DeepSeek R1 | deepseek | $0.55 | $2.19 | 64K |
| 🔶 Command R+ | cohere | $2.5 | $10 | 128K |
| 🔶 Command R | cohere | $0.15 | $0.6 | 128K |
Prices in USD per million tokens. Last updated August 2026. Use the calculator above for exact cost estimates.