LLM Cost Calculator for Batch Processing Pipelines
Estimate LLM costs for large-scale batch jobs - document processing, classification, embeddings, and scheduled AI workloads.
Recommended Setup
Cost Comparison: All Cloud Models
Based on 50M input + 15M output tokens/month
| Model | Provider | Monthly cost |
|---|---|---|
| Gemma 3 4B | $3.20 | |
| Llama 3.1 8B (Groq) | Groq | $3.70 |
| Gemma 3 12B | $3.95 | |
| gpt-oss-120b | OpenAI | $4.70 |
| Gemma 3n 4B | $4.80 | |
| Doubao Seed 2.0 Mini | Doubao (ByteDance) | $5.85 |
| Gemma 3 27B | $6.40 | |
| Gemma 4 26B A4B | $7.95 | |
| DeepSeek V4 Flash | DeepSeek | $8.00 |
| GPT-5 Nano | OpenAI | $8.50 |
| Qwen3.5-Flash | Qwen (Alibaba) | $8.50 |
| Qwen3 14B | Qwen (Alibaba) | $8.60 |
| GLM-4.7 Flash | Zhipu AI (GLM) | $9.00 |
| Doubao Pro 32K | Doubao (ByteDance) | $9.70 |
| Hunyuan TurboS | Hunyuan (Tencent) | $9.70 |
| Llama 4 Scout (Groq) | Groq | $10.60 |
| GPT-4.1 Nano | OpenAI | $11.00 |
| Gemini 2.5 Flash-Lite | $11.00 | |
| Qwen3 30B A3B | Qwen (Alibaba) | $11.25 |
| Qwen3 VL 32B Instruct | Qwen (Alibaba) | $11.30 |
| Gemma 4 31B | $11.40 | |
| Hunyuan T1 | Hunyuan (Tencent) | $15.40 |
| GPT-4o-mini | OpenAI | $16.50 |
| GPT-OSS 120B (Groq) | Groq | $16.50 |
| DeepSeek V3.2 | DeepSeek | $16.60 |
| Qwen3 Coder Next | Qwen (Alibaba) | $17.50 |
| DeepSeek V3 | DeepSeek | $22.00 |
| DeepSeek V3.1 | DeepSeek | $22.35 |
| Qwen3 VL 235B A22B Instruct | Qwen (Alibaba) | $23.20 |
| Qwen3 32B (Groq) | Groq | $23.35 |
| Qwen2.5 VL 72B Instruct | Qwen (Alibaba) | $23.75 |
| Qwen2.5 72B Instruct | Qwen (Alibaba) | $24.00 |
| Qwen3.5-Plus | Qwen (Alibaba) | $24.70 |
| Qwen3.6 Flash | Qwen (Alibaba) | $26.45 |
| GPT-5.4 Nano | OpenAI | $28.75 |
| DeepSeek V3 (Mar 2025) | DeepSeek | $30.00 |
| Gemini 3.1 Flash-Lite | $35.00 | |
| DeepSeek V4 Pro | DeepSeek | $35.05 |
| Qwen3 Coder 480B A35B | Qwen (Alibaba) | $38.00 |
| Llama 3.3 70B (Groq) | Groq | $41.35 |
| GPT-5 Mini | OpenAI | $42.50 |
| GPT-4.1 Mini | OpenAI | $44.00 |
| Qwen3.7 Plus | Qwen (Alibaba) | $44.00 |
| Qwen3.6 Plus | Qwen (Alibaba) | $45.75 |
| Qwen2.5 Coder 32B Instruct | Qwen (Alibaba) | $48.00 |
| Kimi K2.5 | Kimi (Moonshot AI) | $48.50 |
| Qwen3-Max | Qwen (Alibaba) | $49.80 |
| Gemini 2.5 Flash | $52.50 | |
| Llama 3.3 70B (Together) | Together AI | $57.20 |
| R1 0528 | DeepSeek | $57.25 |
| Doubao Seed 2.0 Pro | Doubao (ByteDance) | $59.05 |
| Kimi K2 | Kimi (Moonshot AI) | $63.00 |
| Kimi K2.5 (Together) | Together AI | $67.00 |
| Kimi K2 Thinking | Kimi (Moonshot AI) | $67.50 |
| DeepSeek R1 | DeepSeek | $72.50 |
| Qwen3 Coder Plus | Qwen (Alibaba) | $81.25 |
| Qwen3.5 397B (Together) | Together AI | $84.00 |
| Qwen3 Max Thinking | Qwen (Alibaba) | $97.50 |
| GLM-5 | Zhipu AI (GLM) | $98.00 |
| Claude 3.5 Haiku | Anthropic | $100 |
| GPT-5.4 Mini | OpenAI | $105 |
| Kimi K2.6 | Kimi (Moonshot AI) | $110 |
| Qwen3.7 Max | Qwen (Alibaba) | $119 |
| GLM-5-Turbo | Zhipu AI (GLM) | $120 |
| o4-mini | OpenAI | $121 |
| o3 Mini | OpenAI | $121 |
| Claude Haiku 4.5★ recommended | Anthropic | $125 |
| GLM-5.1 | Zhipu AI (GLM) | $136 |
| DeepSeek V4 Pro (Together) | Together AI | $171 |
| Moonshot V1 (128K) | Kimi (Moonshot AI) | $175 |
| Gemini 3.5 Flash | $210 | |
| GPT-5 | OpenAI | $213 |
| GPT-5 Codex | OpenAI | $213 |
| Gemini 2.5 Pro | $213 | |
| GPT-4.1 | OpenAI | $220 |
| o3 | OpenAI | $220 |
| o4 Mini Deep Research | OpenAI | $220 |
| GPT-4o | OpenAI | $275 |
| Gemini 3.1 Pro Preview | $280 | |
| GPT-5.4 | OpenAI | $350 |
| Claude Sonnet 4.6 | Anthropic | $375 |
| Claude Sonnet 4.5 | Anthropic | $375 |
| Claude Sonnet 4 | Anthropic | $375 |
| Claude Opus 4.7 | Anthropic | $625 |
| Claude Opus 4.6 | Anthropic | $625 |
| Claude Opus 4.8 | Anthropic | $625 |
| Claude Opus 4.5 | Anthropic | $625 |
| GPT-5.5 | OpenAI | $700 |
| o3 Deep Research | OpenAI | $1,100 |
| Claude Opus 4.8 (Fast) | Anthropic | $1,250 |
| o1 | OpenAI | $1,650 |
| Claude Opus 4.1 | Anthropic | $1,875 |
| Claude Opus 4 | Anthropic | $1,875 |
| o3 Pro | OpenAI | $2,200 |
| GPT-5 Pro | OpenAI | $2,550 |
| Claude Opus 4.7 (Fast) | Anthropic | $3,750 |
| Claude Opus 4.6 (Fast) | Anthropic | $3,750 |
| GPT-5.5 Pro | OpenAI | $4,200 |
| GPT-5.4 Pro | OpenAI | $4,200 |
| o1-pro | OpenAI | $16,500 |
Frequently Asked Questions
How much does it cost to process 1 million documents with GPT-4o?
At an average of 800 input tokens and 200 output tokens per document, processing 1M documents with GPT-4o costs approximately $4,000 (800M input tokens � $2.50/M + 200M output tokens � $10/M). Using OpenAI's Batch API at 50% discount, that drops to $2,000. Switching to GPT-4o mini reduces the same workload to about $270 - a 15x cost reduction if quality requirements permit.
What's the cheapest model for bulk classification tasks?
For binary or multi-class classification where completion length is short (10-50 tokens), Gemini 1.5 Flash is typically most cost-effective at $0.075/M input and $0.30/M output - under $20/month for 100M classification tokens. Claude Haiku 4.5 and GPT-4o mini are strong alternatives at $0.15/M input, often offering better structured-output reliability for complex taxonomies with many classes.
When should I switch from API to self-hosted for batch jobs?
Self-hosting becomes cost-competitive around 500M tokens/month for batch workloads with predictable scheduling. At that volume, a dedicated inference server - approximately $500-800/month including hardware amortization, electricity, and tooling - undercuts even Gemini Flash API pricing. The key condition is high hardware utilization: batch jobs that run predictably 8+ hours/day achieve 55-70% GPU utilization, making the economics work.