AI model comparison
Compare listed API pricing, context windows, capabilities, and reported benchmark scores across 20 models.
Pricing and benchmarks change frequently and may use different evaluation settings. Verify provider documentation before production decisions. Latest item verification: June 10, 2026.
Which model should I use?
Answer 3 questions — get a personalised recommendation
What's your primary task?
| Tier | MM | OS | Speed | |||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Claude Fable 5Anthropic | Flagship | 1M | $10$50 out | — | 92.6% | 95% | — | ✓ | — | Slow |
| Claude Haiku 4.5Anthropic | Efficient | 200k | $0.25$1.25 out | 81% | 58% | 40% | 31% | ✓ | — | Fast |
| Claude Opus 4.8Anthropic | Flagship | 200k | $15$75 out | 90% | 86% | 72% | 74% | ✓ | — | Slow |
| Claude Sonnet 4.6Anthropic | Balanced | 200k | $3$15 out | 88% | 78% | 57% | 62% | ✓ | — | Medium |
| Codestral 25.01Mistral AI | Balanced | 256k | $0.30$0.90 out | — | — | — | — | — | — | Fast |
| DeepSeek R1DeepSeek | Reasoning | 128k | $0.55$2.19 out | 90.8% | 71.5% | 49.2% | 79.8% | — | ✓ | Slow |
| DeepSeek V4-ProDeepSeek | Open | 1M | $0.27$1.1 out | 89.5% | 73% | 80.6% | 85% | — | ✓ | Fast |
| Gemini 3.1 ProGoogle | Flagship | 1M | $1.25$10 out | 90.5% | 94.3% | 65% | 87.5% | ✓ | — | Medium |
| Gemini 3.1 UltraGoogle | Flagship | 2M | $2.5$15 out | 91.5% | 91% | 70% | 89% | ✓ | — | Slow |
| Gemini 3.5 FlashGoogle | Efficient | 1M | $1.5$9 out | 88.5% | 80% | 61% | 78% | ✓ | — | Fast |
| GPT-4oOpenAI | Balanced | 128k | $2.5$10 out | 88.7% | 53.6% | 33% | 13% | ✓ | — | Fast |
| GPT-4o miniOpenAI | Efficient | 128k | $0.15$0.60 out | 82% | — | — | — | ✓ | — | Fast |
| GPT-5.5OpenAI | Flagship | 1M | $10$30 out | 92% | 88% | 69% | 81% | ✓ | — | Medium |
| Grok 4.3xAI | Flagship | 256k | $3$15 out | 92.9% | 85.5% | 50% | 93.3% | ✓ | — | Medium |
| Llama 4 MaverickMeta | Open | 1M | $0.27$0.85 out | 85.5% | 69.8% | 38% | 49% | ✓ | ✓ | Fast |
| Llama 4 ScoutMeta | Open | 10M | $0.11$0.34 out | 79.6% | — | — | — | ✓ | ✓ | Fast |
| Mistral Large 2Mistral AI | Balanced | 128k | $2$6 out | 84% | 56.6% | — | — | — | — | Medium |
| Mistral Medium 3.5Mistral AI | Open | 256k | $1.5$7.5 out | 86% | 65% | 77.6% | 70% | ✓ | ✓ | Medium |
| o3OpenAI | Reasoning | 200k | $10$40 out | 91.8% | 87.7% | 71.7% | 88% | ✓ | — | Slow |
| o4-miniOpenAI | Reasoning | 200k | $1.1$4.4 out | 90% | 82.6% | 68.1% | 83% | ✓ | — | Medium |
MMLU
General knowledge across many academic subjects
GPQA
Graduate-level science questions
SWE
Reported SWE-bench Verified issue-resolution rate
AIME
Competition mathematics accuracy
Cost calculator
Estimate monthly API spend using the listed token prices. Discounts, caching, free tiers, and batch pricing are not included.
Tokens per request
1,300
Requests per month
30,000
Cheapest listed option
Llama 4 Scout · $9.81/mo
| Model | Provider | Per request | Monthly |
|---|---|---|---|
| Llama 4 Scoutcheapest | Meta | $0.0003 | $9.81 |
| GPT-4o mini | OpenAI | $0.0006 | $16.65 |
| Llama 4 Maverick | Meta | $0.0008 | $24.45 |
| Codestral 25.01 | Mistral AI | $0.0009 | $26.10 |
| DeepSeek V4-Pro | DeepSeek | $0.0010 | $30.45 |
| Claude Haiku 4.5 | Anthropic | $0.0011 | $33.75 |
| DeepSeek R1 | DeepSeek | $0.0020 | $60.81 |
| o4-mini | OpenAI | $0.0041 | $122.10 |
| Mistral Large 2 | Mistral AI | $0.0058 | $174.00 |
| Mistral Medium 3.5 | Mistral AI | $0.0067 | $202.50 |
| Gemini 3.5 Flash | $0.0080 | $238.50 | |
| Gemini 3.1 Pro | $0.0086 | $258.75 | |
| GPT-4o | OpenAI | $0.0092 | $277.50 |
| Gemini 3.1 Ultra | $0.0132 | $397.50 | |
| Claude Sonnet 4.6 | Anthropic | $0.0135 | $405.00 |
| Grok 4.3 | xAI | $0.0135 | $405.00 |
| GPT-5.5 | OpenAI | $0.0290 | $870.00 |
| o3 | OpenAI | $0.0370 | $1.1k |
| Claude Fable 5 | Anthropic | $0.0450 | $1.4k |
| Claude Opus 4.8 | Anthropic | $0.0675 | $2.0k |