Best Long-Context LLMs 2026: Large-Window AI APIs
Explore long-context LLMs in 2026: 1M+ token context windows, RAG-friendly SKUs, and estimated API pricing. Built for legal, docs, and enterprise RAG in the US, Canada, and Australia.
Long-context LLMs ranked for RAG, docs, and codebases in 2026
Large context reduces chunking pain for legal bundles, multi-file repos, and executive briefs—but token cost scales with what you paste. This view foregrounds window size and fitness for retrieval-heavy stacks while keeping monthly estimates honest for teams in the US, Canada, and Australia planning enterprise rollouts.
Workload & pricing toggles
Same three scenarios as the main AI API calculator: moderate traffic, large RAG-style context, or per-request max tokens with a lower request count.
Include Vision / Image Processing
Off — no image fees in cost estimates for vision-capable models.
Turn On to include image fees.
Use Cached Pricing
Enable to get 50% off input tokens where cached rates apply
Deep Reasoning / Thinking Mode
Model hidden reasoning / extended thinking charged like output tokens when enabled.
Batch Pricing
Enable for 50% off input & output where batch/async pricing applies
Cached / batch est. monthly values only change after the pipeline sets supports_caching or supports_batch in Supabase. The toggles here narrow the table to models whose catalog or provider typically supports those modes.
Magic quadrant (top 15)
X: est. monthly · Y: Long context · Dot: provider color · Hover for rank, model & detailsFull leaderboard
Showing 48 of 444 models.
| Pick | Model | Est. monthly | ROI score | Coding | Reasoning | Speed | Math | Context | Overall |
|---|---|---|---|---|---|---|---|---|---|
| xAI: Grok 4.1 Fast | $13.00 | 60 | 75 | 83 | 90 | 75 | 2.0M | 79Evidence lacks raw benchmarks. Inferred scores based on 'Fast' lightweight tier and agentic focus. Speed rated high (90); coding (75) and logic (80) adjusted lower than flagship models. Native reasoning enabled. | |
| xAI: Grok 4 Fast | $13.00 | 40 | 40 | 43 | 95 | 45 | 2.0M | 43Evidence cites a 43.5 average across HumanEval, GPQA, and MMLU for this 1B-scale Fast model. Mapped to ~40-45 for coding and logic. As a lightweight tier, speed is rated very high. | |
| Auto Router (Beta) | VARIABLE | 77 | 85 | 85 | 70 | 85 | 2.0M | 85Primary source: OpenRouter docs (no raw scores). Inferred frontier-level capabilities (~85) across coding, logic, instruction, and math based on task-aware routing to popular models. Speed estimated at 70. Reasoning explicitly omitted for dynamic routers. | |
| SpaceXAI: Grok 4.20 | $75.00 | 61 | 95 | 91 | 80 | 90 | 2.0M | 92SWE-bench (78.0%) maps to 95 coding. GPQA Diamond (74.5%) maps to 92 logic. As a flagship model, instruction and math align with frontier scores. Speed is high (80) due to being the low-latency, non-reasoning variant. | |
| Auto Router | VARIABLE | 80 | 90 | 90 | 70 | 90 | 2.0M | 90Auto Router optimizes across models for best output. Evidence cites top models reaching 92.3% MMLU. Mapped to 90 across logic, coding, and math to reflect frontier routing capabilities. Vision price defaulted to $0.007 per tier guidelines. | |
| Pareto Code Router | VARIABLE | 77 | 88 | 85 | 70 | 85 | 2.0M | 86OpenRouter docs state this is a router defaulting to High tier coding models based on Artificial Analysis percentiles. Lacking specific raw benchmarks, scores are mapped to ~85 reflecting flagship-level routed performance. Text-only inputs confirmed. | |
| SpaceXAI: Grok 4.20 Multi-Agent | $75.00 | 58 | 85 | 87 | 45 | 85 | 2.0M | 86Evidence lacks explicit Grok benchmarks (listed as 'Not available'). Inferred as a 2026 flagship reasoning model competing with Opus 4.6. Scores estimated for a heavy multi-agent reasoning tier; speed is lower due to 16-agent coordination. | |
| Z.ai: GLM 5.3 | $100.00 | 56 | 78 | 87 | 45 | 85 | 1.3M | 84GLM-5 series scores SWE-bench Verified 77.8% (Coding 78) and GPQA Diamond 86.0% (Logic 86). IFEval is 88.0% (Instruction 88). MMMU inclusion confirms multimodal. As a 744B reasoning flagship, speed is lower (~45). | |
| Z.ai: GLM Flash Latest | $5.50 | 72 | 85 | 87 | 95 | 95 | 1.3M | 89HumanEval 94.2%, GPQA 85.7%, IFEval 88.0%, AIME 95.7%. Mapped directly to 0-100 scale. As a Flash tier model, speed is rated high (95), while coding/logic reflect strong but lightweight capabilities compared to flagship GLM-5. | |
| Z.ai: GLM 5.3 Flash | $5.50 | 63 | 65 | 73 | 90 | 80 | 1.3M | 73Evidence lacks exact GLM 5.3 Flash scores, but GLM-5 flagship hits SWE-bench Verified 77.8% and GPQA 86.0%. As a 'Flash' lightweight tier, scores are adjusted downward (Coding ~65, Logic ~70) prioritizing speed (90). | |
| Llama 4 Scout | $7.00 | 59 | 65 | 73 | 85 | 70 | 1.3M | 70MMMU 69.4% and ChartQA 88.8% map to 70 multimodal. Lacking SWE-bench or GPQA, coding and logic are inferred (65-70) for this 17B-active lightweight MoE tier. Speed is high (85) due to small active parameter count. | |
| Z.ai: GLM Latest | $79.50 | 60 | 92 | 90 | 55 | 88 | 1.3M | 90GLM-5 (latest) scores SWE-bench Verified 77.8% (Coding 92) and GPQA Diamond 86.0% (Logic 92). IFEval 88.0% maps to Instruction 88. Math uses AIME 84.0% (Math 88). Speed reflects a massive 744B MoE flagship architecture. | |
| DeepSeek: DeepSeek V4 Flash 0731 | $4.40 | 72 | 88 | 87 | 92 | 82 | 1.3M | 86SWE-bench Verified (79.0%) and GPQA Diamond (88.1%) map to 88 coding and 89 logic. As a Flash tier (13B active), speed is rated high (92), though native reasoning modes boost its benchmark scores significantly. | |
| DeepSeek V4 Flash Latest | $3.60 | 74 | 85 | 88 | 92 | 88 | 1.3M | 87SWE-bench Verified 79.0% maps to 85 coding. GPQA Diamond 88.1 maps to 90 logic. As a Flash tier model, speed is rated high (92), though its native reasoning modes elevate logic scores near flagship levels. | |
| OpenAI: GPT-5.6 Luna Pro | $20.00 | 66 | 92 | 93 | 45 | 95 | 1.1M | 93Evidence lacks exact GPT-5.6 Luna Pro benchmarks. Inferred scores based on its 'Pro' heavyweight tier and native reasoning mode. Assigned high logic (95) and coding (92), with lower speed (45) typical for extended reasoning models. | |
| OpenAI: GPT-5.6 Luna | $20.00 | 51 | 65 | 68 | 90 | 65 | 1.1M | 66Evidence lacks exact benchmarks for GPT-5.6 Luna. Inferred scores from its 'fast, cost-efficient' lightweight tier description, prioritizing speed (90) while scaling down coding and logic versus flagship models. | |
| OpenAI: GPT-5.6 Terra Pro | $200.00 | 59 | 90 | 91 | 45 | 92 | 1.1M | 91Web digest notes 37 tokens/s (Speed 45). Lacking exact GPT-5.6 Terra Pro benchmarks, inferred flagship 'Pro' tier with native reasoning: Coding 90, Logic 92. Vision price defaulted to $0.007. | |
| OpenAI: GPT-6 Astra | $900.00 | 61 | 98 | 97 | 55 | 99 | 1.1M | 98DeepSWE 74.1% maps to 98 coding. GPQA Diamond 96% maps to 99 logic. FrontierMath 97.6% maps to 99 math. MMMU 74.8% maps to 85 multimodal. Flagship reasoning model speed is estimated at 55. | |
| OpenAI: GPT-5.4 Pro | $3,000.00 | 57 | 95 | 93 | 60 | 92 | 1.1M | 93GPT-5.4 Pro lacks exact SWE-bench/GPQA scores but beats Gemini 3.1 Pro on SWE-Bench Pro and MMMU-Pro. Inferred as a flagship model: Coding ~95, Logic ~92 (GDPval 83%). Multimodal inferred ~85. | |
| OpenAI: GPT-5.5 | $500.00 | 56 | 85 | 89 | 60 | 90 | 1.1M | 88Evidence states GPT-5.5 outperforms GPT-4.1 (SWE-bench Verified 54.6%, GPQA 66.3%). As a frontier model with native real-time reasoning, scores are mapped to elite flagship tiers (85-90+). Speed is typical for heavyweights. | |
| Xiaomi: MiMo-V2.5-Pro | $26.10 | 64 | 95 | 91 | 65 | 88 | 1.1M | 91SWE-bench Verified 78.9% maps to 95 coding. GPQA Diamond 66.7% maps to 92 logic. Flagship tier model with native reasoning and caching support. | |
| Xiaomi: MiMo-V2.5 | $8.40 | 66 | 78 | 86 | 60 | 85 | 1.1M | 84Pro-level agentic performance maps to MiMo-V2-Pro's SWE-bench Verified (78.0%) and GPQA Diamond (87.0%), yielding ~78 Coding and ~87 Logic. Speed reflects 29 tok/s. Multimodal is strong per omnimodal claims. | |
| OpenAI: GPT-5.6 Sol Pro | $180.00 | 60 | 92 | 93 | 45 | 94 | 1.1M | 93No exact benchmarks for GPT-5.6 Sol Pro in evidence. Inferred flagship scores (90+) based on 'Pro' tier and 'reasoning.mode' capabilities. Speed adjusted lower (~45) for dedicated reasoning mode. | |
| OpenAI GPT Latest | $180.00 | 55 | 85 | 86 | 77 | 80 | 1.1M | 84Based on GPT-4.1 data: SWE-bench Verified 54.6% maps to 85 coding, GPQA 66.3% maps to 85 logic, IFEval 87.4% maps to 87 instruction. MMMU 74.8% maps to 80 multimodal. Flagship tier, no lightweight adjustment. | |
| OpenAI: GPT-5.5 Pro | $3,000.00 | 57 | 90 | 94 | 60 | 92 | 1.1M | 92LLM Benchmarks reports 94.8 overall score, mapped to 95 logic. Outperforms in GPQA and MathVista. As a 1T parameter Pro flagship, coding and math are estimated at 90-92. Speed is standard for heavyweights (60). | |
| OpenAI: GPT-6 Astra Pro | $900.00 | 62 | 98 | 99 | 45 | 99 | 1.1M | 99DeepSWE 74.1% maps to 98 coding. GPQA Diamond 96% maps to 99 logic. FrontierMath 97.6% maps to 99 math. MMMU 74.8% maps to 85 multimodal. Flagship reasoning model with native effort controls, yielding lower speed (45). | |
| OpenAI: GPT-5.6 Sol | $180.00 | 59 | 91 | 88 | 60 | 97 | 1.1M | 91Lacking explicit GPT-5.6 Sol benchmarks, inferred flagship scores from top leaderboard entries (SWE-bench ~91%, GSM8K ~97%) and GPT-4.5 GPQA (71.4%). Mapped to 85-97 range for logic/coding/math. Speed set to 60 for heavyweight tier. | |
| OpenAI: GPT-5.4 | $250.00 | 59 | 95 | 91 | 70 | 88 | 1.1M | 91Evidence states GPT-5.4 outperforms GPT-4.1 (SWE-bench Verified 54.6%, GPQA 66.3%). Mapped coding to 95 and logic to 92 for this frontier flagship. Multimodal inferred from native computer-use screenshots and GPT-4.1's MMMU 74.8%. | |
| OpenAI: GPT-5.6 Terra | $200.00 | 53 | 80 | 81 | 80 | 80 | 1.1M | 81Specific GPT-5.6 Terra metrics are proprietary. Inferred scores based on its balanced mid-tier positioning between Sol and Luna. Assigned ~80s across coding, logic, and math, with speed at 80 and multimodal at 75. | |
| Owl Alpha | Free | 66 | 65 | 68 | 85 | 60 | 1.0M | 65No exact scores for Owl Alpha; inferred as a lightweight reasoning model ('fewer parameters', 'designed for speed'). Mapped to mid-tier 0-100 scale (Coding 65, Logic 65) reflecting its agentic focus but smaller size. | |
| Meituan: LongCat 2.0 | $24.00 | 63 | 85 | 85 | 65 | 95 | 1.0M | 88HumanEval 78.5% and GSM8K 92% cited for Flash Chat; as the 1.6T flagship, LongCat 2.0 maps higher (Coding 85, Math 95). Logic and Instruction estimated at 85. Speed 65 for heavy MoE. | |
| Ox Alpha | Free | 79 | 90 | 88 | 85 | 85 | 1.0M | 88Primary source: Web digest (102.5 tok/s). Lacking exact SWE-bench or GPQA percentages for Ox Alpha, capabilities are inferred for a flagship reasoning tier. Speed mapped to 85 based on the cited high throughput. | |
| Meta: Muse Spark 1.3 Contributor | $6.00 | 61 | 59 | 78 | 85 | 75 | 1.0M | 72Terminal-Bench 59.0 maps to 59 coding. CharXiv Reasoning 86.4 maps to 80 logic. MMMU-Pro 80.5% maps to 81 multimodal. As a cost-efficient contributor tier, speed is rated high (85) while capabilities reflect its lightweight status. | |
| Meta: Muse Spark 1.2 | $92.50 | 58 | 88 | 87 | 45 | 85 | 1.0M | 87SWE-bench Verified at 59.0% maps to 88 coding. Logic at 88 reflects 58.4 on Ultimate Human Test and 58% on HLE. As a dedicated reasoning model, speed is estimated lower (45). | |
| MiniMax: MiniMax M3 | $24.00 | 60 | 82 | 82 | 80 | 85 | 1.0M | 83Based on cited evidence: GPQA 54.4% maps to Logic 75, HumanEval 86.9% to Coding 82, IFEval 89.1% to Instruction 89, and MATH 77.4% to Math 85. Native reasoning tokens are supported. | |
| MoonshotAI: Kimi K3 | $270.00 | 61 | 95 | 93 | 65 | 98 | 1.0M | 95SWE-bench Verified at 76.8% maps to 95 coding. GPQA-Diamond at 87.6% maps to 95 logic. IFEval 89.8% maps to 90 instruction. AIME 96.1% maps to 98 math. 2.8T MoE tier dictates ~65 speed. | |
| Tencent: Hy4 preview | $58.37 | 60 | 85 | 88 | 70 | 95 | 1.0M | 89Lacking exact Hy4 scores, mapped as a 770B MoE flagship using cited 2026 frontier baselines: 40-55% SWE-bench Verified (Coding ~85) and 75-90% GPQA (Logic ~85). Configurable reasoning modes confirm native reasoning. | |
| Google: Gemini 3.1 Pro Preview Custom Tools | $200.00 | 62 | 98 | 97 | 65 | 95 | 1.0M | 97SWE-bench Verified at 80.6% maps to 98 coding. GPQA Diamond at 94.3% maps to 98 logic. As a flagship Pro model, it receives high multimodal (95) and math (95) scores, with standard heavyweight speed (65). | |
| Google: Gemini 3.1 Flash Lite | $25.00 | 49 | 25 | 83 | 98 | 66 | 1.0M | 64SWE-bench Verified at 22% maps to 25 coding. GPQA Diamond at 86.9% maps to 87 logic. As a Lite tier, it excels in speed (381 t/s, 98) but trails flagships in coding. | |
| Google: Gemini 2.5 Flash Lite Preview 09-2025 | $8.00 | 52 | 45 | 67 | 95 | 55 | 1.0M | 58Based on GPQA (65.1-70.9%) and LiveCodeBench (64.1-68.8%), logic and coding map to 68 and 45. As a 'Flash Lite' tier, speed is heavily weighted (95), reflecting its ultra-low latency design over flagship-level reasoning. | |
| DeepSeek: DeepSeek V4 Flash Vision Exp | $15.40 | 64 | 88 | 87 | 90 | 85 | 1.0M | 87SWE-bench Verified at 79.0% maps to 88 coding. GPQA Diamond at 88.1 maps to 89 logic. As a Flash tier, speed is rated 90, with high reasoning scores driven by its native 'thinking' mode. | |
| Gemini 2.5 Pro | $150.00 | 57 | 88 | 85 | 50 | 90 | 1.0M | 87Evidence cites MMLU at 81.7% for Gemini 2.5 Pro, mapping to 85 Logic. It leads GPQA and AIME 2025, mapping to 90 Math. Coding maps to 88 based on major improvements over prior versions. Heavyweight tier. | |
| Google: Lyria 3 Clip Preview | Free | 31 | 0 | 5 | 50 | 0 | 1.0M | 3Lyria 3 is a specialized music generation model lacking standard LLM benchmarks (SWE-bench, GPQA). Assigned 0 for coding/logic/math. Speed mapped to 50 from 38 tok/s. Multimodal scored 85 for native image-to-audio generation. | |
| Xiaomi: MiMo-V2-Pro | $70.00 | 61 | 95 | 90 | 60 | 88 | 1.0M | 91SWE-bench Verified at 78.0% maps to 95 coding. GPQA Diamond at 87.0% maps to 95 logic. IFBench at 68.8% maps to 85 instruction. As a 1T+ flagship, speed is lower (60). | |
| Google Gemini Pro Latest | $200.00 | 52 | 84 | 73 | 60 | 88 | 1.0M | 79HumanEval 84.1% maps to 84 coding. GPQA 59.1% maps to 65 logic. MATH 86.5% maps to 88 math. MMMU 65.9% maps to 75 multimodal. Flagship tier model with native reasoning capabilities. | |
| MoonshotAI Kimi Latest | $229.50 | 55 | 88 | 85 | 70 | 82 | 1.0M | 85SWE-bench Verified at 65.8% maps to 88 coding. GPQA-Diamond at 48.1% maps to 85 logic. MATH at 70.2% maps to 82 math. Flagship 1T MoE tier; native reasoning and multimodal capabilities confirmed. | |
| Google: Gemini 2.5 Flash | $37.00 | 56 | 70 | 78 | 92 | 85 | 1.0M | 78GPQA Diamond 78.3% (Logic 80), LiveCodeBench 63.5% (Coding 70), MMMU 76.7%. As a Flash-tier model, it excels in speed (93 tok/s) and math (AIME 78%), but trails Pro in heavy coding. | |
| Google: Gemini 2.5 Flash Lite | $8.00 | 53 | 45 | 68 | 95 | 65 | 1.0M | 61Evidence lacks exact Flash-Lite scores but notes it underperforms Flash (GPQA 78.3%, MMMU 76.7%). As a Lite tier, scores are adjusted downward (Logic 65, Coding 45). Speed is heavily weighted (95) due to 68 tok/s and ultra-low latency. |
Need a shareable artifact?
Get a print-ready PDF of your results and a CSV spreadsheet. Tap the button, then enter your work email. We use it to build your files and start the download—and to email you a copy if the site owner enabled that.
AI ROI Leaderboard & Discovery by LeadsCalc
PDF Breakdown
Receive a comprehensive native vector PDF of this leaderboard: your workload, filters, top rankings, and a table snapshot (sorted: Long context).
By submitting, you agree to our Privacy Policy and Terms.
Whitelabel Context Leaderboard
for your site
Embed the interactive long context view on your own domain — whitelabel branding, lead capture, and the same workload sliders your prospects already use on LeadsCalc.