Best Long-Context LLMs 2026: Large-Window AI APIs
Explore long-context LLMs in 2026: 1M+ token context windows, RAG-friendly SKUs, and estimated API pricing. Built for legal, docs, and enterprise RAG in the US, Canada, and Australia.
Long-context LLMs ranked for RAG, docs, and codebases in 2026
Large context reduces chunking pain for legal bundles, multi-file repos, and executive briefs—but token cost scales with what you paste. This view foregrounds window size and fitness for retrieval-heavy stacks while keeping monthly estimates honest for teams in the US, Canada, and Australia planning enterprise rollouts.
Workload & pricing toggles
Same three scenarios as the main AI API calculator: moderate traffic, large RAG-style context, or per-request max tokens with a lower request count.
Include Vision / Image Processing
Off — no image fees in cost estimates for vision-capable models.
Turn On to include image fees.
Use Cached Pricing
Enable to get 50% off input tokens where cached rates apply
Deep Reasoning / Thinking Mode
Model hidden reasoning / extended thinking charged like output tokens when enabled.
Batch Pricing
Enable for 50% off input & output where batch/async pricing applies
Cached / batch est. monthly values only change after the pipeline sets supports_caching or supports_batch in Supabase. The toggles here narrow the table to models whose catalog or provider typically supports those modes.
Magic quadrant (top 15)
X: est. monthly · Y: Long context · Dot: provider color · Hover for rank, model & detailsFull leaderboard
Showing 48 of 401 models.
| Pick | Model | Est. monthly | ROI score | Coding | Reasoning | Speed | Math | Context | Overall |
|---|---|---|---|---|---|---|---|---|---|
| xAI: Grok 4.1 Fast | $13.00 | 61 | 75 | 83 | 90 | 75 | 2.0M | 79Evidence lacks raw benchmarks. Inferred scores based on 'Fast' lightweight tier and agentic focus. Speed rated high (90); coding (75) and logic (80) adjusted lower than flagship models. Native reasoning enabled. | |
| xAI: Grok 4 Fast | $13.00 | 41 | 40 | 43 | 95 | 45 | 2.0M | 43Evidence cites a 43.5 average across HumanEval, GPQA, and MMLU for this 1B-scale Fast model. Mapped to ~40-45 for coding and logic. As a lightweight tier, speed is rated very high. | |
| Pareto Code Router | VARIABLE | 78 | 88 | 85 | 70 | 85 | 2.0M | 86OpenRouter docs state this is a router defaulting to High tier coding models based on Artificial Analysis percentiles. Lacking specific raw benchmarks, scores are mapped to ~85 reflecting flagship-level routed performance. Text-only inputs confirmed. | |
| Auto Router | VARIABLE | 81 | 90 | 90 | 70 | 90 | 2.0M | 90Auto Router optimizes across models for best output. Evidence cites top models reaching 92.3% MMLU. Mapped to 90 across logic, coding, and math to reflect frontier routing capabilities. Vision price defaulted to $0.007 per tier guidelines. | |
| xAI: Grok 4.20 | $75.00 | 62 | 95 | 91 | 80 | 90 | 2.0M | 92SWE-bench (78.0%) maps to 95 coding. GPQA Diamond (74.5%) maps to 92 logic. As a flagship model, instruction and math align with frontier scores. Speed is high (80) due to being the low-latency, non-reasoning variant. | |
| Auto Router (Beta) | VARIABLE | 78 | 85 | 85 | 70 | 85 | 2.0M | 85Primary source: OpenRouter docs (no raw scores). Inferred frontier-level capabilities (~85) across coding, logic, instruction, and math based on task-aware routing to popular models. Speed estimated at 70. Reasoning explicitly omitted for dynamic routers. | |
| xAI: Grok 4.20 Multi-Agent | $75.00 | 59 | 85 | 87 | 45 | 85 | 2.0M | 86Evidence lacks explicit Grok benchmarks (listed as 'Not available'). Inferred as a 2026 flagship reasoning model competing with Opus 4.6. Scores estimated for a heavy multi-agent reasoning tier; speed is lower due to 16-agent coordination. | |
| Llama 4 Scout | $7.00 | 60 | 65 | 73 | 85 | 70 | 1.3M | 70MMMU 69.4% and ChartQA 88.8% map to 70 multimodal. Lacking SWE-bench or GPQA, coding and logic are inferred (65-70) for this 17B-active lightweight MoE tier. Speed is high (85) due to small active parameter count. | |
| OpenAI: GPT-5.6 Luna Pro | $100.00 | 62 | 92 | 93 | 45 | 95 | 1.1M | 93Evidence lacks exact GPT-5.6 Luna Pro benchmarks. Inferred scores based on its 'Pro' heavyweight tier and native reasoning mode. Assigned high logic (95) and coding (92), with lower speed (45) typical for extended reasoning models. | |
| OpenAI: GPT-5.6 Luna | $100.00 | 47 | 65 | 68 | 90 | 65 | 1.1M | 66Evidence lacks exact benchmarks for GPT-5.6 Luna. Inferred scores from its 'fast, cost-efficient' lightweight tier description, prioritizing speed (90) while scaling down coding and logic versus flagship models. | |
| OpenAI GPT Latest | $500.00 | 55 | 85 | 86 | 77 | 80 | 1.1M | 84Based on GPT-4.1 data: SWE-bench Verified 54.6% maps to 85 coding, GPQA 66.3% maps to 85 logic, IFEval 87.4% maps to 87 instruction. MMMU 74.8% maps to 80 multimodal. Flagship tier, no lightweight adjustment. | |
| OpenAI: GPT-5.6 Sol | $500.00 | 59 | 91 | 88 | 60 | 97 | 1.1M | 91Lacking explicit GPT-5.6 Sol benchmarks, inferred flagship scores from top leaderboard entries (SWE-bench ~91%, GSM8K ~97%) and GPT-4.5 GPQA (71.4%). Mapped to 85-97 range for logic/coding/math. Speed set to 60 for heavyweight tier. | |
| OpenAI: GPT-5.6 Terra Pro | $250.00 | 60 | 90 | 91 | 45 | 92 | 1.1M | 91Web digest notes 37 tokens/s (Speed 45). Lacking exact GPT-5.6 Terra Pro benchmarks, inferred flagship 'Pro' tier with native reasoning: Coding 90, Logic 92. Vision price defaulted to $0.007. | |
| OpenAI: GPT-5.4 Pro | $3,000.00 | 58 | 95 | 93 | 60 | 92 | 1.1M | 93GPT-5.4 Pro lacks exact SWE-bench/GPQA scores but beats Gemini 3.1 Pro on SWE-Bench Pro and MMMU-Pro. Inferred as a flagship model: Coding ~95, Logic ~92 (GDPval 83%). Multimodal inferred ~85. | |
| OpenAI: GPT-5.5 | $500.00 | 57 | 85 | 89 | 60 | 90 | 1.1M | 88Evidence states GPT-5.5 outperforms GPT-4.1 (SWE-bench Verified 54.6%, GPQA 66.3%). As a frontier model with native real-time reasoning, scores are mapped to elite flagship tiers (85-90+). Speed is typical for heavyweights. | |
| OpenAI: GPT-5.4 | $250.00 | 60 | 95 | 91 | 70 | 88 | 1.1M | 91Evidence states GPT-5.4 outperforms GPT-4.1 (SWE-bench Verified 54.6%, GPQA 66.3%). Mapped coding to 95 and logic to 92 for this frontier flagship. Multimodal inferred from native computer-use screenshots and GPT-4.1's MMMU 74.8%. | |
| OpenAI: GPT-5.6 Terra | $250.00 | 54 | 80 | 81 | 80 | 80 | 1.1M | 81Specific GPT-5.6 Terra metrics are proprietary. Inferred scores based on its balanced mid-tier positioning between Sol and Luna. Assigned ~80s across coding, logic, and math, with speed at 80 and multimodal at 75. | |
| Xiaomi: MiMo-V2.5-Pro | $26.10 | 65 | 95 | 91 | 65 | 88 | 1.1M | 91SWE-bench Verified 78.9% maps to 95 coding. GPQA Diamond 66.7% maps to 92 logic. Flagship tier model with native reasoning and caching support. | |
| Xiaomi: MiMo-V2.5 | $8.40 | 67 | 78 | 86 | 60 | 85 | 1.1M | 84Pro-level agentic performance maps to MiMo-V2-Pro's SWE-bench Verified (78.0%) and GPQA Diamond (87.0%), yielding ~78 Coding and ~87 Logic. Speed reflects 29 tok/s. Multimodal is strong per omnimodal claims. | |
| OpenAI: GPT-5.6 Sol Pro | $500.00 | 60 | 92 | 93 | 45 | 94 | 1.1M | 93No exact benchmarks for GPT-5.6 Sol Pro in evidence. Inferred flagship scores (90+) based on 'Pro' tier and 'reasoning.mode' capabilities. Speed adjusted lower (~45) for dedicated reasoning mode. | |
| OpenAI: GPT-5.5 Pro | $3,000.00 | 58 | 90 | 94 | 60 | 92 | 1.1M | 92LLM Benchmarks reports 94.8 overall score, mapped to 95 logic. Outperforms in GPQA and MathVista. As a 1T parameter Pro flagship, coding and math are estimated at 90-92. Speed is standard for heavyweights (60). | |
| Meituan: LongCat 2.0 | $24.00 | 64 | 85 | 85 | 65 | 95 | 1.0M | 88HumanEval 78.5% and GSM8K 92% cited for Flash Chat; as the 1.6T flagship, LongCat 2.0 maps higher (Coding 85, Math 95). Logic and Instruction estimated at 85. Speed 65 for heavy MoE. | |
| Owl Alpha | Free | 67 | 65 | 68 | 85 | 60 | 1.0M | 65No exact scores for Owl Alpha; inferred as a lightweight reasoning model ('fewer parameters', 'designed for speed'). Mapped to mid-tier 0-100 scale (Coding 65, Logic 65) reflecting its agentic focus but smaller size. | |
| Google: Gemini 3.6 Flash | $135.00 | 48 | 74 | 63 | 95 | 78 | 1.0M | 69Based on Flash tier evidence: HumanEval 74.3% (Coding 74), GPQA 51.0% (Logic 55), MATH 77.9% (Math 78). As a lightweight model, it scores lower than flagships but excels in speed (95). Multimodal supported (MMMU 62.3%). | |
| MiniMax: MiniMax M3 | $24.00 | 61 | 82 | 82 | 80 | 85 | 1.0M | 83Based on cited evidence: GPQA 54.4% maps to Logic 75, HumanEval 86.9% to Coding 82, IFEval 89.1% to Instruction 89, and MATH 77.4% to Math 85. Native reasoning tokens are supported. | |
| MoonshotAI: Kimi K3 | $270.00 | 62 | 95 | 93 | 65 | 98 | 1.0M | 95SWE-bench Verified at 76.8% maps to 95 coding. GPQA-Diamond at 87.6% maps to 95 logic. IFEval 89.8% maps to 90 instruction. AIME 96.1% maps to 98 math. 2.8T MoE tier dictates ~65 speed. | |
| Thinking Machines: Inkling | $80.50 | 58 | 80 | 85 | 65 | 85 | 1.0M | 84Evidence lacks exact Inkling scores. Inferred from 975B MoE flagship tier: mapped to ~80-85 for coding/logic/math based on general frontier benchmark references (MATH 75-90%). Vision supported; default frontier price used. | |
| Z.ai: GLM 5.2 | $51.90 | 62 | 92 | 89 | 55 | 90 | 1.0M | 90GLM-5 series scores SWE-bench Verified 77.8% (mapped to 92 Coding) and GPQA Diamond 86.0% (mapped to 90 Logic). IFEval 88.0% maps to 88 Instruction. As a flagship reasoning model, it receives high capability scores. | |
| Google Gemini Pro Latest | $200.00 | 53 | 84 | 73 | 60 | 88 | 1.0M | 79HumanEval 84.1% maps to 84 coding. GPQA 59.1% maps to 65 logic. MATH 86.5% maps to 88 math. MMMU 65.9% maps to 75 multimodal. Flagship tier model with native reasoning capabilities. | |
| DeepSeek: DeepSeek V4 Pro | $26.10 | 66 | 98 | 91 | 70 | 88 | 1.0M | 92SWE-bench Verified at 81.0% maps to 98 coding. GPQA Diamond at 66.3% maps to 92 logic. Flagship MoE tier justifies high scores; speed is standard for large MoE. | |
| Gemini 2.5 Pro | $150.00 | 58 | 88 | 85 | 50 | 90 | 1.0M | 87Evidence cites MMLU at 81.7% for Gemini 2.5 Pro, mapping to 85 Logic. It leads GPQA and AIME 2025, mapping to 90 Math. Coding maps to 88 based on major improvements over prior versions. Heavyweight tier. | |
| Google: Gemini 3.5 Flash Lite | $37.00 | 54 | 70 | 73 | 95 | 75 | 1.0M | 73Lacking exact 3.5 Flash-Lite scores, inferred from Gemini 3 Flash (SWE-bench Verified 78.0%, GPQA 90.4%). Mapped to 0-100 scale with strict penalties for the 'Lite' tier, reducing coding and logic while maximizing speed. | |
| DeepSeek: DeepSeek V4 Flash | $5.63 | 66 | 68 | 79 | 90 | 85 | 1.0M | 78V4 flagship claims 80%+ SWE-bench; Flash tier (13B active) lacks explicit scores but is inferred ~68 for coding. Logic and Math scaled down for Flash efficiency. Speed rated 90 for fast inference design. | |
| Google: Gemini 3.1 Flash Lite | $25.00 | 50 | 25 | 83 | 98 | 66 | 1.0M | 64SWE-bench Verified at 22% maps to 25 coding. GPQA Diamond at 86.9% maps to 87 logic. As a Lite tier, it excels in speed (381 t/s, 98) but trails flagships in coding. | |
| Google: Gemini 2.5 Flash Lite Preview 09-2025 | $8.00 | 52 | 45 | 67 | 95 | 55 | 1.0M | 58Based on GPQA (65.1-70.9%) and LiveCodeBench (64.1-68.8%), logic and coding map to 68 and 45. As a 'Flash Lite' tier, speed is heavily weighted (95), reflecting its ultra-low latency design over flagship-level reasoning. | |
| Google: Lyria 3 Clip Preview | Free | 31 | 0 | 5 | 50 | 0 | 1.0M | 3Lyria 3 is a specialized music generation model lacking standard LLM benchmarks (SWE-bench, GPQA). Assigned 0 for coding/logic/math. Speed mapped to 50 from 38 tok/s. Multimodal scored 85 for native image-to-audio generation. | |
| Xiaomi: MiMo-V2-Pro | $70.00 | 62 | 95 | 90 | 60 | 88 | 1.0M | 91SWE-bench Verified at 78.0% maps to 95 coding. GPQA Diamond at 87.0% maps to 95 logic. IFBench at 68.8% maps to 85 instruction. As a 1T+ flagship, speed is lower (60). | |
| MoonshotAI Kimi Latest | $270.00 | 56 | 88 | 85 | 70 | 82 | 1.0M | 85SWE-bench Verified at 65.8% maps to 88 coding. GPQA-Diamond at 48.1% maps to 85 logic. MATH at 70.2% maps to 82 math. Flagship 1T MoE tier; native reasoning and multimodal capabilities confirmed. | |
| Google: Gemini 2.5 Flash | $37.00 | 56 | 70 | 78 | 92 | 85 | 1.0M | 78GPQA Diamond 78.3% (Logic 80), LiveCodeBench 63.5% (Coding 70), MMMU 76.7%. As a Flash-tier model, it excels in speed (93 tok/s) and math (AIME 78%), but trails Pro in heavy coding. | |
| Meta: Muse Spark 1.1 | $92.50 | 62 | 92 | 91 | 45 | 92 | 1.0M | 92SWE-bench Verified at 59.0% maps to 92 coding. MMMU-Pro at 80.5% maps to 95 multimodal. Ultimate Human Test (58.4) maps to 92 logic. Flagship reasoning model; speed adjusted to 45. | |
| Google: Gemini 2.5 Flash Lite | $8.00 | 54 | 45 | 68 | 95 | 65 | 1.0M | 61Evidence lacks exact Flash-Lite scores but notes it underperforms Flash (GPQA 78.3%, MMMU 76.7%). As a Lite tier, scores are adjusted downward (Logic 65, Coding 45). Speed is heavily weighted (95) due to 68 tok/s and ultra-low latency. | |
| Google: Gemini 3.5 Flash | $150.00 | 57 | 85 | 88 | 95 | 75 | 1.0M | 84SWE-bench Verified 78.0% and GPQA Diamond 90.4% map to 85 and 90. Despite being a lightweight Flash tier (speed 95), explicit evidence dictates high capability scores, though typically lower than Pro. | |
| Poolside: Laguna S 2.1 | $6.00 | 64 | 82 | 74 | 88 | 70 | 1.0M | 75OpenRouter cites 70.2% on Terminal-Bench 2.1, mapped to 82 coding. As an 8B active parameter MoE (Small tier), it achieves high speed but moderate logic and math compared to heavyweight flagships. | |
| Google: Gemini 3.1 Pro Preview | $200.00 | 58 | 88 | 89 | 65 | 88 | 1.0M | 88Evidence lacks raw benchmarks. Inferred scores based on 'Pro' flagship tier and 'frontier reasoning' claims, assigning high 80s for coding, logic, and math. Speed estimated at 65 for a heavyweight. | |
| Google: Gemini 2.5 Pro Preview 06-05 | $150.00 | 63 | 96 | 93 | 45 | 96 | 1.0M | 95SWE-bench (59.6%) maps to 96 coding. GPQA (86.4%) maps to 96 logic. AIME (88.0%) maps to 96 math. MMMU (82.0%) maps to 90 multimodal. Flagship tier model with native reasoning; speed adjusted for thinking overhead. | |
| Google: Gemini 3.1 Pro Preview Custom Tools | $200.00 | 63 | 98 | 97 | 65 | 95 | 1.0M | 97SWE-bench Verified at 80.6% maps to 98 coding. GPQA Diamond at 94.3% maps to 98 logic. As a flagship Pro model, it receives high multimodal (95) and math (95) scores, with standard heavyweight speed (65). | |
| Google Gemini Flash Latest | $135.00 | 48 | 70 | 65 | 95 | 75 | 1.0M | 69Evidence cites HumanEval 74.3% (Coding ~70), GPQA 51.0% (Logic ~60), and MMMU 62.3% (Multimodal ~65). As a lightweight Flash tier, scores are adjusted lower than Pro flagships, while Speed is rated high (~95) for its class. | |
| Google: Gemini 3 Flash Preview | $50.00 | 61 | 88 | 89 | 95 | 85 | 1.0M | 88GPQA Diamond 90.4% maps to Logic 92; SWE-bench 78% maps to Coding 88. As a Flash-tier model, Speed is rated very high (95). Multimodal inferred at 80 due to extensive video/audio/image support. |
Need a shareable artifact?
Get a print-ready PDF of your results and a CSV spreadsheet. Tap the button, then enter your work email. We use it to build your files and start the download—and to email you a copy if the site owner enabled that.
AI ROI Leaderboard & Discovery by LeadsCalc
PDF Breakdown
Receive a comprehensive native vector PDF of this leaderboard: your workload, filters, top rankings, and a table snapshot (sorted: Long context).
By submitting, you agree to our Privacy Policy and Terms.
Whitelabel Context Leaderboard
for your site
Embed the interactive long context view on your own domain — whitelabel branding, lead capture, and the same workload sliders your prospects already use on LeadsCalc.