Long-context LLMs (200k+ tokens)

19 long-context llms (200k+ tokens) in the Reviuws directory, with specs, pricing and community reviews.

  • Grok 4.3 — xai: Cheaper long-context Grok with a 1M-token window. $1.25/$2.50 per 1M tokens under 200k prompt tokens.
  • Claude Fable 5 — anthropic: Anthropic's most capable widely released model, built for long-running agents. 1M-token context, 128k max output, always-on adaptive thinking. $10/$50 per 1M to
  • Claude Opus 5 — anthropic: Anthropic's recommended default for complex agentic coding and enterprise work. 1M context, 128k max output, $5/$25 per 1M tokens.
  • Claude Sonnet 5 — anthropic: The best Claude balance of speed and intelligence. 1M context, $3/$15 per 1M tokens list price (introductory $2/$10 through August 31, 2026).
  • Claude Haiku 4.5 — anthropic: Claude Haiku 4.5 is Anthropic's fastest and most cost-effective model, optimized for near-instantaneous response times. It delivers exceptional performance on h
  • GPT-5.6 Sol — openai: GPT-5.6 Sol represents OpenAI's absolute pinnacle of machine intelligence, specifically engineered to tackle ultra-complex reasoning challenges and autonomous a
  • GPT-5.6 Terra — openai: Balances intelligence and cost in the GPT-5.6 family. 1.05M context, $2.50/$15 per 1M tokens.
  • GPT-5.6 Luna — openai: GPT-5.6 tuned for cost-sensitive, high-volume workloads. 1.05M context, $1/$6 per 1M tokens.
  • GPT-5.5 — openai: GPT-5.5 represents OpenAI's next-generation frontier model, specifically optimized for highly complex reasoning, advanced coding tasks, and multi-step analytica
  • GPT-5.4 mini — openai: GPT-5.4 mini is OpenAI's highly efficient and cost-effective model optimized for high-volume automated tasks. It excels in driving specialized coding sub-agents
  • GPT-5.4 nano — openai: GPT-5.4 nano is OpenAI's highly optimized, lightweight model designed for high-throughput, low-latency text processing tasks. Positioned as the most cost-effect
  • Gemini 3.6 Flash — google: Google's most intelligent model built for speed, with strong search grounding. $1.50/$7.50 per 1M tokens; free tier available.
  • Gemini 3.5 Flash — google: Google's Gemini 3.5 Flash is a highly optimized, cost-effective model designed specifically for high-speed, high-volume conversational tasks and multi-turn agen
  • Gemini 3.5 Flash-Lite — google: Google's most cost-efficient GA model for high-volume agentic tasks, translation and data processing. $0.30/$2.50 per 1M tokens.
  • Gemini 3.1 Pro (Preview) — google: Gemini 3.1 Pro is a highly advanced multimodal reasoning model from Google, specifically optimized for complex chat interactions and deep analytical coding. Equ
  • Grok 4.5 — xai: xAI's frontier model for coding and agentic work. 500k context, $2/$6 per 1M tokens under 200k prompt tokens ($4/$12 above).
  • DeepSeek V4 Pro — deepseek: Mixture-of-experts model with 1.6T total and 49B active parameters, hybrid compressed attention, and a 1M-token context window.
  • Llama 4 Maverick — meta: Meta's natively multimodal MoE with 400B total and 17B active parameters, under the Llama 4 Community License.
  • Llama 4 Scout — meta: Efficient Llama 4 MoE: 109B total / 17B active parameters with an industry-leading 10M-token context window.