Best code AI models

52 code models compared on specs, pricing and community reviews.

  • o4-mini — openai: o4-mini delivers strong reasoning performance at lower cost and latency than o3. It supports tool use and is optimized for math and coding. It's designed for hi
  • Gemini 2.5 Pro — google: Gemini 2.5 Pro is Google's flagship model with native multimodality and a 1M token context window, featuring built-in 'thinking' for complex reasoning. It leads
  • Gemini 2.5 Flash-Lite — google: Gemini 2.5 Flash-Lite is optimized for high-throughput, latency-sensitive tasks at the lowest cost in the Gemini 2.5 family. It retains a 1M token context windo
  • Claude Opus 4.1 — anthropic: Claude Opus 4.1 is Anthropic's top-tier model, excelling at agentic coding, complex reasoning and long-horizon tasks. It supports a 200k context window and visi
  • Mistral Large 2 — mistral: Mistral Large 2 is a 123B parameter dense model with strong multilingual, coding and reasoning capabilities and a 128k context window. It's available under a re
  • Codestral 2508 — mistral: Codestral is fine-tuned specifically for code generation, completion and understanding across 80+ programming languages, with a 256k context window. It supports
  • Gemini 2.5 Flash — google: Gemini 2.5 Flash balances speed, cost and reasoning quality with a configurable thinking budget. It supports a 1M token context and full multimodal input. It's
  • Claude Sonnet 4.5 — anthropic: Claude Sonnet 4.5 offers strong coding and agentic capabilities at a mid-tier price point, with a 200k context window (1M in beta). It's positioned as Anthropic
  • Mistral Small 3.2 — mistral: Mistral Small 3.2 is a 24B parameter open-weight model with multimodal support, tuned for instruction following and reduced repetition errors. It offers a 128k
  • DeepSeek-V3.1 — deepseek: DeepSeek-V3.1 is a 671B parameter (37B active) mixture-of-experts model combining fast and thinking modes in one model. It offers strong coding and reasoning at
  • DeepSeek-R1 — deepseek: DeepSeek-R1 is a reasoning-focused model trained with large-scale reinforcement learning, achieving performance comparable to OpenAI's o1 on math and coding ben
  • Qwen3-235B-A22B — alibaba: Qwen3-235B-A22B is a mixture-of-experts model with 235B total/22B active parameters, supporting seamless switching between thinking and non-thinking modes. It s
  • Qwen 3 235B — alibaba: Alibaba's Qwen 3 235B is an exceptionally powerful open-weight model optimized for elite multilingual understanding and advanced programming tasks. Built on a m
  • GPT-5.5 — openai: GPT-5.5 represents OpenAI's next-generation frontier model, specifically optimized for highly complex reasoning, advanced coding tasks, and multi-step analytica
  • Claude Sonnet 5 — anthropic: Claude Sonnet 5 by Anthropic delivers an exceptional balance of speed and high-tier intelligence, specifically optimized for advanced chat and complex coding ta
  • Claude Sonnet 4.5 — anthropic: Claude 4.5 Sonnet is Anthropic's premier mid-tier model designed to deliver elite-level reasoning, coding, and comprehension capabilities. It excels at deeply u
  • Grok 4.5 — xai: Grok 4.5 is xAI's frontier model engineered specifically for advanced coding, complex reasoning, and agentic workflows. Featuring a massive 500k token context w
  • GPT-5.6 Terra — openai: GPT-5.6 Terra is a balanced powerhouse within OpenAI's latest model family, engineered to deliver a cost-effective blend of advanced intelligence and efficiency
  • Claude Fable 5 — anthropic: Claude Fable 5 is Anthropic's premier model designed to power highly sophisticated, autonomous, long-running agents. It features a massive 1M-token context wind
  • Claude Opus 5 — anthropic: Claude Opus 5 is Anthropic's flagship model designed specifically for complex agentic coding and heavy enterprise workloads. Featuring an expansive 1-million-to
  • DeepSeek V4 Pro — deepseek: DeepSeek V4 Pro is an advanced mixture-of-experts model engineered for high-performance chat and coding tasks, leveraging 1.6 trillion total parameters with onl
  • GPT-5.6 Sol — openai: GPT-5.6 Sol represents OpenAI's absolute pinnacle of machine intelligence, specifically engineered to tackle ultra-complex reasoning challenges and autonomous a
  • Gemini 3.1 Pro (Preview) — google: Gemini 3.1 Pro is a highly advanced multimodal reasoning model from Google, specifically optimized for complex chat interactions and deep analytical coding. Equ
  • Mistral Large 3 — mistral: Mistral Large 3 is Mistral AI's premier European frontier model, engineered to deliver top-tier reasoning, advanced coding capabilities, and highly sophisticate
  • Qwen3.5-397B-A17B — alibaba: Alibaba's Qwen3.5-397B-A17B is the debut model of the Qwen3.5 family, leveraging a massive Mixture of Experts architecture with 397 billion total and 17 billion
  • Llama 4 Maverick — meta: Llama 4 Maverick is a state-of-the-art natively multimodal Mixture-of-Experts (MoE) model developed by Meta, engineered to excel in high-performance chat and co
  • Gemini 3.6 Flash — google: Gemini 3.6 Flash is Google's premier speed-optimized model, engineered to deliver rapid responses without sacrificing deep intelligence or coding capabilities.
  • GPT-5 — openai: OpenAI's flagship unified model that combines fast responses with built-in deep reasoning, replacing the separate GPT-4o and o-series split. It set new state-of
  • Grok 3 — xai: Grok 3 introduced 'Think' mode for extended reasoning and DeepSearch for web-integrated answers, with a 131k context window. It was trained on xAI's Colossus su
  • Command A — cohere: Command A is Cohere's most capable model, optimized for enterprise use cases like RAG, tool use and agents, with a 256k context window. It runs efficiently on j
  • Jamba 1.6 — ai21: Jamba 1.6 combines Mamba and Transformer architectures in a mixture-of-experts design, offering a 256k context window with efficient long-context inference. It
  • Phi-4 — microsoft: Phi-4 is a 14B parameter dense model trained with a focus on data quality, achieving strong performance on reasoning and math benchmarks relative to its size. I
  • Nemotron-4 340B — nvidia: Nemotron-4 340B is a dense 340B parameter model optimized to generate high-quality synthetic training data for other LLMs. It's released under the NVIDIA Open M
  • Llama-3.1-Nemotron-Ultra-253B — nvidia: This model is a pruned and post-trained derivative of Llama 3.1 405B, optimized by Nvidia for reasoning tasks with toggleable reasoning mode. It supports a 128k
  • Kimi K2 — moonshot ai: Kimi K2 is a 1T parameter (32B active) mixture-of-experts model trained for agentic tool-use and coding, released with open weights under a modified MIT license
  • GLM-4.5 — zhipu ai: GLM-4.5 is a 355B parameter (32B active) mixture-of-experts model designed for agentic, reasoning and coding tasks with a hybrid thinking mode. It's open-weight
  • MiniMax M1 — minimax: MiniMax M1 is a 456B parameter (45.9B active) hybrid-attention MoE reasoning model supporting up to 1M tokens of context, trained with efficient large-scale rei
  • ERNIE 4.5 — baidu: ERNIE 4.5 is a family of mixture-of-experts models (up to 424B parameters) supporting text and multimodal understanding, released with open weights for several
  • Granite 3.3 8B — ibm: Granite 3.3 8B is IBM's open-weight model tuned for enterprise use cases like RAG, function calling and fill-in-the-middle code, with a 128k context window. It'
  • GPT-5.1-Codex-Max — openai: OpenAI's frontier agentic coding model, built on an updated reasoning core and the first trained to work coherently across multiple context windows through comp
  • Claude Opus 4.5 — anthropic: Claude 4.5 Opus is Anthropic's flagship model designed to tackle the most demanding cognitive tasks, offering unmatched depth in reasoning and precision. It is
  • Llama 4 405B — meta: Llama 4 405B represents Meta's frontier-class open-weights model, offering state-of-the-art general reasoning, coding, and chat capabilities. Designed to compet
  • Claude 3.5 Sonnet — anthropic: Claude 3.5 Sonnet introduced computer use (agentic screen control) and improved coding performance over Claude 3 Opus at a lower price. It supports a 200k conte
  • Gemini 3 Pro — google: Google's most capable model at launch, with state-of-the-art reasoning and deep multimodal understanding across text, image, audio and video. It powers the Gemi
  • DeepSeek R2 — deepseek: DeepSeek R2 is a highly efficient, reasoning-focused open-weights model designed specifically to excel in complex mathematical synthesis and advanced coding tas
  • Qwen3-Max — alibaba: Alibaba's largest Qwen release, a trillion-parameter-scale mixture-of-experts model with a production thinking mode for coding, reasoning and agentic tasks. It
  • GPT-4.1 mini — openai: GPT-4.1 mini balances cost and capability, offering a 1M token context window at a fraction of GPT-4.1's price. It maintains strong coding and reasoning perform
  • DeepSeek-V3.2 — deepseek: Open-weight successor to V3.2-Exp that introduces DeepSeek Sparse Attention for efficient long-context reasoning and agentic work. A high-compute Speciale varia
  • GPT-4.1 — openai: GPT-4.1 offers a 1M token context window with major improvements in coding and instruction following over GPT-4o. It's available in standard, mini and nano size
  • o3 — openai: o3 is a reasoning-focused model that uses extended chain-of-thought to solve complex math, science and coding problems. It supports tool use during reasoning. I
  • GLM-4.6 — zhipu ai: Zhipu's upgrade to GLM-4.5, expanding context from 128k to 200k tokens with better agentic, reasoning and coding behaviour. It has become a widely adopted open
  • Mistral Medium 3 — mistral: An enterprise-focused multimodal model that delivers frontier-level coding, reasoning and vision quality at a fraction of flagship pricing. Mistral positions it