Best text AI models
44 text models compared on specs, pricing and community reviews.
- o4-mini — openai: o4-mini delivers strong reasoning performance at lower cost and latency than o3. It supports tool use and is optimized for math and coding. It's designed for hi
- Gemini 2.5 Pro — google: Gemini 2.5 Pro is Google's flagship model with native multimodality and a 1M token context window, featuring built-in 'thinking' for complex reasoning. It leads
- Gemini 2.5 Flash-Lite — google: Gemini 2.5 Flash-Lite is optimized for high-throughput, latency-sensitive tasks at the lowest cost in the Gemini 2.5 family. It retains a 1M token context windo
- Claude Opus 4.1 — anthropic: Claude Opus 4.1 is Anthropic's top-tier model, excelling at agentic coding, complex reasoning and long-horizon tasks. It supports a 200k context window and visi
- Claude Haiku 4.5 — anthropic: Claude Haiku 4.5 is optimized for speed and affordability while retaining solid reasoning ability inherited from the Claude 4 family. It supports a 200k context
- Mistral Large 2 — mistral: Mistral Large 2 is a 123B parameter dense model with strong multilingual, coding and reasoning capabilities and a 128k context window. It's available under a re
- Codestral 2508 — mistral: Codestral is fine-tuned specifically for code generation, completion and understanding across 80+ programming languages, with a 256k context window. It supports
- Gemini 2.5 Flash — google: Gemini 2.5 Flash balances speed, cost and reasoning quality with a configurable thinking budget. It supports a 1M token context and full multimodal input. It's
- Gemma 3 — google: Gemma 3 is a family of open-weight models (1B-27B) with multimodal and multilingual support, derived from Gemini research. It supports a 128k context window. It
- Claude Sonnet 4.5 — anthropic: Claude Sonnet 4.5 offers strong coding and agentic capabilities at a mid-tier price point, with a 200k context window (1M in beta). It's positioned as Anthropic
- Llama 3.3 70B — meta: Llama 3.3 70B delivers performance comparable to Llama 3.1 405B at a fraction of the size, focused purely on text tasks. It supports a 128k context window and m
- Mistral Small 3.2 — mistral: Mistral Small 3.2 is a 24B parameter open-weight model with multimodal support, tuned for instruction following and reduced repetition errors. It offers a 128k
- DeepSeek-V3.1 — deepseek: DeepSeek-V3.1 is a 671B parameter (37B active) mixture-of-experts model combining fast and thinking modes in one model. It offers strong coding and reasoning at
- DeepSeek-R1 — deepseek: DeepSeek-R1 is a reasoning-focused model trained with large-scale reinforcement learning, achieving performance comparable to OpenAI's o1 on math and coding ben
- Qwen2.5-VL-72B — alibaba: Qwen2.5-VL-72B provides advanced visual understanding including document parsing, video comprehension and object grounding. It supports a 128k context window an
- Qwen3-235B-A22B — alibaba: Qwen3-235B-A22B is a mixture-of-experts model with 235B total/22B active parameters, supporting seamless switching between thinking and non-thinking modes. It s
- Grok 3 — xai: Grok 3 introduced 'Think' mode for extended reasoning and DeepSearch for web-integrated answers, with a 131k context window. It was trained on xAI's Colossus su
- Command A — cohere: Command A is Cohere's most capable model, optimized for enterprise use cases like RAG, tool use and agents, with a 256k context window. It runs efficiently on j
- Jamba 1.6 — ai21: Jamba 1.6 combines Mamba and Transformer architectures in a mixture-of-experts design, offering a 256k context window with efficient long-context inference. It
- Phi-4 — microsoft: Phi-4 is a 14B parameter dense model trained with a focus on data quality, achieving strong performance on reasoning and math benchmarks relative to its size. I
- Phi-4-multimodal — microsoft: Phi-4-multimodal is a 5.6B parameter model that unifies text, image and audio understanding in a single small model with a 128k context window. It's released un
- Nemotron-4 340B — nvidia: Nemotron-4 340B is a dense 340B parameter model optimized to generate high-quality synthetic training data for other LLMs. It's released under the NVIDIA Open M
- Llama-3.1-Nemotron-Ultra-253B — nvidia: This model is a pruned and post-trained derivative of Llama 3.1 405B, optimized by Nvidia for reasoning tasks with toggleable reasoning mode. It supports a 128k
- Amazon Nova Pro — amazon: Amazon Nova Pro is a multimodal model supporting text, image and video understanding with a 300k token context window, available exclusively via AWS Bedrock. It
- Amazon Nova Micro — amazon: Amazon Nova Micro is a text-only model optimized for the lowest latency and cost within the Nova family, with a 128k context window. It's available via AWS Bedr
- Kimi K2 — moonshot ai: Kimi K2 is a 1T parameter (32B active) mixture-of-experts model trained for agentic tool-use and coding, released with open weights under a modified MIT license
- GLM-4.5 — zhipu ai: GLM-4.5 is a 355B parameter (32B active) mixture-of-experts model designed for agentic, reasoning and coding tasks with a hybrid thinking mode. It's open-weight
- MiniMax M1 — minimax: MiniMax M1 is a 456B parameter (45.9B active) hybrid-attention MoE reasoning model supporting up to 1M tokens of context, trained with efficient large-scale rei
- ERNIE 4.5 — baidu: ERNIE 4.5 is a family of mixture-of-experts models (up to 424B parameters) supporting text and multimodal understanding, released with open weights for several
- Sonar Pro — perplexity: Sonar Pro is a real-time web-search-grounded model built on top of open-weight LLMs, providing cited answers with a 200k context window. It's optimized for comp
- Granite 3.3 8B — ibm: Granite 3.3 8B is IBM's open-weight model tuned for enterprise use cases like RAG, function calling and fill-in-the-middle code, with a 128k context window. It'
- OLMo 2 32B — allen institute for ai: OLMo 2 32B is the largest model in Allen Institute's fully open OLMo family, released with training data, code, and checkpoints for full reproducibility. It's c
- Gemma 2 9B — google: Gemma 2 9B is a dense open-weight model trained with knowledge distillation from larger models, offering strong performance for its size. It supports an 8k cont
- Cohere Rerank 3.5 — cohere: Rerank 3.5 improves search relevance by reordering candidate documents using reasoning and better multilingual understanding. It supports 100+ languages and a 4
- Claude 3.5 Sonnet — anthropic: Claude 3.5 Sonnet introduced computer use (agentic screen control) and improved coding performance over Claude 3 Opus at a lower price. It supports a 200k conte
- Grok 2 — xai: Grok 2 improved reasoning, coding and multilingual capabilities over Grok 1.5, and introduced image generation through a FLUX.1 partnership. It's available on X
- Reka Flash 3 — reka ai: Reka Flash 3 is a 21B parameter multimodal model with reasoning capabilities, trained via reinforcement learning and released under Apache 2.0. It supports a 12
- Pixtral Large — mistral: Pixtral Large is a 124B parameter model combining a 1B vision encoder with Mistral Large 2's text backbone, supporting a 128k context window. It excels at docum
- Aya Expanse 32B — cohere: Aya Expanse 32B is an open-weight model covering 23 languages, developed by Cohere's research lab to close multilingual performance gaps in open models. It supp
- GPT-4o mini — openai: GPT-4o mini is a cost-efficient version of GPT-4o designed for high-volume tasks. It supports text and vision inputs with a 128k context window. It offers stron
- GPT-4.1 mini — openai: GPT-4.1 mini balances cost and capability, offering a 1M token context window at a fraction of GPT-4.1's price. It maintains strong coding and reasoning perform
- GPT-4o — openai: GPT-4o is OpenAI's natively multimodal model supporting text, image and audio inputs with fast, low-latency responses. It powers ChatGPT and the API with a 128k
- GPT-4.1 — openai: GPT-4.1 offers a 1M token context window with major improvements in coding and instruction following over GPT-4o. It's available in standard, mini and nano size
- o3 — openai: o3 is a reasoning-focused model that uses extended chain-of-thought to solve complex math, science and coding problems. It supports tool use during reasoning. I