Find, compare and review AI models

Reviuws tracks 89 AI models across 29 providers with specs, pricing per million tokens, benchmark scores and community reviews.

  • SDXL — stability: Stable Diffusion XL (SDXL) is the definitive open-source standard for text-to-image generation, renowned for its massive community-driven ecosystem. Utilizing a
  • Runway Gen-4 — runway: Runway Gen-4 represents the pinnacle of production-grade AI video generation, offering filmmakers and creators unprecedented control over physics, camera moveme
  • Whisper Large v3 — openai: Whisper Large v3 is OpenAI's state-of-the-art open-source speech recognition model, trained on millions of hours of diverse audio data to deliver industry-leadi
  • text-embedding-3-small — openai: Text-embedding-3-small is OpenAI's highly efficient and cost-effective embedding model designed to convert textual data into numerical vectors. It offers a sign
  • GPT Image 2 — openai: GPT Image 2 is OpenAI's state-of-the-art visual model designed for advanced image generation and editing directly through an API. It allows developers and creat
  • o4-mini — openai: o4-mini delivers strong reasoning performance at lower cost and latency than o3. It supports tool use and is optimized for math and coding. It's designed for hi
  • Gemini 2.5 Pro — google: Gemini 2.5 Pro is Google's flagship model with native multimodality and a 1M token context window, featuring built-in 'thinking' for complex reasoning. It leads
  • Gemini 2.5 Flash-Lite — google: Gemini 2.5 Flash-Lite is optimized for high-throughput, latency-sensitive tasks at the lowest cost in the Gemini 2.5 family. It retains a 1M token context windo
  • text-embedding-004 — google: text-embedding-004 generates vector representations for text used in search, clustering and classification. It's accessible via the Gemini API. It supports task
  • Claude Opus 4.1 — anthropic: Claude Opus 4.1 is Anthropic's top-tier model, excelling at agentic coding, complex reasoning and long-horizon tasks. It supports a 200k context window and visi
  • Claude Haiku 4.5 — anthropic: Claude Haiku 4.5 is optimized for speed and affordability while retaining solid reasoning ability inherited from the Claude 4 family. It supports a 200k context
  • Mistral Large 2 — mistral: Mistral Large 2 is a 123B parameter dense model with strong multilingual, coding and reasoning capabilities and a 128k context window. It's available under a re
  • Codestral 2508 — mistral: Codestral is fine-tuned specifically for code generation, completion and understanding across 80+ programming languages, with a 256k context window. It supports
  • Whisper-1 — openai: Whisper is a general-purpose speech recognition model trained on diverse multilingual audio. It's available open-source and via OpenAI's API. It handles transcr
  • Gemini 2.5 Flash — google: Gemini 2.5 Flash balances speed, cost and reasoning quality with a configurable thinking budget. It supports a 1M token context and full multimodal input. It's
  • Gemma 3 — google: Gemma 3 is a family of open-weight models (1B-27B) with multimodal and multilingual support, derived from Gemini research. It supports a 128k context window. It
  • Claude Sonnet 4.5 — anthropic: Claude Sonnet 4.5 offers strong coding and agentic capabilities at a mid-tier price point, with a 200k context window (1M in beta). It's positioned as Anthropic
  • Llama 3.3 70B — meta: Llama 3.3 70B delivers performance comparable to Llama 3.1 405B at a fraction of the size, focused purely on text tasks. It supports a 128k context window and m
  • Mistral Small 3.2 — mistral: Mistral Small 3.2 is a 24B parameter open-weight model with multimodal support, tuned for instruction following and reduced repetition errors. It offers a 128k
  • DeepSeek-V3.1 — deepseek: DeepSeek-V3.1 is a 671B parameter (37B active) mixture-of-experts model combining fast and thinking modes in one model. It offers strong coding and reasoning at
  • DeepSeek-R1 — deepseek: DeepSeek-R1 is a reasoning-focused model trained with large-scale reinforcement learning, achieving performance comparable to OpenAI's o1 on math and coding ben
  • Qwen2.5-VL-72B — alibaba: Qwen2.5-VL-72B provides advanced visual understanding including document parsing, video comprehension and object grounding. It supports a 128k context window an
  • Grok 4.3 — xai: Cheaper long-context Grok with a 1M-token window. $1.25/$2.50 per 1M tokens under 200k prompt tokens.
  • Qwen3-235B-A22B — alibaba: Qwen3-235B-A22B is a mixture-of-experts model with 235B total/22B active parameters, supporting seamless switching between thinking and non-thinking modes. It s
  • Claude Fable 5 — anthropic: Anthropic's most capable widely released model, built for long-running agents. 1M-token context, 128k max output, always-on adaptive thinking. $10/$50 per 1M to
  • Claude Opus 5 — anthropic: Anthropic's recommended default for complex agentic coding and enterprise work. 1M context, 128k max output, $5/$25 per 1M tokens.
  • Claude Sonnet 5 — anthropic: The best Claude balance of speed and intelligence. 1M context, $3/$15 per 1M tokens list price (introductory $2/$10 through August 31, 2026).
  • Claude Haiku 4.5 — anthropic: Claude Haiku 4.5 is Anthropic's fastest and most cost-effective model, optimized for near-instantaneous response times. It delivers exceptional performance on h
  • GPT-5.6 Sol — openai: GPT-5.6 Sol represents OpenAI's absolute pinnacle of machine intelligence, specifically engineered to tackle ultra-complex reasoning challenges and autonomous a
  • GPT-5.6 Terra — openai: Balances intelligence and cost in the GPT-5.6 family. 1.05M context, $2.50/$15 per 1M tokens.
  • GPT-5.6 Luna — openai: GPT-5.6 tuned for cost-sensitive, high-volume workloads. 1.05M context, $1/$6 per 1M tokens.
  • GPT-5.5 — openai: GPT-5.5 represents OpenAI's next-generation frontier model, specifically optimized for highly complex reasoning, advanced coding tasks, and multi-step analytica
  • GPT-5.4 mini — openai: GPT-5.4 mini is OpenAI's highly efficient and cost-effective model optimized for high-volume automated tasks. It excels in driving specialized coding sub-agents
  • GPT-5.4 nano — openai: GPT-5.4 nano is OpenAI's highly optimized, lightweight model designed for high-throughput, low-latency text processing tasks. Positioned as the most cost-effect
  • GPT-Realtime 2.1 — openai: OpenAI's current realtime speech-to-speech model with reasoning and tool use.
  • Sora 2 — openai: OpenAI's text-to-video model with synchronized audio, billed per second of generated video.
  • Gemini 3.6 Flash — google: Google's most intelligent model built for speed, with strong search grounding. $1.50/$7.50 per 1M tokens; free tier available.
  • Gemini 3.5 Flash — google: Google's Gemini 3.5 Flash is a highly optimized, cost-effective model designed specifically for high-speed, high-volume conversational tasks and multi-turn agen
  • Gemini 3.5 Flash-Lite — google: Google's most cost-efficient GA model for high-volume agentic tasks, translation and data processing. $0.30/$2.50 per 1M tokens.
  • Gemini 3.1 Pro (Preview) — google: Gemini 3.1 Pro is a highly advanced multimodal reasoning model from Google, specifically optimized for complex chat interactions and deep analytical coding. Equ
  • Nano Banana 2 (Gemini 3.1 Flash Image) — google: Gemini 3.1 Flash Image — fast image generation and editing. Roughly $0.067 per 1K image, $0.151 per 4K image.
  • Veo 3.1 — google: Google's latest video model with native audio. Standard $0.40/sec at 720p–1080p; Fast from $0.10/sec.
  • Gemini Embedding 2 — google: Google's first multimodal embedding model — text, image, video, audio and PDFs in one space. $0.20 per 1M text input tokens.
  • Grok 4.5 — xai: xAI's frontier model for coding and agentic work. 500k context, $2/$6 per 1M tokens under 200k prompt tokens ($4/$12 above).
  • Mistral Large 3 — mistral: Mistral Large 3 is Mistral AI's premier European frontier model, engineered to deliver top-tier reasoning, advanced coding capabilities, and highly sophisticate
  • DeepSeek V4 Pro — deepseek: Mixture-of-experts model with 1.6T total and 49B active parameters, hybrid compressed attention, and a 1M-token context window.
  • Qwen3.5-397B-A17B — alibaba: Alibaba's first Qwen3.5 release: a 397B-parameter MoE with 17B active parameters, Apache 2.0 licensed, hosted as Qwen3.5-Plus on Model Studio.
  • Llama 4 Maverick — meta: Meta's natively multimodal MoE with 400B total and 17B active parameters, under the Llama 4 Community License.
  • Llama 4 Scout — meta: Efficient Llama 4 MoE: 109B total / 17B active parameters with an industry-leading 10M-token context window.
  • FLUX.2 [pro] — black-forest-labs: FLUX 2 Pro by Black Forest Labs is a top-tier text-to-image model engineered engineered specifically for elite-level photorealism and precise visual rendering.
  • FLUX.2 [dev] — black-forest-labs: FLUX 2 Dev is an advanced open-weights image generation model developed by Black Forest Labs, designed to deliver state-of-the-art visual quality and prompt adh
  • text-embedding-3-large — openai: Text-embedding-3-large is OpenAI's flagship embedding model, engineered to provide highly accurate vector representations for complex semantic search and retrie
  • Grok 3 — xai: Grok 3 introduced 'Think' mode for extended reasoning and DeepSearch for web-integrated answers, with a 131k context window. It was trained on xAI's Colossus su
  • Command A — cohere: Command A is Cohere's most capable model, optimized for enterprise use cases like RAG, tool use and agents, with a 256k context window. It runs efficiently on j
  • Embed v4 — cohere: Embed v4 generates unified embeddings for text, images and mixed documents (like PDFs with charts), supporting a 128k token context. It's designed for enterpris
  • Jamba 1.6 — ai21: Jamba 1.6 combines Mamba and Transformer architectures in a mixture-of-experts design, offering a 256k context window with efficient long-context inference. It
  • Phi-4 — microsoft: Phi-4 is a 14B parameter dense model trained with a focus on data quality, achieving strong performance on reasoning and math benchmarks relative to its size. I
  • Phi-4-multimodal — microsoft: Phi-4-multimodal is a 5.6B parameter model that unifies text, image and audio understanding in a single small model with a 128k context window. It's released un
  • Nemotron-4 340B — nvidia: Nemotron-4 340B is a dense 340B parameter model optimized to generate high-quality synthetic training data for other LLMs. It's released under the NVIDIA Open M
  • Llama-3.1-Nemotron-Ultra-253B — nvidia: This model is a pruned and post-trained derivative of Llama 3.1 405B, optimized by Nvidia for reasoning tasks with toggleable reasoning mode. It supports a 128k