Live benchmark explorer
Every recorded Reviuws benchmark run in one interactive view — filter by test, modality and provider, then re-rank the leaderboard by accuracy score, measured latency or input price per million tokens.
Benchmarks covered
- Reasoning
- Coding
- Instruction Following
- Factuality
- MMLU-Pro
- GPQA Diamond
- Humanity's Last Exam
- AIME 2024
- SWE-bench Verified
- LiveCodeBench
- Terminal-Bench
- HumanEval
- MMMU
- MathVista
- IFEval
- MGSM
- LMArena (Chatbot Arena) Elo
- ARC-AGI-2
Top scoring models
- Gemini 3 Pro — google
- Veo 3.1 — google
- GPT-4o — openai
- o3 — openai
- Claude Haiku 4.5 — anthropic
- DeepSeek-R1 — deepseek
- Claude Sonnet 4.5 — anthropic
- Gemini 2.5 Flash — google
- Claude Opus 4.5 — anthropic
- Sora 2 — openai
- Grok 4 — xai
- Mistral Medium 3 — mistral
- GPT-5 — openai
- Llama 4 Maverick — meta
- Kimi K2 — moonshot ai
- FLUX.1.1 [pro] — black forest labs
- Midjourney V7 — midjourney
- DeepSeek-V3.2 — deepseek
- Qwen3-Max — alibaba
- GPT-5.5 — openai
- Claude Fable 5.1 — anthropic
- Claude Opus 5.5 — anthropic
- Grok 4.7 — xai
- SDXL — stability
- Runway Gen-4 — runway