Live benchmark explorer

Every recorded Reviuws benchmark run in one interactive view — filter by test, modality and provider, then re-rank the leaderboard by accuracy score, measured latency or input price per million tokens.

Benchmarks covered

Top scoring models

  1. Gemini 3 Pro — google
  2. Veo 3.1 — google
  3. GPT-4o — openai
  4. o3 — openai
  5. Claude Haiku 4.5 — anthropic
  6. DeepSeek-R1 — deepseek
  7. Claude Sonnet 4.5 — anthropic
  8. Gemini 2.5 Flash — google
  9. Claude Opus 4.5 — anthropic
  10. Sora 2 — openai
  11. Grok 4 — xai
  12. Mistral Medium 3 — mistral
  13. GPT-5 — openai
  14. Llama 4 Maverick — meta
  15. Kimi K2 — moonshot ai
  16. FLUX.1.1 [pro] — black forest labs
  17. Midjourney V7 — midjourney
  18. DeepSeek-V3.2 — deepseek
  19. Qwen3-Max — alibaba
  20. GPT-5.5 — openai
  21. Claude Fable 5.1 — anthropic
  22. Claude Opus 5.5 — anthropic
  23. Grok 4.7 — xai
  24. SDXL — stability
  25. Runway Gen-4 — runway