Claude 4.5 Sonnet keeps the LLM crown for a fifth week runningDeepSeek V3.1 cuts API pricing and the market reprices overnightYOLO26 lands as the new edge-vision default for retailMistral Large 3 ships with EU-only data residency — a procurement resetFLUX.2 brings open-weights image generation closer to Midjourney qualityVeo 3 doubles its clip length, but the editing workflow is still the bottleneckQwen 3 VL goes long-context — and the multimodal eval landscape shiftsElevenLabs v3 ships twelve new languages, with caveats for production useXGBoost still beats every transformer your team will deploy this quarterThe hidden cost of frontier-tier models when you actually run them at scaleClaude 4.5 Sonnet keeps the LLM crown for a fifth week runningDeepSeek V3.1 cuts API pricing and the market reprices overnightYOLO26 lands as the new edge-vision default for retailMistral Large 3 ships with EU-only data residency — a procurement resetFLUX.2 brings open-weights image generation closer to Midjourney qualityVeo 3 doubles its clip length, but the editing workflow is still the bottleneckQwen 3 VL goes long-context — and the multimodal eval landscape shiftsElevenLabs v3 ships twelve new languages, with caveats for production useXGBoost still beats every transformer your team will deploy this quarterThe hidden cost of frontier-tier models when you actually run them at scale

Large Language Models, ranked.

Text in, text out. Chat, write, reason.

Which language model should I use for my text-based business task, and what will it cost?

Use case

All use cases · 43 models shown

value sweet spot$0.20$0.50$1.0$2.0$5.0$10$200102030405060708090100COST ($ PER 1M TOKENS, BLENDED) → cheaper is leftCOMPETENCY (0–100)Claude Sonnet 5Command R+DeepSeek V4 Flash 0731DeepSeek V4 Pro 0813GPT-5.3 CodexGPT-5.4 miniGPT-5.4 nanoGPT-5.5GPT-5.5 InstantGPT-5.6 LunaGemma 3 27BGrok 4.3Hy3KAT Coder Pro V2Kimi K2 ThinkingKimi K2.7 CodeMiMo-V2.5MiniMax-M3Mistral Large 3Muse Spark 1.2Nex-N2-ProQwen3.6 35B A3BQwen3.8 27BSolar Pro 4Step 3.7 Flasho3-proClaude Fable 5Claude Opus 5Gemini 3.7 FlashGrok 4.6Kimi K3Qwen3.8 2.4T A95B
Cloud (API) Local / open weights Both (cloud + local) Frontier this weekBubble size = speed 
Scores are weighted averages of underlying business-relevant benchmarks. How we score →UPDATED · TUE 25 AUG 2026

Cloud or local?

Same question, two answers. The cloud leader for raw quality vs the open-weights leader for self-hosting.

Cloud

via API

Better if you want top-tier quality with zero infrastructure, fast setup, and elastic scale — and sending requests to a managed API is acceptable for your data.

Claude 4.5 SonnetCATEGORY LEADER · COMPETENCY 96/100
Cost
$7.50 / 1M tok
Speed
78 tok/s · 0.6s TTFT
Accuracy
1.8% hallucination
Context
~400 pages
To deploy
API key, 5 min

Local

open weights

Better if data must stay in-house for compliance or privacy, you run steady high volume where self-hosting is cheaper, or you need full control and predictable cost.

Llama 4.1 70BOPEN-WEIGHTS LEADER · COMPETENCY 72/100
Cost
$0.75 / 1M tok*
Speed
64 tok/s on H100
Accuracy
3.6% hallucination
Context
~256 pages
To deploy
~$2/hr GPU

Not sure which LLM model fits for you?

A short, no-jargon questionnaire that maps your task, volume, and accuracy needs to shortlist three models with monthly cost estimates.

Find your model ↗