Claude 4.5 Sonnet keeps the LLM crown for a fifth week runningDeepSeek V3.1 cuts API pricing and the market reprices overnightYOLO26 lands as the new edge-vision default for retailMistral Large 3 ships with EU-only data residency — a procurement resetFLUX.2 brings open-weights image generation closer to Midjourney qualityVeo 3 doubles its clip length, but the editing workflow is still the bottleneckQwen 3 VL goes long-context — and the multimodal eval landscape shiftsElevenLabs v3 ships twelve new languages, with caveats for production useXGBoost still beats every transformer your team will deploy this quarterThe hidden cost of frontier-tier models when you actually run them at scaleClaude 4.5 Sonnet keeps the LLM crown for a fifth week runningDeepSeek V3.1 cuts API pricing and the market reprices overnightYOLO26 lands as the new edge-vision default for retailMistral Large 3 ships with EU-only data residency — a procurement resetFLUX.2 brings open-weights image generation closer to Midjourney qualityVeo 3 doubles its clip length, but the editing workflow is still the bottleneckQwen 3 VL goes long-context — and the multimodal eval landscape shiftsElevenLabs v3 ships twelve new languages, with caveats for production useXGBoost still beats every transformer your team will deploy this quarterThe hidden cost of frontier-tier models when you actually run them at scale

Multimodal (Vision + Language), ranked.

Models that read documents, charts, images.

Which AI understands my documents, images, and charts?

Use case

All use cases · 2 models shown

Coverage expanding. We track 2 models in Multimodal so far, and add more as reliable, business-relevant data appears. The chart still shows what we have today.
value sweet spot$0.50$1.0$2.0$5.0$10$200102030405060708090100COST ($ PER 1M TOKENS, BLENDED) → cheaper is leftMULTIMODAL CAPABILITY (0–100)MiniCPM-V 4Qwen 3 VL
Cloud (API) Local / open weights Both (cloud + local) Frontier this weekBubble size = speed 
Scores are weighted averages of underlying business-relevant benchmarks. How we score →UPDATED · TUE 25 AUG 2026

Cloud or local?

Same question, two answers. The cloud leader for raw quality vs the open-weights leader for self-hosting.

Cloud

via API

Better if you want top-tier quality with zero infrastructure, fast setup, and elastic scale — and sending requests to a managed API is acceptable for your data.

Gemini 2.5 ProCATEGORY LEADER · COMPETENCY 94/100
Cost
$6 / 1M tok
Speed
86 tok/s
Accuracy
2.4%
Context
~1.5M tok
To deploy
API key

Local

open weights

Better if data must stay in-house for compliance or privacy, you run steady high volume where self-hosting is cheaper, or you need full control and predictable cost.

Qwen 3 VLOPEN-WEIGHTS LEADER · COMPETENCY 80/100
Cost
$1.40 / 1M tok*
Speed
76 tok/s
Accuracy
3.4%
Context
128k tok
To deploy
~$2/hr GPU

Not sure which Multimodal model fits for you?

A short, no-jargon questionnaire that maps your task, volume, and accuracy needs to shortlist three models with monthly cost estimates.

Find your model ↗