Models head to head — pick two or three. Same scoring, same week. Search to swap, add, or remove a model and the matchup recomputes. Hover or tap a stat to see how it shifted.
If quality matters most
AQwen 3 VL
Best raw competency at 35/100 and the lowest hallucination rate in the set. Pay the premium when accuracy compounds.
If cost matters most
OMiniCPM-V 4
64% cheaper than Qwen, 0 points lower on competency. The value pick for high-volume workloads.
AlibabaAQwen 3 VL
★ FrontierOpen weights
OpenBMBOMiniCPM-V 4
Open weights
CompetencyWeighted MULTIMODAL score
35/100
flat W21 · rank 4
Best in class
35/100
flat W21 · rank 6 · −0 vs winner
Cost (input)Per 1M tokens
$1.40
0% cheaper than Qwen
Cheapest
$0.50*
self-hosted — GPU rental, not per-token
Cost (output)Per 1M tokens
$4.20
0% cheaper
Cheapest
$1.50*
self-hosted — GPU rental, not per-token
SpeedTokens / second
76 tok/s
local inference, batch 1
88 tok/s
local inference, batch 1
Fastest
AccuracyHallucination rate
3.4%
lowest in category
Most accurate
4.2%
+0.8 pp vs winner
ContextMax input tokens
~256 pages
~256 pages context window
Largest
~128 pages
~128 pages context window
DeploymentWhere it runs
Cloud or local
Self-host on your GPUs
Most flexible
Local only
Self-host on your GPUs
Data residencyCompliance
Anywhere
runs inside your perimeter
Anywhere
runs inside your perimeter
Most flexible
LicenseWhat you can do
Open weights
commercial OK
Most permissive
Open weights
commercial OK
Where each one wins.
Sub-competency scores by business task. Bars highlight the row leader in accent.
Document Understanding30% of category weight
AQwen 3 VL
35
OMiniCPM-V 4
35
Chart & Data25% weight
AQwen 3 VL
35
OMiniCPM-V 4
35
Visual QA25% weight
AQwen 3 VL
35
OMiniCPM-V 4
35
Image + Text Reasoning20% weight
AQwen 3 VL
35
OMiniCPM-V 4
35
AQwen 3 VL
✓ Best for
Strong at Document Understanding for Alibaba.
Strong at Chart & Data for Alibaba.
⚠ Watch out for
Validate fit for your specific workload before committing.