Models head to head — pick two or three. Same scoring, same week. Search to swap, add, or remove a model and the matchup recomputes. Hover or tap a stat to see how it shifted.
If quality matters most
XGrok 4
Best raw competency at 64.2/100. Pay the premium when accuracy compounds.
If cost matters most
AQwen 3 Max
60% cheaper than Grok, 16.2 points lower on competency. The value pick for high-volume workloads.
xAIXGrok 4
Cloud · API
AlibabaAQwen 3 Max
Open weights
CompetencyWeighted LLM score
64.2/100
↑ +0 W21 · rank 42
Best in class
48/100
flat W21 · rank 89 · −16.2 vs winner
Cost (input)Per 1M tokens
$3.00
~3× the cheapest in the row
$1.20
60% cheaper than Grok
Cheapest
Cost (output)Per 1M tokens
$15.00
~3× the cheapest in the row
$6.00
60% cheaper
Cheapest
SpeedTokens / second
0 tok/s
TTFT 0s
Fastest cloud
0 tok/s
TTFT 0s
AccuracyHallucination rate
—
no data
—
no data
ContextMax input tokens
~256 pages
~256 pages context window
Largest
~256 pages
~256 pages context window
DeploymentWhere it runs
Cloud only
API key · quick setup
Cloud or local
Self-host on your GPUs
Most flexible
Data residencyCompliance
Vendor regions
verify region availability
Anywhere
runs inside your perimeter
Most flexible
LicenseWhat you can do
Proprietary
commercial OK · ToS applies
Open weights
commercial OK
Most permissive
Where each one wins.
Sub-competency scores by business task. Bars highlight the row leader in accent.
Customer service40% of category weight
XGrok 4
88
AQwen 3 Max
81
Research & analysis22% weight
XGrok 4
85
AQwen 3 Max
78
Summarization18% weight
XGrok 4
84
AQwen 3 Max
77
Code review10% weight
XGrok 4
81
AQwen 3 Max
74
Document understanding6% weight
XGrok 4
78
AQwen 3 Max
71
XGrok 4
✓ Best for
Strong at Customer Service for xAI.
Strong at Content Writing for xAI.
⚠ Watch out for
Validate fit for your specific workload before committing.