Models head to head — pick two or three. Same scoring, same week. Search to swap, add, or remove a model and the matchup recomputes. Hover or tap a stat to see how it shifted.
If quality matters most
XGrok 4
Best raw competency at 64.2/100. Pay the premium when accuracy compounds.
If cost matters most
OGPT-5
-367% cheaper than Grok, 5.7 points lower on competency. The value pick for high-volume workloads.
OpenAIOGPT-5
Cloud · API
xAIXGrok 4
Cloud · API
CompetencyWeighted LLM score
58.5/100
flat W21 · rank 2 · −5.7 vs winner
64.2/100
↑ +0 W21 · rank 42
Best in class
Cost (input)Per 1M tokens
$14.00
~5× the cheapest in the row
$3.00
79% cheaper than GPT-5
Cheapest
Cost (output)Per 1M tokens
$42.00
~3× the cheapest in the row
$15.00
64% cheaper
Cheapest
SpeedTokens / second
76 tok/s
local inference, batch 1
Fastest cloud
0 tok/s
TTFT 0s
AccuracyHallucination rate
2.1%
lowest in category
Most accurate
—
no data
ContextMax input tokens
~256 pages
~256 pages context window
Largest
~256 pages
~256 pages context window
DeploymentWhere it runs
Cloud only
API key · quick setup
Most flexible
Cloud only
API key · quick setup
Data residencyCompliance
Vendor regions
verify region availability
Most flexible
Vendor regions
verify region availability
LicenseWhat you can do
Proprietary
commercial OK · ToS applies
Most permissive
Proprietary
commercial OK · ToS applies
Where each one wins.
Sub-competency scores by business task. Bars highlight the row leader in accent.
Customer service40% of category weight
OGPT-5
96
XGrok 4
88
Research & analysis22% weight
OGPT-5
93
XGrok 4
85
Summarization18% weight
OGPT-5
92
XGrok 4
84
Code review10% weight
OGPT-5
89
XGrok 4
81
Document understanding6% weight
OGPT-5
86
XGrok 4
78
OGPT-5
✓ Best for
Strong at Customer Service for OpenAI.
Strong at Content Writing for OpenAI.
⚠ Watch out for
Validate fit for your specific workload before committing.