← Back to dashboard

Compare.

Models head to head — pick two or three. Same scoring, same week. Search to swap, add, or remove a model and the matchup recomputes. Hover or tap a stat to see how it shifted.

If quality matters most

Grok 4

Best raw competency at 64.2/100. Pay the premium when accuracy compounds.

If cost matters most

GPT-5

-367% cheaper than Grok, 5.7 points lower on competency. The value pick for high-volume workloads.

OpenAIGPT-5
Cloud · API
xAIGrok 4
Cloud · API
CompetencyWeighted LLM score
58.5/100
flat W21 · rank 2 · −5.7 vs winner
64.2/100
↑ +0 W21 · rank 42
Best in class
Cost (input)Per 1M tokens
$14.00
~5× the cheapest in the row
$3.00
79% cheaper than GPT-5
Cheapest
Cost (output)Per 1M tokens
$42.00
~3× the cheapest in the row
$15.00
64% cheaper
Cheapest
SpeedTokens / second
76 tok/s
local inference, batch 1
Fastest cloud
0 tok/s
TTFT 0s
AccuracyHallucination rate
2.1%
lowest in category
Most accurate
no data
ContextMax input tokens
~256 pages
~256 pages context window
Largest
~256 pages
~256 pages context window
DeploymentWhere it runs
Cloud only
API key · quick setup
Most flexible
Cloud only
API key · quick setup
Data residencyCompliance
Vendor regions
verify region availability
Most flexible
Vendor regions
verify region availability
LicenseWhat you can do
Proprietary
commercial OK · ToS applies
Most permissive
Proprietary
commercial OK · ToS applies

Where each one wins.

Sub-competency scores by business task. Bars highlight the row leader in accent.

Customer service40% of category weight
GPT-5
96
Grok 4
88
Research & analysis22% weight
GPT-5
93
Grok 4
85
Summarization18% weight
GPT-5
92
Grok 4
84
Code review10% weight
GPT-5
89
Grok 4
81
Document understanding6% weight
GPT-5
86
Grok 4
78

GPT-5

✓ Best for
  • Strong at Customer Service for OpenAI.
  • Strong at Content Writing for OpenAI.
⚠ Watch out for
  • Validate fit for your specific workload before committing.
Open the full profile →

Grok 4

✓ Best for
  • Strong at Customer Service for xAI.
  • Strong at Content Writing for xAI.
⚠ Watch out for
  • Validate fit for your specific workload before committing.
Open the full profile →

Cost vs quality, charted.

The selected models on the value frontier. Up-and-to-the-left wins.

$0.5$2$5$10$15+10090807060COST $/1M TOK →COMPETENCY /100 ↑GPT-5 · 58Grok 4 · 64
Filled = cloud / APIOutline = open weightsLower-left = best value

Still picking? We can shorten the list.

Ten questions, one answer. Cost estimates, quality bar, and the eval gates we'd add this sprint.