← Back to dashboard

Compare.

Models head to head — pick two or three. Same scoring, same week. Search to swap, add, or remove a model and the matchup recomputes. Hover or tap a stat to see how it shifted.

If quality matters most

Grok 4.6

Best raw competency at 93.2/100. Pay the premium when accuracy compounds.

If cost matters most

Claude Opus 5

-150% cheaper than Grok, 0.1 points lower on competency. The value pick for high-volume workloads.

If you need control

GPT-5.6 Sol

Self-host inside your perimeter — . Lower quality, but no vendor lock-in and no data leaves your network.

SpaceXAIGrok 4.6
AnthropicClaude Opus 5
OpenAIGPT-5.6 Sol
CompetencyWeighted LLM score
93.2/100
flat W21 · rank 1
Best in class
93.1/100
flat W21 · rank 2 · −0.1 vs winner
92/100
flat W21 · rank 3 · −1.2 vs winner
Cost (input)Per 1M tokens
$2.00
60% cheaper than Claude
Cheapest
$5.00
~3× the cheapest in the row
$4.00
~2× the cheapest in the row
Cost (output)Per 1M tokens
$6.00
76% cheaper
Cheapest
$25.00
~4× the cheapest in the row
$20.00
~3× the cheapest in the row
SpeedTokens / second
53 tok/s
TTFT 35.7s
53 tok/s
TTFT 35s
62 tok/s
TTFT 97.6s
Fastest cloud
AccuracyHallucination rate
no data
no data
no data
ContextMax input tokens
— context window
Largest
— context window
— context window
DeploymentWhere it runs
Most flexible
Data residencyCompliance
Most flexible
LicenseWhat you can do
Most permissive

Where each one wins.

Sub-competency scores by business task. Bars highlight the row leader in accent.

Customer Service20% of category weight
Grok 4.6
95
Claude Opus 5
95
GPT-5.6 Sol
93
Content Writing15% weight
Grok 4.6
96
Claude Opus 5
95
GPT-5.6 Sol
94
Research & Analysis16% weight
Grok 4.6
96
Claude Opus 5
96
GPT-5.6 Sol
95
Summarization12% weight
Grok 4.6
97
Claude Opus 5
97
GPT-5.6 Sol
96
Translation7% weight
Grok 4.6
98
Claude Opus 5
98
GPT-5.6 Sol
98
Reasoning16% weight
Grok 4.6
96
Claude Opus 5
95
GPT-5.6 Sol
95
Knowledge14% weight
Grok 4.6
75
Claude Opus 5
77
GPT-5.6 Sol
75

Cost vs quality, charted.

The selected models on the value frontier. Up-and-to-the-left wins.

$0.5$2$5$10$15+10090807060COST $/1M TOK →COMPETENCY /100 ↑Grok 4.6 · 93Claude Opus 5 · 93GPT-5.6 Sol · 92
Filled = cloud / APIOutline = open weightsLower-left = best value

Still picking? We can shorten the list.

Ten questions, one answer. Cost estimates, quality bar, and the eval gates we'd add this sprint.