Every model.
Every category.
One chart.

Pick a category. Pick the business task. We re-weight the underlying benchmarks and redraw competency vs cost. Bubble size = speed. Filled = cloud, outline = open weights.

Which language model should I use for my text-based business task, and what will it cost?

Use case

All use cases · 42 models shown

value sweet spot$0.20$0.50$1.0$2.0$5.0$10$200102030405060708090100COST ($ PER 1M TOKENS, BLENDED) → cheaper is leftCOMPETENCY (0–100)Claude Sonnet 5Command R+DeepSeek V4 Flash 0731GPT-5.3 CodexGPT-5.4 miniGPT-5.4 nanoGPT-5.5GPT-5.6 LunaGemma 3 27BHy3KAT Coder Pro V2Kimi K2 ThinkingKimi K2.7 CodeMiMo-V2.5MiniMax-M3Mistral Large 3Muse Spark 1.2Nex-N2-ProQwen3.6 35B A3BQwen3.8 MaxSolar Pro 4Step 3.7 Flasho3-proClaude Fable 5Claude Opus 5GLM-5.3Grok 4.6Kimi K3
Cloud (API) Local / open weights Both (cloud + local) Frontier this weekBubble size = speed 
Scores are weighted averages of underlying business-relevant benchmarks. How we score →UPDATED · FRI 28 AUG 2026

Cloud or local?

Same question, two answers. The cloud leader for raw quality vs the open-weights leader for self-hosting.

Cloud

via API

Better if you want top-tier quality with zero infrastructure, fast setup, and elastic scale — and sending requests to a managed API is acceptable for your data.

Claude 4.5 SonnetCATEGORY LEADER · COMPETENCY 96/100
Cost
$7.50 / 1M tok
Speed
78 tok/s · 0.6s TTFT
Accuracy
1.8% hallucination
Context
~400 pages
To deploy
API key, 5 min

Local

open weights

Better if data must stay in-house for compliance or privacy, you run steady high volume where self-hosting is cheaper, or you need full control and predictable cost.

Llama 4.1 70BOPEN-WEIGHTS LEADER · COMPETENCY 72/100
Cost
$0.75 / 1M tok*
Speed
64 tok/s on H100
Accuracy
3.6% hallucination
Context
~256 pages
To deploy
~$2/hr GPU

Not sure which model fits for you?

A short, no-jargon questionnaire that maps your task, volume, and accuracy needs to shortlist three models with monthly cost estimates.

Find your model ↗

Top stories of the week.

● AI Wins

Ox Alpha: The Anonymous Model That Just Blew Up Enterprise AI's Moat Argument

An anonymous model called Ox Alpha appeared on OpenRouter and quietly matched or beat GPT-4-class performance, with no brand, no press release, and no pedigree. If a mystery model can do that, the "you need us specifically" argument that frontier labs have been selling to enterprise buyers just got a lot harder to make.

Read the story →
● AI Tools

The Ox Alpha Moment: When an Anonymous Model Matches GPT-4 Class Output, Your Vendor Lock-In Is Already a Liability

Z AI's Ox Alpha surfaced on public benchmarks with GPT-4-class performance and a price tag that undercuts the major labs, and nobody saw it coming. If a model nobody had heard of last quarter can compete at the frontier, the case for paying premium rates to a single incumbent just got a lot harder to make.

Read the story →
● AI Tools

Replit Makes Model Routing the Default: The Moat Is No Longer the Model

Replit has made intelligent model routing its default behavior, automatically selecting the best AI model for each task rather than locking users to a single provider. This is less a product update and more a signal: the platform layer is quietly becoming more valuable than the model underneath it.

Read the story →

The weekly verdict.

A short, plain-English read on the week in AI — the decisions worth your time, not the hype.

Tools. Workflows. Custom.
Three paths, plainly explained.

"Which model?" is usually the third question. The first two are how much you're willing to build, and how much of your own data you'd put behind it. Here's the spectrum from off-the-shelf to fully bespoke.

PREMADE TOOLS

Buy the app,
use it Monday.

ChatGPT for the team, Cursor for the devs, Midjourney for the design lead. Someone else picked the model, wrapped it in a UI, and runs the infrastructure. You pay per seat. You can be productive this afternoon.

WORKFLOW · CLOUD · LOCAL · HYBRID

Build the workflow,
pick your model.

You’re embedding a model into a process — support routing, document review, content pipelines. Call a cloud API for raw quality, self-host open weights for control and cost, or do both. Most real systems end up hybrid.

CUSTOM TRAINED MODEL

Train it on your data,
only if you must.

Most teams don’t need this. RAG + a frontier model handles 80% of the cases people think need a custom model. If your domain is genuinely esoteric (legal, medical, niche industrial) and you have labeled data, fine-tuning a 7B–14B open model is the move.

We don't do
average.
You shouldn't
either.

We run AI strategy and implementation engagements for teams that want a sharper edge and less hedge. Engagements start with a 15-minute fit call — short, honest, no pitch deck.

Not every team is a fit. If we're not the right firm for what you're trying to do, we'll tell you in the first 10 minutes — and point you toward someone who is.

One email,
every week.

The week's verdict. News worth your time, and a chart you'll forward to your team.