Every model.
Every category.
One chart.

Pick a category. Pick the business task. We re-weight the underlying benchmarks and redraw competency vs cost. Bubble size = speed. Filled = cloud, outline = open weights.

Which language model should I use for my text-based business task, and what will it cost?

Use case

All use cases · 43 models shown

value sweet spot$0.20$0.50$1.0$2.0$5.0$10$200102030405060708090100COST ($ PER 1M TOKENS, BLENDED) → cheaper is leftCOMPETENCY (0–100)Claude Sonnet 5Command R+DeepSeek V4 Flash 0731DeepSeek V4 Pro 0813GPT-5.3 CodexGPT-5.4 miniGPT-5.4 nanoGPT-5.5GPT-5.5 InstantGPT-5.6 LunaGemma 3 27BGrok 4.3Hy3KAT Coder Pro V2Kimi K2 ThinkingKimi K2.7 CodeMiMo-V2.5MiniMax-M3Mistral Large 3Muse Spark 1.2Nex-N2-ProQwen3.6 35B A3BQwen3.8 27BSolar Pro 4Step 3.7 Flasho3-proClaude Fable 5Claude Opus 5Gemini 3.7 FlashGrok 4.6Kimi K3Qwen3.8 2.4T A95B
Cloud (API) Local / open weights Both (cloud + local) Frontier this weekBubble size = speed 
Scores are weighted averages of underlying business-relevant benchmarks. How we score →UPDATED · TUE 25 AUG 2026

Cloud or local?

Same question, two answers. The cloud leader for raw quality vs the open-weights leader for self-hosting.

Cloud

via API

Better if you want top-tier quality with zero infrastructure, fast setup, and elastic scale — and sending requests to a managed API is acceptable for your data.

Claude 4.5 SonnetCATEGORY LEADER · COMPETENCY 96/100
Cost
$7.50 / 1M tok
Speed
78 tok/s · 0.6s TTFT
Accuracy
1.8% hallucination
Context
~400 pages
To deploy
API key, 5 min

Local

open weights

Better if data must stay in-house for compliance or privacy, you run steady high volume where self-hosting is cheaper, or you need full control and predictable cost.

Llama 4.1 70BOPEN-WEIGHTS LEADER · COMPETENCY 72/100
Cost
$0.75 / 1M tok*
Speed
64 tok/s on H100
Accuracy
3.6% hallucination
Context
~256 pages
To deploy
~$2/hr GPU

Not sure which model fits for you?

A short, no-jargon questionnaire that maps your task, volume, and accuracy needs to shortlist three models with monthly cost estimates.

Find your model ↗

Top stories of the week.

● Pricing

Inference Costs Are Falling Fast. Are You Still Paying Last Quarter's Prices?

AI inference costs have been falling sharply as Together AI, Groq, Fireworks, and Cerebras wage an aggressive price war. If you haven't audited your API spend since last quarter, you're almost certainly overpaying.

Read the story →
● AI Tools

The API Bill Is Coming Due: Why Smart Enterprises Are Running LLMs Locally

As frontier API prices compound with token volumes, the total cost of ownership math for local LLM deployment via Ollama and llama.cpp is quietly flipping in favor of on-premises for high-volume enterprise workloads. This is not a hobbyist trend, it is a procurement conversation.

Read the story →
● AI Tools

What a Stripe–OpenRouter Deal Would Mean: Model Routing as the Next Payment Rail

Speculation is growing that Stripe could target the model-routing layer, the infrastructure sitting between applications and AI models, as its next strategic acquisition. For AI platform vendors and operators, the thesis is worth taking seriously regardless of whether a deal materialises.

Read the story →

The weekly verdict.

A short, plain-English read on the week in AI — the decisions worth your time, not the hype.

Tools. Workflows. Custom.
Three paths, plainly explained.

"Which model?" is usually the third question. The first two are how much you're willing to build, and how much of your own data you'd put behind it. Here's the spectrum from off-the-shelf to fully bespoke.

PREMADE TOOLS

Buy the app,
use it Monday.

ChatGPT for the team, Cursor for the devs, Midjourney for the design lead. Someone else picked the model, wrapped it in a UI, and runs the infrastructure. You pay per seat. You can be productive this afternoon.

WORKFLOW · CLOUD · LOCAL · HYBRID

Build the workflow,
pick your model.

You’re embedding a model into a process — support routing, document review, content pipelines. Call a cloud API for raw quality, self-host open weights for control and cost, or do both. Most real systems end up hybrid.

CUSTOM TRAINED MODEL

Train it on your data,
only if you must.

Most teams don’t need this. RAG + a frontier model handles 80% of the cases people think need a custom model. If your domain is genuinely esoteric (legal, medical, niche industrial) and you have labeled data, fine-tuning a 7B–14B open model is the move.

We don't do
average.
You shouldn't
either.

We run AI strategy and implementation engagements for teams that want a sharper edge and less hedge. Engagements start with a 15-minute fit call — short, honest, no pitch deck.

Not every team is a fit. If we're not the right firm for what you're trying to do, we'll tell you in the first 10 minutes — and point you toward someone who is.

One email,
every week.

The week's verdict. News worth your time, and a chart you'll forward to your team.