The AI budget conversation has shifted. It's no longer "can we afford to use AI?" It's "why are we still paying frontier prices for tasks that cheaper models handle just as well?" That question is now driving real procurement decisions.
What Happened
Demand for AI infrastructure has forced an efficiency reckoning across the industry. As compute costs ballooned and ROI timelines stretched, operators started asking harder questions about model selection. Meanwhile, Snowflake moved to address AI costs directly through model routing, automatically directing queries to the cheapest model capable of handling them. That's not a feature, that's a philosophy shift.
The platform data backs this up. Look at what's sitting near the top of the competency rankings right now:
- Grok 4.6: competency score 98/100, blended cost $3/1M tokens (input $2, output $6)
- Veo 3: competency score 98/100, blended cost $4.5/1M tokens (input $4.5, output $13.5)
- Claude Opus 5: competency score 97/100, blended cost $10/1M tokens (input $5, output $25)
- Cartesia Sonic: competency score 100/100, blended cost $50/1M tokens (input $50, output $150)
Two of the four highest-performing models on the board cost under $5/1M tokens blended. That's not a rounding error, that's a structural shift in the value curve.
Why It Matters
For most enterprise workloads, the difference between a 97 and a 100 competency score is invisible to end users. What isn't invisible is a 3x to 17x difference in token cost at scale. If your team is routing every task through a premium model out of habit or vendor lock-in, you're leaving serious money on the table.
Scaling AI infrastructure is already painful enough without overpaying per token. And as one sharp editorial framing put it, the goal is to replace the overhead, not the output quality. Routing cheaper models to handle the scaffolding work is exactly that.
What To Do
This isn't about switching everything overnight. It's about being deliberate.
- Audit your task mix. Classify your AI workloads by complexity. Most enterprise tasks (summarization, classification, drafting, extraction) don't need a 100-point model.
- Run a cost-competency comparison. If a model scoring 98 costs $3/1M and your current model costs $10-50/1M, the burden of proof is on the expensive one.
- Ask your vendor about routing. If they don't offer model routing or tiered pricing, that's a negotiating point, or a reason to look elsewhere.
- Renegotiate with data. Usage logs plus a cost-per-task breakdown is leverage. Bring it to your next renewal conversation.
FAQ
Q: Does a lower-cost model actually perform well enough for real business tasks? A: For the majority of enterprise use cases, yes. Models scoring 97-98 on competency benchmarks handle most production workloads without meaningful quality loss. The gap between 97 and 100 matters in edge cases, not in bulk operations.
Q: What is model routing and should we be using it? A: Model routing automatically sends each query to the most cost-efficient model capable of answering it. Snowflake is building this natively into its platform. If your AI stack doesn't do this, you're probably overpaying on a significant share of your requests.
Q: How do I know which tasks need a premium model? A: Tasks requiring deep multi-step reasoning, nuanced judgment, or highly specialized domain knowledge benefit most from top-tier models. Routine extraction, summarization, and classification rarely do. Start by tagging your highest-volume tasks and testing a cheaper model against your current output quality bar.
Q: Is this just a cost-cutting exercise, or does it improve AI strategy? A: Both. Routing the right model to the right task reduces cost and improves reliability, because a model optimized for a narrower task often outperforms a generalist on that specific job. It's a better architecture, not just a cheaper one.
The goal is to replace the scaffold, not the artist.
The Deep View
Bottom Line
The era of defaulting to the biggest model available is over for anyone serious about AI ROI. With high-competency models now available at $3-5/1M tokens blended, the question isn't whether to optimize, it's how fast you can audit your current stack and renegotiate accordingly. Your next vendor conversation should start with a cost-per-task breakdown, not a feature list.