Inference costs have been falling sharply, and the providers doing the cutting are not the names on most enterprise contracts. If your API bill hasn't moved, your vendor relationship has gotten a lot more expensive by comparison.
What Happened
A cluster of specialized inference providers, including Together AI, Groq, Fireworks, and Cerebras, have been aggressively undercutting frontier model pricing to win workloads away from the big labs. The mechanism is straightforward: these providers run open-weight models on purpose-built or highly optimized hardware, stripping out the research overhead baked into frontier lab pricing. Smaller models are increasingly competitive on capability, which means the "we need the biggest model" justification is getting harder to defend on a spreadsheet.
Meanwhile, frontier labs are being forced to respond, with moats shrinking faster than anyone publicly admits. The result is a market where the price you locked in six months ago is almost certainly not the best available price today.
Why It Matters
This is a cost-structure conversation, not a technology one. Consider what the current landscape looks like on capability vs. cost using verified platform data:
- Grok 4.6: competency score 97/100, blended cost $3/1M tokens (input $2, output $6)
- Claude Opus 5: competency score 97/100, blended cost $10/1M tokens (input $5, output $25)
- Veo 3: competency score 98/100, blended cost $4.5/1M tokens (input $4.5, output $13.5)
- Cartesia Sonic: competency score 100/100, blended cost $50/1M tokens (input $50, output $150)
Two models at identical competency scores (97) carry a 3x price difference at the blended rate. That gap is not explained by quality. It's explained by who you're buying from and whether you've shopped recently. Infrastructure choices compound: the teams that optimized their compute stack early are now running meaningfully lower unit economics than those who defaulted to the obvious vendor.
What To Do
You don't need to rip and replace your stack. You need a 90-minute audit.
- Pull your last 90 days of API invoices and calculate your effective cost per 1M tokens by model and provider.
- Map your workloads by sensitivity: production customer-facing calls, internal tooling, and batch jobs have very different risk tolerances for switching.
- Run a parallel test on one non-critical workload with a lower-cost provider. Groq and Fireworks both offer free tiers sufficient for a real benchmark.
- Check model competency vs. cost tradeoffs before assuming you need the most expensive option. A 97-competency model at $3/1M and a 97-competency model at $10/1M are not the same purchase decision.
- Renegotiate or switch on batch workloads first. These carry the lowest switching risk and the highest volume, so the savings show up fast.
FAQ
Q: Do I have to migrate my entire stack to capture savings? No. Start with batch or internal workloads where latency and reliability requirements are lower. A partial migration can still cut your bill materially without touching production.
Q: Are cheaper inference providers less reliable? Reliability varies by provider and workload type. Groq and Fireworks have enterprise SLAs and are used in production by large teams. Evaluate uptime and rate limits for your specific use case rather than assuming cost correlates with reliability.
Q: How do I compare models fairly if I'm not an ML engineer? Focus on competency scores (a normalized capability benchmark) and blended token cost. Two models with the same competency score at different prices are a straightforward cost decision, not a technical one.
Q: What if my contract locks me in with a current provider? Check your contract for volume commitment terms vs. minimum spend. Many enterprise API agreements have flexibility on model selection even within a vendor. You may be able to shift to a cheaper model tier without breaking the agreement.
Bottom Line
The inference price war is real, it is moving fast, and most enterprise API budgets haven't caught up. Run the audit, test one workload on a cheaper provider this week, and stop treating last quarter's pricing as a baseline.