Methodology.
How the scores are made.

Every number on ai.hiero.dev is built to be explainable in one sentence. Here is exactly how competency, cost, speed, and confidence are calculated, and where the underlying data comes from.

1. Competency (the vertical axis)

Each model has per-task sub-competency scores (e.g. for LLMs: summarisation, research, customer service, reasoning, code). The competency we plot is a weighted average of those sub-scores. The weights come from the business task or industry you select — change the task and the chart re-weights and redraws. With no task selected we use a balanced default weighting.

2. Cost (the horizontal axis)

Cost is normalised per category: $/1M tokens for text models, $/image, $/minute for audio and video. For self-hosted open-weights models we show an estimated GPU $/hour equivalent so cloud and local options sit on the same axis.

3. Speed (bubble size) & deployment (fill)

Bubble size encodes throughput/latency. A filled bubble is a cloud/API model; an outline bubble is an open-weights model you can self-host. The deploy filter lets you isolate either.

4. Confidence

Each score carries a confidence level — Green, Amber, or Red — derived from how much corroborating source data we have for that model and how fresh it is. Low coverage or stale data lowers confidence even when a score looks favourable.

5. Data sources

We aggregate, normalise, and weight signals from multiple public sources, including:

  • Artificial Analysis — pricing, speed, and quality indices
  • LMArena — human-preference Elo
  • SWE-bench & CodeSOTA — coding capability
  • Hugging Face — open-weights popularity and metadata
  • Provider documentation and additional benchmark feeds

We do not simply reprint these numbers. Each source is trust-weighted, scores are combined into a normalised competency, and our analysis — not the raw leaderboard — is what we publish.

6. Update cadence

Pricing, speed, and scores refresh daily from fast sources; slower benchmarks refresh weekly. New models are detected automatically and reviewed before publication, and a model is marked the latest in its family when a newer version supersedes it.

7. What we deliberately leave out

If a metric can't be explained to a business owner in one sentence, it isn't on the dashboard. We favour decision-useful clarity over benchmark completeness.

See the chart ↗