Every number on ai.hiero.dev is built to be explainable in one sentence. Here is exactly how competency, cost, speed, and confidence are calculated, and where the underlying data comes from.
Each model has per-task sub-competency scores (e.g. for LLMs: summarisation, research, customer service, reasoning, code). The competency we plot is a weighted average of those sub-scores. The weights come from the business task or industry you select — change the task and the chart re-weights and redraws. With no task selected we use a balanced default weighting.
Cost is normalised per category: $/1M tokens for text models, $/image, $/minute for audio and video. For self-hosted open-weights models we show an estimated GPU $/hour equivalent so cloud and local options sit on the same axis.
Bubble size encodes throughput/latency. A filled bubble is a cloud/API model; an outline bubble is an open-weights model you can self-host. The deploy filter lets you isolate either.
Each score carries a confidence level — Green, Amber, or Red — derived from how much corroborating source data we have for that model and how fresh it is. Low coverage or stale data lowers confidence even when a score looks favourable.
We aggregate, normalise, and weight signals from multiple public sources, including:
We do not simply reprint these numbers. Each source is trust-weighted, scores are combined into a normalised competency, and our analysis — not the raw leaderboard — is what we publish.
Pricing, speed, and scores refresh daily from fast sources; slower benchmarks refresh weekly. New models are detected automatically and reviewed before publication, and a model is marked the latest in its family when a newer version supersedes it.
If a metric can't be explained to a business owner in one sentence, it isn't on the dashboard. We favour decision-useful clarity over benchmark completeness.