← Back to The Race
Ranking methodology
Current status: placeholderRankings are sorted by release date (newest first) so the field is well-defined, not because that is a credible ranking signal. Treat every rank shown on Hivig as illustrative until this page describes a real, sourced formula.
What the real methodology needs to define
- Which benchmarks count, and how they're weighted (e.g. LMSYS Chatbot Arena Elo, Artificial Analysis quality index, task-specific evals).
- How ties and missing benchmark data are handled.
- Refresh cadence and a staleness rule, matched to the 72-hour data refresh.
- What counts as the "same model" across dated snapshot releases, so a rank delta means something consistent.
Full detail lives in RANKING_METHODOLOGY.md in the project source.