← Back to The Race

Ranking methodology

Current status: placeholderRankings are sorted by release date (newest first) so the field is well-defined, not because that is a credible ranking signal. Treat every rank shown on Hivig as illustrative until this page describes a real, sourced formula.

What the real methodology needs to define

  • Which benchmarks count, and how they're weighted (e.g. LMSYS Chatbot Arena Elo, Artificial Analysis quality index, task-specific evals).
  • How ties and missing benchmark data are handled.
  • Refresh cadence and a staleness rule, matched to the 72-hour data refresh.
  • What counts as the "same model" across dated snapshot releases, so a rank delta means something consistent.

Full detail lives in RANKING_METHODOLOGY.md in the project source.