The GOAT Standard — v1.0

How goat ranks the greatest of all time: what we measure, how we score it, and how the AI Panel and fans get their say.

The ranking is never final. The most interesting part of goat is the conversation, not the ranking.

1. Principles

  1. Three lenses, never blended. Every category is ranked three separate ways, shown side by side:
    • The Machine: measured record only, computed by published formula, within one sport at a time.
    • The Panel: independent judges (AI models first, credentialed humans later), each filing a signed, dated ballot with reasons.
    • The Championship: the fans' vote. Disagreement between lenses is a feature. It is where the debate starts.
  2. Transparent. Every input, weight, and formula is published. Anyone can recompute a score from the page.
  3. Reproducible. Scores are a pure function of (data snapshot, standard version). Same inputs, same output.
  4. Versioned. Weights and indicators change only through a new version of the standard. Old versions stay readable. Every score states its version.
  5. Honest about uncertainty. Every score carries a data-coverage figure and a rank range (how far the rank moves when weights are perturbed).
  6. Era-fair. Prefer measures that are already relative to contemporaries. An honor that did not exist in someone's era counts as not available, not as zero.
  7. Sourced. Every data point has a source and an as-of date.

2. The six pillars

Every category is described through the same six pillars. Pillars 1–4 are measured (the Machine). Pillars 5–6 need judgment (the Panel).

# Pillar Question it answers Lens
1 Achievement How many of the sport's top prizes did they win? Machine
2 Dominance How far above their contemporaries were they? Machine
3 Peak How high was their single best stretch? Machine
4 Longevity How long did they stay elite? Machine
5 Impact Did they change the sport itself? Panel
6 Character Are they worth emulating? What can we learn from them? Panel

Competition (strength of rivals beaten, head-to-heads against other greats) will become a measured pillar once rivalry data is collected. Until then, the Panel weighs it inside its judgment.

3. The process

  1. Define the category. Be specific: "men's singles tennis, Open Era", not "tennis". Mixed categories need an explicit reason.
  2. Define eligibility. Who is in the pool, and what is the minimum record to be considered? Publish the rule.
  3. Pick indicators for each measured pillar. Two to five per category, each with a source. For each indicator, record:
    • which pillar it feeds
    • its floor (the value that scores 0, e.g. 50% for a win rate)
    • whether it is era-limited (did not exist for part of the pool)
  4. Normalize. For each indicator, score = 100 × (value − floor) / (best − floor), where best is the best value in the eligible pool. The all-time leader in an indicator scores 100.
  5. Aggregate. Pillar score = mean of its available indicator scores. GOAT score = weighted mean of available pillar scores, re-normalized over the pillars that have data.
  6. Measure uncertainty.
    • Coverage = share of indicator slots with data.
    • Rank range = min/max rank across 500 seeded random perturbations of the pillar weights (each weight × a factor in [0.5, 1.5]).
  7. Convene the Panel. Each judge gets the same packet: this standard, the category, the candidates, and their data. Each returns a ranking, Impact and Character ratings (0–10), and a rationale per pick. Ballots are stored verbatim, dated, and never edited.
  8. Open the Championship. Fans vote (one ballot per person per category; ballots can change). Fan results are shown on their own, never folded into the Machine.
  9. Publish all three lenses side by side, with sources, version, and rank ranges.
  10. Revise. Data refreshes re-run the Machine automatically. Changes to weights or indicators require a new standard version with a changelog.

4. Default weights (v1.0)

Pillar Weight
Achievement 35%
Dominance 25%
Peak 20%
Longevity 20%

Why this order: championships are the currency every sport agrees on (Achievement). Dominance comes next because it is the most era-fair signal. Peak and Longevity split the classic "best ever vs. greatest career" argument evenly.

5. Comparing across sports

The Machine never ranks across sports. Its indicators only mean something inside their own sport: a golf major is not a tennis Slam, and a score of 90 in golf and 90 in tennis are not the same achievement. Normalizing them onto one scale would look precise while comparing unlike things.

Cross-sport rankings are therefore judgment calls, made by the Panel (and, once voting opens, the Championship). Each judge picks and ranks its top 10 across all sports and gives a reason for every pick.

Changelog