What Uncertainty Buys a Rating System
Elo gives a competitor one number. Glicko, Glicko-2 and TrueSkill give it two: a rating, and some measure of how sure the system is about that rating. Everybody knows the second number is there. Fewer people know what it actually does to a prediction, and the answer surprised me enough when I was working through elote’s implementations that I think it’s worth a post.
The short version: in Glicko, your own uncertainty has no effect at all on your predicted score. Only your opponent’s does. TrueSkill disagrees, and uses both. That difference is not a detail, it changes what the number means and what you can do with it.
The second number
The three systems name it differently and compute it differently, but it plays the same role.
Glicko calls it the rating deviation, RD. It starts at 350 for a new competitor, shrinks as results come in, and inflates again during inactivity. Roughly, the true rating is within two RD of the reported one.
Glicko-2 adds volatility, which tracks how erratic a competitor’s results have been. A player with steady results and a player with wild swings can have the same rating and the same RD but different volatility, and the volatile one moves further per result.
TrueSkill models skill as a Gaussian and carries sigma, the standard deviation of that belief, alongside mu, the mean.
Whose uncertainty moves the prediction
Here is the part worth internalizing. Take a 2500-rated Glicko player against a 1500-rated one, and vary only the underdog’s RD:
opponent RD 30 -> 0.9968
opponent RD 50 -> 0.9966
opponent RD 100 -> 0.9959
opponent RD 200 -> 0.9923
opponent RD 350 -> 0.9792
The prediction softens as the opponent gets less known. A 1000-point gap against a settled opponent is a 99.7% proposition; against a brand-new opponent it drops to 97.9%. That is a big move in the space that matters, roughly a sixfold increase in the chance of an upset.
Now hold the opponent fixed at RD 30 and vary the favorite’s own RD instead:
own RD 30 -> 0.9968
own RD 50 -> 0.9968
own RD 100 -> 0.9968
own RD 200 -> 0.9968
own RD 350 -> 0.9968
Nothing. Not a rounding difference, exactly nothing.
That is not a quirk of the implementation, it falls out of the formula. Glicko’s expected score is
E = 1 / (1 + 10^(-g(RD_j) * (r_i - r_j) / 400))
where g attenuates the rating gap toward zero as its argument grows. The only RD in that expression is RD_j, the opponent’s. Your own uncertainty never enters. It governs how much your rating moves after the result, not what the system expects to happen.
Once you see it, the logic is clear enough. The question “how likely is A to beat B” is a question about how good B is. Your uncertainty about A is already priced into A’s rating being the number it is.
The probabilities don’t sum to one
This has a consequence that looks alarming and isn’t. Take an 1800-rated player who stops playing, against a settled 1500:
gap their RD they beat 1500 1500 beats them
0 mo 30.0 0.8480 0.1520
3 mo 329.6 0.8480 0.2327
6 mo 350.0 0.8480 0.2395
12 mo 350.0 0.8480 0.2395
At rest the two directions sum to exactly 1. After six months of inactivity they sum to 1.088.
Both numbers are correct by Glicko’s own definition. When the 1800 player is the subject, the system attenuates by the opponent’s RD of 30, which is small, so the prediction is unchanged. When the 1500 player is the subject, it attenuates by the returning player’s RD of 350, which is large, so their chances get talked up. Each direction asks about a different opponent’s uncertainty, and those are now different quantities.
So Glicko’s expected_score is not a symmetric function of its two arguments once the RDs differ. If you are storing predictions, or feeding them to something that assumes a proper probability distribution over two outcomes, that matters. Pick a direction and be consistent, or normalize deliberately and know that you have.
I like this example because it is the kind of thing you only find by evaluating both directions and adding them up. It is worth doing that once for any rating system you rely on.
TrueSkill makes the other choice
TrueSkill’s expected score comes from the distribution of the performance difference, and for two players its variance is
2*beta^2 + sigma_i^2 + sigma_j^2
Both sigmas are in there. So uncertainty on either side softens the prediction. Same mu gap of 30 against 25, varying both sigmas together:
sigma 8.333 -> 0.6476
sigma 5.000 -> 0.7059
sigma 2.500 -> 0.7653
sigma 1.000 -> 0.7936
And varying only the opponent’s, holding the favorite at sigma 1.0:
opp sigma 8.333 -> 0.6866
opp sigma 5.000 -> 0.7385
opp sigma 2.500 -> 0.7784
opp sigma 1.000 -> 0.7936
Both columns move, which is the opposite of what Glicko does. It also means TrueSkill’s expected score is symmetric, and the two directions do sum to one.
That same variance expression is what TrueSkill’s match quality uses, which is a nice internal consistency check: one class computing the same physical quantity in two places should agree with itself. When I was going through these implementations, asserting that equality caught more than a dozen shape-checking tests would have.
What the uncertainty is for besides prediction
Two other things fall out of tracking sigma or RD.
A conservative rating you can display. TrueSkill reports mu - 3*sigma, which is a lower bound you can be fairly confident the player exceeds:
mu=25 sigma=8.333 -> rating 0.001
mu=25 sigma=5.000 -> rating 10.000
mu=25 sigma=2.500 -> rating 17.500
mu=25 sigma=1.000 -> rating 22.000
A brand-new player sits at essentially zero no matter how good they might be, and climbs as the system becomes convinced. That is exactly the behavior you want on a leaderboard, because it makes a player earn their position by playing rather than by getting lucky twice. It also means the displayed number can be negative for someone with very few games, which is worth knowing before it shows up in your UI.
A cold start that admits it’s a cold start. Two brand-new Glicko players 300 points apart:
both RD 350 -> 0.7605
both RD 30 -> 0.8480
Same gap, very different confidence. Elo cannot express that distinction, because it has nowhere to put it.
Choosing
If you need a single displayable number that is honest about inexperience, TrueSkill’s conservative estimate gives you one for free, and its predictions respond to uncertainty on both sides.
If you want predictions that stay stable for your established players regardless of their own layoffs, Glicko’s asymmetry is a feature. A veteran’s expected score against a known opponent does not drift just because they took the summer off, while everyone else’s expectations against them do soften.
Either way, the test worth writing is the one I ran to make these tables: hold everything fixed, vary one uncertainty at a time, and check that the number moves in the direction and by the magnitude the paper says. Both directions, and add them up. All of the code above is in elote, and every number in this post came out of running it.
Stay in the loop
Get notified when I publish new posts. No spam, unsubscribe anytime.