Ask most people how good their model is and they'll quote accuracy: how often the top pick won. For betting, that's close to useless. Back the favourite every week and you'll look very "accurate", but the odds already price the favourite in, so you earn nothing. What pays is getting the probability right, especially where yours differs from the market's.
That's calibration. When you say 40%, does it happen about 40% of the time? Three scores measure it properly, and all of them are "proper" scoring rules, meaning your expected score is best when you report what you genuinely believe.
Brier score (one yes/no event) = (p - o)^2
70% forecast, it happens: (0.7 - 1)^2 = 0.09
70% forecast, it doesn't: (0.7 - 0)^2 = 0.49
Always saying 50%: 0.25 (the "know-nothing" mark)
Log loss = -ln(probability you gave to what actually happened)
80% on the winner: 0.22 8% on the winner: 2.53
RPS (three-way, ordered Home / Draw / Away)
= 1/2 x [ (pH - oH)^2 + (pH + pD - oH - oD)^2 ]
Brier, introduced by Glenn Brier in 1950, is the mean squared error of your probabilities, and lower is better. Log loss is brutal on confident misses. The Rank Probability Score respects the fact that a home win is "closer" to a draw than to an away win, and Constantinou and Fenton argued in 2012 that it's the rule that properly assesses football forecasting models.
Take an illustrative match with three forecasters: a sensible model (55/25/20), the market (45/28/27) and an overconfident model (80/12/8).
| Result |
Model RPS |
Market RPS |
Overconfident RPS |
Overconfident log loss |
| Home win |
0.121 |
0.188 |
0.023 |
0.22 |
| Draw |
0.171 |
0.138 |
0.323 |
2.12 |
| Away win |
0.471 |
0.368 |
0.743 |
2.53 |
Illustrative numbers. Lower is better.
The overconfident forecaster looks a genius when the home side wins and gets hammered otherwise. Accuracy would score the 55% and 80% forecasts identically. Only averages over hundreds of matches reveal who's calibrated, so pair the scores with a calibration table: bucket your forecasts (0-10%, 10-20% and so on) and check the hit rate in each.
Why it matters for profit
Walsh and Joshi put this to the test on NBA data, training models over several seasons and running betting experiments on one held-out season with published odds. Choosing models by calibration returned an average +34.69%, against -35.17% when choosing by accuracy. It's basketball and a single test season, so take it as strong backing for the principle rather than a recipe. But it's a principle we'd build every model around: score probabilities, not picks.