Skip to content
PickAuditor
Learn · Metric

Brier score

How the Brier score grades a stated probability, what 0.25 means, why overconfidence is punished, and when PickAuditor computes it.

Published

Most services publish a pick and a price. Some also publish a probability — “Arsenal 61%” — and a probability can be graded in a way a bare pick cannot. The Brier score is the standard tool for that. PickAuditor computes it only where a service states a win probability, which is why many profiles show a dash in that column rather than a number.

The definition

For one pick with stated probability p and outcome o (1 if the pick won, 0 if it lost):

Brier = (p − o)²

A service's Brier score is the mean of that over its graded picks that carry a stated probability. It runs from 0 (every probability was 1.0 on a winner or 0.0 on a loser) to 1 (the reverse). Lower is better. Pushes have no outcome and are left out.

Example · hypothetical figures · four picks
Stated 0.70, won: (0.70 − 1)²
0.090
Stated 0.70, won
0.090
Stated 0.70, lost: (0.70 − 0)²
0.490
Stated 0.60, lost
0.160
Mean of the four
0.208

Two wins and two losses at an average stated probability of 0.675. The service was more confident than these four results justified, and the score says so before anyone looks at the money.

The number to remember: 0.25

A forecaster who says 50% for every pick scores (0.5 − 1)² = 0.25 on a win and (0.5 − 0)² = 0.25 on a loss. Their Brier score is 0.25 whatever happens. That is the baseline for a binary outcome: a score above 0.25 means the stated probabilities carried less information than saying nothing, and a score below it means they carried some.

The differences that matter are small in absolute terms. A forecaster who is perfectly calibrated at 60% — says 60% and wins 60% of the time — has an expected score of 0.6 × 0.16 + 0.4 × 0.36 = 0.240. That is a real edge over 0.250, and it looks like a rounding error. Read Brier scores to three decimals and with n, and treat gaps of 0.01 over fewer than a few hundred picks as noise.

Overconfidence costs more than caution

Example · hypothetical figures · two forecasters, same hit rate
Says 60%, wins 60%: 0.6 × 0.16 + 0.4 × 0.36
0.240
Says 90%, wins 60%: 0.6 × 0.01 + 0.4 × 0.81
0.330
Says 50% every time
0.250

The same picks, the same results. The forecaster who overstated confidence scores worse than one who said nothing, because the squared error punishes a confident miss (0.81) far more than it rewards a confident hit (0.01). This is what makes Brier useful for reading marketing: a service that attaches 85–95% to every pick will either have a hit rate to match or a Brier score that shows the gap.

Calibration and resolution

The score combines two things. Calibration asks whether events stated at 70% happened about 70% of the time. Resolution asks whether the probabilities were spread out and informative — 80% on some picks, 35% on others — rather than hovering near the base rate. A forecaster can be well calibrated with no resolution (always says 52%, wins 52%) and score close to 0.25; another can be sharp but miscalibrated and score worse. A calibration chart, which plots stated probability against observed hit rate in bins, separates the two; the single score does not.

Why the column is often empty

PickAuditor records a probability only when a page frames the number as a win probability. Many services show a confidence label instead — stars, “high”, “92% confidence” — and a confidence is not a probability: it is not clear what it would mean for it to be calibrated. Those labels are stored as text and do not feed the score. A service that publishes probabilities on some picks and not others gets a Brier score over the picks that had one, and the n next to the score is that smaller count, not the total. The rule that applies to every figure applies here: below 30 scored picks the column shows the sample and is not ranked.

Base rates and fair comparison

Example · hypothetical figures · no skill, two different markets
Favourites at 1.20 · says 83%, wins 83%: 0.83 × 0.17² + 0.17 × 0.83²
0.141
Coin-flip spreads · says 50%, wins 50%
0.250

Neither forecaster knows anything the price did not already say, yet one scores 0.141 and the other 0.250. Comparing the two says nothing about skill; it says one of them picks favourites.

Compare Brier within a market, and against the market's own implied probability rather than against 0.25. Read it beside CLV, which asks the related question of whether the service's view moved the market; see closing line value.

Log score

PickAuditor also stores the logarithmic score per pick: the natural log of the probability the service assigned to the outcome that happened, floored at one in a million. It punishes extreme misses harder than Brier does — saying 99% and losing costs −4.6 rather than 0.98 — and rewards them less. It is kept for calibration work and is not shown on profiles.

On the site

The leaderboard can be ordered by Brier score — ascending, since lower is better — over rankable services that have one. Profiles show it with its n and period, and the methodology has the exact computation. Where a service states no probabilities, the column is a dash, not a zero: it is a fact about what the service publishes, not a judgement.