Why one big score ruins a projection

Two, twenty-three, two. Nine, nine, nine. Identical averages, entirely different players — and for three gameweeks this model could not tell them apart.

Advertisement

Before Gameweek 4, Bruno Fernandes had returned 2, 23 and 2. On season totals that is 27 points in three matches, which put him among the best attacking midfielders in the game. The model agreed enthusiastically: it ranked him first at 7.55 expected points and recommended him as captain, away at Manchester City.

Anyone who had actually watched the three matches knew that was absurd. The interesting question is not that the model was wrong — it is exactly how it was wrong, because the fix turned out to apply to far more than one player.

Averages destroy the evidence you need

A season total throws away the one thing that tells you whether a rate is real: whether it happened repeatedly. Three matches producing 0.3 expected goal involvement each, and three matches producing 0.9, nothing and nothing, give the same total and the same per-90 rate. The first is a player doing something consistently. The second is one afternoon and a lot of silence.

The second is also far more likely to be an accident. Football produces lopsided afternoons constantly — a deflection, a penalty, an opponent down to ten. If a player's whole case rests on one of them, you do not have three matches of evidence. You have roughly one.

Counting matches that count

So the model stopped counting appearances and started counting effective matches. For a player's expected goal involvement across their recent matches, it computes the square of the sum divided by the sum of the squares.

The behaviour is easier than the arithmetic. Four identical matches give exactly 4. Four matches where everything happened in one give a number close to 1. Everything real sits in between. It is the same idea as an effective sample size: not how many times you looked, but how many independent looks you actually got.

That number then decides how much a player's own figures are trusted against the baseline for their position. A player's observed rates carry a weight of m / (m + 5), where m is effective matches. The rest is pulled toward what an average player in that position does.

For Bruno, four appearances currently count as 1.78 effective matches. That gives his own numbers about 26% weight instead of the 44% four honest matches would have earned. The model is not calling him bad. It is declining to pay the premium his season total implies, on the grounds that it has barely seen him do it.

Who the model distrusts right now

These are live figures, after four completed gameweeks. Points are the player's actual returns; effective matches is what the model thinks those four appearances are worth as evidence.

PlayerTeamPoints, GW1–4AppearancesEffective
B.FernandesMUN2, 23, 2, 241.78
ØdegaardARS11, 3, 10, 341.95
WissaNEW4, 8, 1, 242.05
HaalandMCI2, 13, 9, 943.55
SchadeBRE3, 10, 2, 1543.42
Dewsbury-HallEVE11, 2, 2, 243.90

The part that catches people out

Look at the last two rows again. Dewsbury-Hall scored 11 and then three 2s — about as lopsided as a points record gets — and the model treats his four appearances as 3.90 matches of evidence, the most trusted figure in the table. Ødegaard returned 11, 3, 10 and 3, which looks like a player delivering regularly, and gets 1.95.

That is not a mistake. The measure reads expected goal involvement, not points. Points are the outcome; chances are the process. Ødegaard's returns were spread across four weeks while the underlying involvement behind them was concentrated. Dewsbury-Hall's returns were concentrated while his involvement was steady — his 11 arrived without much underneath it, and he has been doing the same quiet work since.

A scoreline tells you what happened. It is a poor guide to what a player is likely to do next, and the two disagree in both directions more often than is comfortable.

Why quiet games count in full

One asymmetry is deliberate. A player with no output at all is not discounted — their matches count as whole matches.

Three quiet games are not missing evidence. They are evidence, and it is good evidence: this player does not threaten. The concentration discount only applies when there is output to be lopsided about. Treating an absence of chances as an absence of information would let every anonymous forward keep a flattering baseline forever.

Where the number means nothing

It is worth saying plainly where this measure stops working, because it is printed on the player pages and it looks authoritative everywhere.

Sixty-eight players with four or more appearances have negligible expected involvement — under 0.1 per 90 — and twelve of them show exactly 1.00 effective matches, six of those goalkeepers. That is not a finding about those players. A keeper who picked up one stray 0.02 of expected involvement across four matches gets a concentration score of 1 by arithmetic, not by meaning.

It costs them almost nothing, because attacking rates are a tiny part of what a goalkeeper or a defensive defender is projected to score — their points come from saves, clean sheets and defensive actions, which are measured separately and unaffected. But the number on the page means less for them than it appears to, and it would be dishonest to display it without saying so.

What it does not claim

Effective matches cannot tell a lucky afternoon from a brilliant one. A player who took a hat-trick of genuine chances created by genuine movement gets discounted the same as one who fell over a deflection. The measure says only that the evidence is thin, not that the player is bad.

That is the right kind of wrong. A model that tried to judge which big afternoons deserved belief would be guessing, and would be guessing with exactly the bias that makes people captain the player who scored 23 last week.

Bruno today

He is projected at 5.06 for Gameweek 6, still owned by 39% of the game at £12.0m, and he scored 2 in Gameweek 4 after the change dropped him from first to tenth. One gameweek does not vindicate anything, and it was said at the time that it would have looked just as good by luck.

The argument does not rest on that week. It rests on the fact that 2, 23, 2 and 9, 9, 9 are different, and any model that reports them as the same is not measuring what it claims to measure. The rest of the method is on the how it works page, and what the model predicted before each deadline is on the scorecard page, unedited.

Advertisement