Track record
Every gameweek the model publishes its predictions before the deadline and picks a real FPL team with them. After the last match settles, this page marks both against what actually happened — the weeks it got wrong included — and nothing here is edited afterwards.
So far: ahead of the average manager in 4 of 4 gameweeks (+38 points in total). A baseline beat the model on 6 measures: GW2, last season's points per game on spotting 10+ hauls; GW2, top-10k effective ownership on spotting 6+ hauls; GW2, top-10k effective ownership on spotting 10+ hauls; GW4, last season's points per game on points per player; GW4, top-10k effective ownership on spotting 10+ hauls; GW4, fpl's own projection (ep_next) on spotting 10+ hauls. 4 gameweeks in, read it as an early record, not a verdict.
Last settled GW4 · scored against the run of 12 Sept, 11:50 UTC · pre-registered · checkpoints at GW10, GW19, GW38
Points vs the average manager
289 to the average manager’s 251 · above the average in 4 of 4 gameweeks · Triple Captain in GW1
Against a sample of top-10k managers: −16.9 in GW2, −1.3 in GW3, −6.7 in GW4 (before hits · no sample for GW1). This early, the sample is chosen on the very weeks it is scored on, so a deficit is overstated.
Ordering the players
Spotting 6+ hauls · season
Overall rank
The model's own team, week by week
One real FPL entry, picked by the model before every deadline: what it predicted for the squad it submitted, what the squad scored, and where the average manager and a sample of the top 10k landed the same week.
Model said · scored · the field
| GW | Pts | vs avg | vs 10k | total | rank | predicted | chip |
|---|---|---|---|---|---|---|---|
| 59 | +9 | — | 59 | 1.99m | 65.7 | TC | |
| 91 | +10 | −16.9 | 150 | 1.98m | 59.2 | — | |
| 56 | +5 | −1.3 | 206 | 1.68m | 59.8 | — | |
| 83 | +14 | −6.7 | 289 | 1.33m | 63.0 | — |
Select a gameweek to see the team below. A dash means the figure could not be computed for that week; it is never a zero.
The team in GW483 ptsShow the teamHide
How accurate were the predictions?
League-wide, not just our team: every player the model expected to start, scored on what he then did. Points per player first, then whether the model ordered the players the right way and spotted the hauls — each week against a baseline that knows nothing but last season's points per game.
| GW | Points per playerMAE | Big missesRMSE | Orderingrank ρ | Predicted → scoredΣ xP → Σ pts | Playersn | Spotting 6+AUC | Spotting 10+AUC |
|---|---|---|---|---|---|---|---|
| GW1seed | 2.75 | 3.35 | 0.290.09–0.45 | 332 → 302 | 95 | — | — |
| GW2 | 2.37 | 3.17 | 0.360.23–0.48 | 680 → 666 | 203 | 67%†57–75 | 72%†57–85 |
| GW3 | 2.30 | 2.99 | 0.160.03–0.28 | 681 → 685 | 199 | 69%†60–77 | 64%†46–80 |
| GW4 | 2.64 | 3.44 | 0.310.18–0.43 | 686 → 715 | 205 | 70%62–78 | 58%†41–76 |
| SeasonGW1–GW4 | 2.48 | 3.23 | — | 2378 → 2368 | 702 | — | — |
† fewer than 50 hauls that week — too few to claim a result (GW2 6+: 43 · GW2 10+: 13 · GW3 6+: 42 · GW3 10+: 12 · GW4 10+: 14).
Ordering and the haul figures are per gameweek and are not averaged; the season comparison on the same players is in the tiles above. Likely starters only: players the model expected to play 60+ minutes (predicted E[minutes] > 60 (n=702 of 2458 player-fixtures)).
95% percentile intervals from a row-level bootstrap (1000 resamples, seeded). Player-fixtures within a gameweek share fixtures and clean sheets, so these intervals understate the true width — read them as a floor on the uncertainty, not the uncertainty. A haul AUC on fewer than 50 hauls is too few to claim.
What went wrong each week, and how the baselines did
GW46 misses · beaten by last season's points per game on points per player; top-10k effective ownership on spotting 10+ hauls; fpl's own projection (ep_next) on spotting 10+ haulsOpen
- A worse week than usual: the model missed by 2.64 points a player, against 2.48 across the season so far.
- Last season's points per game was closer than the model on points per player: 2.71 to 2.71 on the 162 players both scored.
- Top-10k effective ownership was better than the model at spotting 10+ hauls this week: on the 205 players both scored, it ranked a hauler above a non-hauler 66% of the time, the model 58%.
- FPL's own projection (ep_next) was better than the model at spotting 10+ hauls this week: on the 205 players both scored, it ranked a hauler above a non-hauler 74% of the time, the model 58%.
- Spotting 10+ hauls was no better than a coin toss this week: 58%, with an interval (41%–76%) that includes 50%.
- The team finished 6.7 behind a sample of top-10k managers (average 89.7, 289 sampled). This early, the sample is chosen on the very weeks it is scored on, so the gap is overstated.
- The model was closer than last season's points per game on the big misses: 3.45 to 3.51 on the 162 players both scored.
- The model ordered the players better than last season's points per game: 0.30 to 0.24 on the 162 players both scored.
- The model was better than last season's points per game at spotting 6+ hauls: 67% to 62% on the 162 players both scored.
- The model was better than last season's points per game at spotting 10+ hauls: 62% to 56% on the 162 players both scored.
- The model ordered the players better than top-10k effective ownership: 0.31 to 0.18 on the 205 players both scored.
- The model was better than top-10k effective ownership at spotting 6+ hauls: 70% to 61% on the 205 players both scored.
- The model ordered the players better than fpl's own projection (ep_next): 0.31 to 0.26 on the 205 players both scored.
- The model was better than fpl's own projection (ep_next) at spotting 6+ hauls: 70% to 60% on the 205 players both scored.
- The team scored 83, 14 above the average manager's 69.
- Only 14 players scored 10+ this week — too few to claim a result on the 10+ figure.
Scored against the run of 12 Sept, 11:50 UTC · rows a0497656…c5bd · full report →
Against the baselines
| Measure | Model | Baseline | Verdict |
|---|---|---|---|
| points per player | 2.71 | 2.71 | baseline ahead |
| big misses | 3.45 | 3.51 | model ahead |
| ordering the players | 0.30 | 0.24 | model ahead |
| spotting 6+ hauls | 67% | 62% | model ahead |
| spotting 10+ hauls | 62% | 56% | model ahead |
| Measure | Model | Baseline | Verdict |
|---|---|---|---|
| ordering the players | 0.31 | 0.18 | model ahead |
| spotting 6+ hauls | 70% | 61% | model ahead |
| spotting 10+ hauls | 58% | 66% | baseline ahead |
| Measure | Model | Baseline | Verdict |
|---|---|---|---|
| ordering the players | 0.31 | 0.26 | model ahead |
| spotting 6+ hauls | 70% | 60% | model ahead |
| spotting 10+ hauls | 58% | 74% | baseline ahead |
GW34 misses · no baseline aheadOpen
- Its ordering of the players barely matched the week's: rank correlation 0.16, where 1.00 would be a perfect order and 0 no relationship at all.
- Spotting 10+ hauls was no better than a coin toss this week: 64%, with an interval (46%–80%) that includes 50%.
- The team finished 1.3 behind a sample of top-10k managers (average 57.3, 170 sampled). This early, the sample is chosen on the very weeks it is scored on, so the gap is overstated.
- The best team the model could legally have bought scored 51 — 40% of the 127 a perfect-hindsight XI would have made.
- A better week than usual: 2.30 points of miss per player, against 2.48 across the season so far.
- The model was closer than last season's points per game on points per player: 2.40 to 2.48 on the 156 players both scored.
- The model was closer than last season's points per game on the big misses: 3.08 to 3.25 on the 156 players both scored.
- The model ordered the players better than last season's points per game: 0.20 to 0.07 on the 156 players both scored.
- The model was better than last season's points per game at spotting 6+ hauls: 68% to 56% on the 156 players both scored.
- The model was better than last season's points per game at spotting 10+ hauls: 64% to 43% on the 156 players both scored.
- The model ordered the players better than top-10k effective ownership: 0.16 to 0.09 on the 199 players both scored.
- The model was better than top-10k effective ownership at spotting 6+ hauls: 69% to 55% on the 199 players both scored.
- The model was better than top-10k effective ownership at spotting 10+ hauls: 64% to 49% on the 199 players both scored.
- The team scored 56, 5 above the average manager's 51.
- Only 42 players scored 6+ this week — too few to claim a result on the 6+ figure.
- Only 12 players scored 10+ this week — too few to claim a result on the 10+ figure.
Scored against the run of 4 Sept, 16:27 UTC · rows 6f8b387e…5a77 · full report →
Against the baselines
| Measure | Model | Baseline | Verdict |
|---|---|---|---|
| points per player | 2.40 | 2.48 | model ahead |
| big misses | 3.08 | 3.25 | model ahead |
| ordering the players | 0.20 | 0.07 | model ahead |
| spotting 6+ hauls | 68% | 56% | model ahead |
| spotting 10+ hauls | 64% | 43% | model ahead |
| Measure | Model | Baseline | Verdict |
|---|---|---|---|
| ordering the players | 0.16 | 0.09 | model ahead |
| spotting 6+ hauls | 69% | 55% | model ahead |
| spotting 10+ hauls | 64% | 49% | model ahead |
FPL's own projection (ep_next) — not available this gameweek: no FPL ep_next capture in this gameweek's pre-deadline window.
GW25 misses · beaten by last season's points per game on spotting 10+ hauls; top-10k effective ownership on spotting 6+ hauls; top-10k effective ownership on spotting 10+ haulsOpen
- Last season's points per game was better than the model at spotting 10+ hauls this week: on the 163 players both scored, it ranked a hauler above a non-hauler 82% of the time, the model 76%.
- Top-10k effective ownership was better than the model at spotting 6+ hauls this week: on the 203 players both scored, it ranked a hauler above a non-hauler 67% of the time, the model 67%.
- Top-10k effective ownership was better than the model at spotting 10+ hauls this week: on the 203 players both scored, it ranked a hauler above a non-hauler 74% of the time, the model 72%.
- The team finished 16.9 behind a sample of top-10k managers (average 107.9, 60 sampled). This early, the sample is chosen on the very weeks it is scored on, so the gap is overstated.
- 12 points sat on the bench and did not count (no Bench Boost).
- A better week than usual: 2.37 points of miss per player, against 2.48 across the season so far.
- The model was closer than last season's points per game on points per player: 2.34 to 2.39 on the 163 players both scored.
- The model was closer than last season's points per game on the big misses: 3.21 to 3.25 on the 163 players both scored.
- The model ordered the players better than last season's points per game: 0.43 to 0.23 on the 163 players both scored.
- The model was better than last season's points per game at spotting 6+ hauls: 71% to 66% on the 163 players both scored.
- The model ordered the players better than top-10k effective ownership: 0.36 to 0.27 on the 203 players both scored.
- The team scored 91, 10 above the average manager's 81.
- Only 43 players scored 6+ this week — too few to claim a result on the 6+ figure.
- Only 13 players scored 10+ this week — too few to claim a result on the 10+ figure.
Scored against the run of 28 Aug, 17:22 UTC · rows dc666d46…9013 · full report →
Against the baselines
| Measure | Model | Baseline | Verdict |
|---|---|---|---|
| points per player | 2.34 | 2.39 | model ahead |
| big misses | 3.21 | 3.25 | model ahead |
| ordering the players | 0.43 | 0.23 | model ahead |
| spotting 6+ hauls | 71% | 66% | model ahead |
| spotting 10+ hauls | 76% | 82% | baseline ahead |
| Measure | Model | Baseline | Verdict |
|---|---|---|---|
| ordering the players | 0.36 | 0.27 | model ahead |
| spotting 6+ hauls | 67% | 67% | baseline ahead |
| spotting 10+ hauls | 72% | 74% | baseline ahead |
FPL's own projection (ep_next) — not available this gameweek: no FPL ep_next capture in this gameweek's pre-deadline window.
GW1seed5 misses · no baseline aheadOpen
- A worse week than usual: the model missed by 2.75 points a player, against 2.48 across the season so far.
- Too optimistic overall: it predicted 332 points for 95 likely starters, who scored 302 — 10% high.
- 11 points sat on the bench and did not count (no Bench Boost).
- Captain Mbeumo scored 2 — tripled to 6 by the Triple Captain.
- The best team the model could legally have bought scored 38 — 26% of the 148 a perfect-hindsight XI would have made.
- The model was closer than last season's points per game on points per player: 2.75 to 2.93 on the 95 players both scored.
- The model was closer than last season's points per game on the big misses: 3.35 to 3.38 on the 95 players both scored.
- The model ordered the players better than last season's points per game: 0.29 to 0.27 on the 95 players both scored.
- The team scored 59, 9 above the average manager's 50.
- No haul figures for this week: the predictions it is scored against did not keep haul probabilities, so there is nothing to mark. Not a zero.
- Gameweek 1 is marked against the predictions committed before the season started (the "seed"), because live runs were not being kept yet.
- No top-10k sample was taken for this week, so there is no elite comparison.
Scored against the seed of 18 Aug, 21:51 UTC · rows 4d566e22…dc02 · full report →
Against the baselines
| Measure | Model | Baseline | Verdict |
|---|---|---|---|
| points per player | 2.75 | 2.93 | model ahead |
| big misses | 3.35 | 3.38 | model ahead |
| ordering the players | 0.29 | 0.27 | model ahead |
Top-10k effective ownership — not available this gameweek: no top-10k capture before this gameweek's deadline.
FPL's own projection (ep_next) — not available this gameweek: no FPL ep_next capture in this gameweek's pre-deadline window.
Every stated population, and the reliability tablesThe auditor’s view: the same error on four groups of players, and whether the model’s probabilities mean what they say.OpenClose
Points per player, on every group we could score it on
| Group | GW1 | GW2 | GW3 | GW4 | Season |
|---|---|---|---|---|---|
| Every player the model projected | 1.49n=584 | 1.18n=602 | 1.08n=618 | 1.18n=654 | 1.23n=2458 |
| Players the model gave any minutes | 1.67n=519 | 1.27n=556 | 1.29n=518 | 1.41n=549 | 1.41n=2142 |
| Players who actually playedhindsight | 2.21n=301 | 1.89n=299 | 1.91n=290 | 2.13n=307 | 2.04n=1197 |
| Likely startersthe headline | 2.75n=95 | 2.37n=203 | 2.30n=199 | 2.64n=205 | 2.48n=702 |
About half of all projected rows score exactly 0, which is why the every-player figure is small and is never shown alone. The players-who-played row selects on the outcome — a player who did not play is not in it — and is marked hindsight for that reason. Comparing with anyone else’s error figure only means something on the same group.
When the model says 70%, does it happen 70% of the time?
| Predicted | Observed | 95% range | n |
|---|---|---|---|
| 0% | 0% | 0–1 | 492 |
| 1% | 0% | 0–2 | 246 |
| 3% | 3% | 1–6 | 245 |
| 7% | 13% | 9–18 | 246 |
| 16% | 25% | 20–31 | 246 |
| 35% | 44% | 38–50 | 245 |
| 65% | 65% | 59–71 | 246 |
| 85% | 89% | 84–92 | 246 |
| 93% | 94% | 90–96 | 246 |
| Predicted | Observed | 95% range | n |
|---|---|---|---|
| 0% | 0% | 0–1 | 378 |
| 0% | 0% | 0–2 | 185 |
| 1% | 1% | 0–4 | 187 |
| 2% | 2% | 1–5 | 187 |
| 5% | 3% | 1–6 | 188 |
| 10% | 8% | 5–13 | 187 |
| 15% | 13% | 9–18 | 187 |
| 22% | 25% | 19–31 | 187 |
| 34% | 37% | 30–44 | 188 |
| Predicted | Observed | 95% range | n |
|---|---|---|---|
| 0% | 0% | 0–1 | 378 |
| 0% | 0% | 0–2 | 185 |
| 0% | 0% | 0–2 | 187 |
| 0% | 0% | 0–2 | 191 |
| 1% | 1% | 0–3 | 185 |
| 2% | 2% | 1–5 | 187 |
| 4% | 2% | 1–5 | 188 |
| 5% | 6% | 4–11 | 186 |
| 10% | 11% | 7–16 | 187 |
Bins of the model’s stated probability against the rate at which it happened, with Wilson 95% intervals. A calibrated forecaster’s observed column tracks its predicted column; the range says how much a small bin can be trusted.
Who beat the model's number, and who fell short
The players furthest above and below their predictions so far this season.
Points scored minus the model’s prediction, added up over the settled gameweeks. Only weeks the model expected the player to start (60+ minutes) count. Two seasons of form fit in a fortnight of variance; read it as colour, not evidence.