What does "frozen" mean for an FPL forecast?
A frozen forecast is a projection written down before the Gameweek deadline and never edited afterwards. At SquadPilot, thirty minutes before each deadline the inputs are refreshed, the model runs, and every player's expected minutes and expected points go into an append-only ledger. Each row carries a UTC timestamp, the data snapshot it was built from, and the model version that produced it.
Append-only is the important part. There is no code path that rewrites a frozen row. If the model says a defender will score 2.3 and he scores 17, that 2.3 stays on the record next to the 17. Outcomes are recorded later as separate observations and joined to the frozen row at read time, so the forecast and the result can be compared by anyone.
That is the difference between a forecast and a story told about one afterwards. Hindsight can improve the story. It cannot touch the row.
Why does a public score matter more than a track record?
Every FPL tool says it is accurate. Very few publish a number you can check, and a track record nobody can audit is really just a marketing line. A public score is evidence, and evidence has to include the weeks that went badly.
The Decision Ledger publishes the error after every Gameweek: the average miss per player, how well calibrated the start probabilities were, the error broken out by what players actually did, and the single worst call of the week, named. The Gameweek 1 miss on De Cuyper is the example the home page leads with, because a tool that never shows you its worst week has chosen not to.
The second reason a public score matters is that it changes how the model gets built. When every call is graded in the open, the incentive is to be calibrated rather than confident. An impressive number that cannot survive the scorecard is worth nothing here.
What did the Gameweek 1 ledger show?
Gameweek 1 of 2026/27 was the first frozen forecast SquadPilot published. Six hundred
player projections were locked before kick-off and graded once FPL confirmed the Gameweek.
Measured as mean absolute error, the average distance between projected and actual points
per player, the frozen forecast scored 1.4995 against 1.5983 for FPL's own free
ep_next estimate on the same players.
How is the error actually measured?
Four numbers, each with a published definition:
- Points MAE: the average absolute gap between frozen expected points and the points the player actually scored.
- xMins MAE: the same gap for minutes. Most points error is minutes error, so it gets reported on its own rather than buried in the points figure.
- Start Brier score: whether the start probabilities were calibrated. Players given a 70% chance of starting should start about 70% of the time.
- Biggest miss: the worst single call of the Gameweek, by name.
Error is also broken out by what each player did (did not play, blanked, returned, hauled), because an average can look respectable while missing every haul, and hauls are the outcomes that move your rank. That segment is where a points model has the most to prove, and the ledger says as much.
What is the forecast compared against?
An error figure needs something beside it before it means anything. So every frozen forecast
is scored against two free baselines any manager already has: ep_next, the
next-Gameweek estimate FPL publishes for every player, and points-per-game, the
assume-nothing-changed guess. Both run on exactly the same players in the same Gameweek.
Beating a broken baseline proves nothing, which is why the comparison gets published a second time with FPL's repeated default values removed. When the honest number is worse than the headline, the honest number goes on the page as well.
Are your own decisions scored as well?
Yes. Lock a decision in the app before the deadline, the model's recommendation or your own override, and it goes on the record alongside the projection that informed it. After the Gameweek, both paths are graded on the same terms. The ledger cuts both ways: it tells you when the model was wrong, and it tells you when you were.
Most managers find this the uncomfortable part, and the useful one. Gut calls feel better than they score. A ledger of your own overrides is the only way to find out whether your read of the game is worth backing over the numbers. Sometimes it is. SquadPilot never makes a move on your account. The decision is yours, and so is the grade.
How do you use a frozen forecast to make better FPL decisions?
- Read the caveats before the headline. Any scorecard worth reading publishes the number that flatters it least.
- Find the segment your decision lives in. Captaincy is a haul question; the bench is a minutes question. The average error tells you neither.
- Treat a start probability as a probability. A 70% starter will sit one week in three. The ledger shows whether that 70% was honest.
- Freeze your decision so it gets scored. An unrecorded call cannot teach you anything.
- Check the ledger after the whistle. Especially in the weeks you overrode the model.
You manage the squad. We map the decisions. The ledger is where that map gets checked against what actually happened, in public, every Gameweek.