Beat Your League

The number next to a player

Every start in your report that carries a percentage is a prediction we record and grade. Before we sold one, we replayed eleven NFL seasons and graded 10,041 of those predictions against the real box scores. Our own grading rule — written down before the run, so it couldn't be argued with afterwards — says the result does not let us tell you the number is right. Here is everything it found anyway, including where it fails.

What the number means

The unit, exactly A percentage on a slot is the probability that the player we seat there outscores the best eligible alternative left on your bench at that slot — under our model, under your league's scoring, if the two don't both sit out. It is never "a good start" in general. It's this player over that player, in your lineup.

What was graded

Seasons 2014 through 2024. Each call is the model's first choice at a slot against its own second choice, graded on what the two actually scored. Weeks 1–3 are out (a call needs three games of record behind both players), weeks 17–18 are out (fantasy seasons are over), and ties are set aside rather than counted either way.

10,041calls graded
9,721decided — 320 ties set aside
11seasons, 838 to 956 calls each
64.6%of decided calls landed

That last figure is the one most products would put in a headline. Read on for why ours doesn't.

Bucket by bucket

The question that matters isn't how often we were right overall — it's whether the number on each call meant what it said. When we put 60–65% on a call, those calls should land about that often. Here is every band, with what we said and what happened. Intervals are wide on purpose: inside one week the same benched player is the alternative at several slots, and one real game moves many calls at once, so the calls are not independent and the math doesn't pretend they are.

When we said…CallsDecidedWe said, on average They landed95% intervalVerdict
50–55%2,7142,62752.4%51.6%50–54%passes
55–60%2,3112,23357.4%58.9%57–61%passes
60–65%1,9461,88762.5%66.3%64–68%off
65–70%1,4071,35767.3%74.1%72–77%off
70–80%1,4171,37574.1%82.3%80–84%off
80–90%24424082.9%91.2%88–95%off
Two of six pass. The four that fail all fail the same way: the calls landed more often than the number said. Where we wrote 67.3% the calls landed 74.1% of the time; where we wrote 74.1% they landed 82.3%. The rule counts that as a failure, and it should — a number that is wrong in a flattering direction is still wrong, and we are not allowed to round it in our own favour. What it does tell you is that the number is shy, not noisy: the most confident tenth of calls landed 86.1% of the time and the least confident tenth 50.5%, a 35.6-point spread, so the order the numbers put calls in is real even where the numbers themselves run low.

What the rule lets us say — and what it doesn't

The grading rule was written before a single number was computed, and its clauses are all required rather than weighed against each other. This run cleared the error threshold (3.6% expected calibration error) and the spread threshold for a stronger grade, and failed on the count of passing bands. So, on every page of this site: no claim that the number is right. The percentage in your report prints as a recorded prediction, and the public record is where it gets settled — one graded call at a time, from mid-September, never edited after the fact.

Two more limits that travel with the number. It was graded for PPR scoring, 12-team leagues and the standard lineup; half-PPR, standard scoring, superflex and other league sizes are on the list to be graded on their own, and until they are, the same number prints for those setups without a grading behind it. And a team defense never carries a percentage at all: no defense was graded here, so that slot shows a projection and no number.

The first weeks, graded on their own

The excluded weeks got their own preregistered run before launch: for weeks 2 and 3 only, a player's own prior-season record enters at half weight, and the same grading machinery ran on weeks 2–3 of all eleven seasons — 1,692 calls. 61.2% of decided calls landed; the error across bands was 2.5%; four of five judgeable bands passed, and the one that failed, failed high again — where we wrote 73.2%, the calls landed 84.5% of the time. Under the rule's own terms that earns the middle grade: the figures may be stated as facts beside the failures, and the product ships week-2-and-3 numbers for the setup it measured — full PPR, 12 teams, the standard lineup — with "last season counted in" printed on every row that leans on it. Other setups wait for their own runs. And week 1 stays numberless for everyone: there is no week-0 injury report to check anyone against, so no model earns a say.

Every number we print gets graded in public — wins and misses, dated, against the real box score.
Set up your team Read a real report