Every start in your report that carries a percentage is a prediction we record and grade. Before we sold one, we replayed eleven NFL seasons and graded 10,041 of those predictions against the real box scores. Our own grading rule — written down before the run, so it couldn't be argued with afterwards — says the result does not let us tell you the number is right. Here is everything it found anyway, including where it fails.
Seasons 2014 through 2024. Each call is the model's first choice at a slot against its own second choice, graded on what the two actually scored. Weeks 1–3 are out (a call needs three games of record behind both players), weeks 17–18 are out (fantasy seasons are over), and ties are set aside rather than counted either way.
That last figure is the one most products would put in a headline. Read on for why ours doesn't.
The question that matters isn't how often we were right overall — it's whether the number on each call meant what it said. When we put 60–65% on a call, those calls should land about that often. Here is every band, with what we said and what happened. Intervals are wide on purpose: inside one week the same benched player is the alternative at several slots, and one real game moves many calls at once, so the calls are not independent and the math doesn't pretend they are.
| When we said… | Calls | Decided | We said, on average | They landed | 95% interval | Verdict |
|---|---|---|---|---|---|---|
| 50–55% | 2,714 | 2,627 | 52.4% | 51.6% | 50–54% | passes |
| 55–60% | 2,311 | 2,233 | 57.4% | 58.9% | 57–61% | passes |
| 60–65% | 1,946 | 1,887 | 62.5% | 66.3% | 64–68% | off |
| 65–70% | 1,407 | 1,357 | 67.3% | 74.1% | 72–77% | off |
| 70–80% | 1,417 | 1,375 | 74.1% | 82.3% | 80–84% | off |
| 80–90% | 244 | 240 | 82.9% | 91.2% | 88–95% | off |
The grading rule was written before a single number was computed, and its clauses are all required rather than weighed against each other. This run cleared the error threshold (3.6% expected calibration error) and the spread threshold for a stronger grade, and failed on the count of passing bands. So, on every page of this site: no claim that the number is right. The percentage in your report prints as a recorded prediction, and the public record is where it gets settled — one graded call at a time, from mid-September, never edited after the fact.
Two more limits that travel with the number. It was graded for PPR scoring, 12-team leagues and the standard lineup; half-PPR, standard scoring, superflex and other league sizes are on the list to be graded on their own, and until they are, the same number prints for those setups without a grading behind it. And a team defense never carries a percentage at all: no defense was graded here, so that slot shows a projection and no number.
The excluded weeks got their own preregistered run before launch: for weeks 2 and 3 only, a player's own prior-season record enters at half weight, and the same grading machinery ran on weeks 2–3 of all eleven seasons — 1,692 calls. 61.2% of decided calls landed; the error across bands was 2.5%; four of five judgeable bands passed, and the one that failed, failed high again — where we wrote 73.2%, the calls landed 84.5% of the time. Under the rule's own terms that earns the middle grade: the figures may be stated as facts beside the failures, and the product ships week-2-and-3 numbers for the setup it measured — full PPR, 12 teams, the standard lineup — with "last season counted in" printed on every row that leans on it. Other setups wait for their own runs. And week 1 stays numberless for everyone: there is no week-0 injury report to check anyone against, so no model earns a say.