Nobody had published an answer, so we measured one — and we'll lead with the part that stings: on the decisions both sides could call, Sleeper's free projections beat our own numbers, 68.8% to 64.4%. Here's the whole test, including why we still couldn't just use theirs.
We took a frozen set of real start-sit decisions from a full 2018 league season — the set our earlier league study graded, decided before this test existed, and a different measurement from the grading record the product publishes today — and asked Sleeper's projection feed (which is Rotowire-sourced and free) to make the same calls. Same head-to-heads, same real box scores, same decision rule. No cherry-picking is possible because the question set was locked first.
| Season | Decisions both could call | Our hit rate | Sleeper's hit rate | They were right, we were wrong | We were right, they were wrong |
|---|---|---|---|---|---|
| 2018 | 368 | 64.4% | 68.8% | 47 | 31 |
Read the gap with its uncertainty: 47 to 31 on disagreements gives a two-sided p-value of 0.089 — suggestive, not conclusive, on one season of one league. We don't quote the 68.8-vs-64.4 gap without that number, and neither should you. Also honest: the archive has no usable records for 2017, and whether archived projections were ever revised after the fact is unknowable from outside — we graded what the archive holds.
| Season | Player-weeks | Our average miss | Sleeper's average miss |
|---|---|---|---|
| 2018 | 1494 | 6.61 points | 6.39 points |
Because of what the feed can't see. It carries roughly 400–520 players in any given week — every era we checked — and real bench decisions constantly involve players outside that list. On our frozen decision set it simply had no opinion on 626 of 994 head-to-heads: nearly two out of three real decisions, unanswerable from the feed alone.
Its one unambiguous win is Week 1, where it projects most starters while any history-based approach is still waiting for a game to be played. Its structural loss is depth, which is exactly where start-sit decisions live.
Nothing quietly. Our reports keep publishing our own numbers, with the calls graded in public — and this page stays up, including the table where the free feed beats us, because one league-season of evidence is enough to decide what to test next, not enough to bury. If we ever blend the feed in, the ranges and confidence rules get re-tested first; until then, what you see in a report is what has been measured.