Methodology and accuracy

How the numbers are made, and how well they work

Every figure on this page is read out of the evaluation artifact the site publishes, at the moment the page is built. Nothing here is typed in by hand, and nothing is left out because it is unflattering.

The evaluation

The weekly model is trained on 2000-2016, has its settings chosen on 2017-2021, and is measured on 2022-2025, which it never sees during training or selection. The comparison is against a Five-game rolling player average, which is the honest thing to beat: it is what you would get from no model at all.

Model version 177.0. Evaluation generated 2026-09-14 00:54:57 UTC. 35 features.

Results, with the noise floor applied

A difference smaller than the test can resolve is not evidence. The last two columns are the smallest difference this evaluation can distinguish from chance, and whether the measured difference clears it.

MetricThis modelBaselineDifferenceSmallest difference the test can resolveVerdict
Mean absolute error4.65615.0490+7.78%1.48%measurable
Lineup regret97.09108.13+10.21%3.52%measurable
Points captured0.77640.7532+3.08%1.26%measurable
Ranking quality (nDCG)0.73520.7102+3.52%1.29%measurable
Rank correlation (Spearman)0.53310.4903+8.73%1.97%measurable
Pairwise accuracy0.58230.5566+4.62%2.44%measurable
Close-call accuracy0.52260.5198+0.54%0.67%not measurable

6 of 7 metrics show a difference the test can resolve. 1 do not: Close-call accuracy.

The one that matters most is the one we cannot claim

On close-call accuracy the model is +0.54% against the baseline, and this evaluation cannot resolve a difference smaller than 0.67%. So the correct statement is that the model is not measurably better than a five-game rolling player average on that metric.

This matters more than the rest of the table, because the artifact itself names closeCallAccuracy the preferred start/sit measure. closeCallAccuracy is 3.6x more precise than pairwiseAccuracy and measures the comparison a user actually makes; pointsCaptured is 2.8x more precise than lineupRegret and is the same quantity normalised. Both were adopted in v178 on measured precision.

The promotion gate, including what it fails

The model has a set of criteria it is supposed to clear before it is treated as an improvement. It does not clear them. The rankings are published anyway, and that is a decision rather than a pass.

CriterionResult
MAE Improvement of at least 5%passes
nDCG Improvement of at least 3%passes
Lineup Regret Improvement of at least 5%passes
Pairwise Accuracy of at least 60%FAILS
No Position MAE Regression over 1%passes

Read this before trusting a ranking. The gate above is not passed. The projections are better than a rolling average on error and on lineup regret by margins the test can resolve, and they are served on that basis. They are not a model that has met its own bar, and anyone telling you a fantasy model is solved is selling something.

Known limitations, from the artifact

What the season-long board is

A separate model from the weekly one. It ranks the draft pool for a specific league rather than projecting one week, and it is updated on Tuesday mornings from what has actually happened this season instead of being left at what was known in August. The served board was last updated 2026-09-16 23:23:52 UTC.

What we do not do

Judge it yourself

The weekly rankings preview is public and updated before every slate, and so is the survivor board and the waiver wire. Compare them against what happens. That is the only test that counts, and it is the reason the previews are readable without an account.

Pricing and what each tier includes: /pricing.