Every class was kept out of training, then scored against the real NFL draft and the first four seasons of pro production — the rookie-contract window, so every class is graded on the same ruler and recent classes count. (Full careers get their say on the re-grade pages.) Draftanomics's board grades out ahead of the analyst consensus and a step shy of NFL war rooms — win, lose, or draw against the league, you see every hit and every whiff below. The sample, everywhere on this page: 11 draft classes graded, 5 with complete four-year windows.
The picks the model nailed and the ones it whiffed on are named — no cherry-picking.
Over the 2018-2022 window, Draftanomics' Career Score correlated at 0.51 with actual career value — well clear of the 0.38 posted by analyst consensus. NFL war rooms still win the argument at 0.60, and that gap makes sense: they've got medicals and private workouts no public model touches. The real story isn't that Draftanomics is infallible. It's that public consensus, the thing most fans anchor to, is the weakest predictor of the three. That's the gap worth arguing about.
Written by our models from the live board; every name and number checked
Out of sample · how well each board predicts real NFL career value
Each number is the out-of-sample rank correlation between a board's pre-draft ordering and what prospects actually became — their real NFL Career Score — our open-data career-value index (snaps, durability, second-contract market value) — pooled across 2018–2022. 0 is a coin flip; 1 is perfect foresight. Draftanomics's board (0.51) ranks ahead of the entire public analyst field (0.38) and lands a step shy of NFL war rooms (0.60), who grade with medicals, interviews, and private workouts no public model ever sees. Every class was held out of training before it was scored. The sample: 11 draft classes graded, 5 with complete four-year windows.
How well does the model predict?
MODERATE
All four tiers, averaged?0.623
STRONG
Spotting elite careers?0.704
MODERATE
Spotting busts?0.676
The score behind each word: how often the model correctly ranks two prospects' outcomes against each other — its AUC, where 0.5 is a coin flip and 1.0 is perfect. The model scores prospects into four outcome tiers (bust / role / starter / elite), with 8 post-draft features excluded — the model only sees pre-draft signal.
Leave-one-year-out:0.617MODERATEmacro AUC when the model never trains on the holdout year. This is the production-realistic number. We publish both because the gap between them is honest.
Starter tier?
WEAK0.596
Role tier?
AT CHANCE0.518
Bust vs not — pre-draft?
MODERATE0.686
Training set
1,663 prospects
Per-position diagnostics and known weak spots are in the model-trust section below. Model trained 2026-06-06.
Technical validation details
Evaluation scheme: stratified k-fold plus leave-one-year-out cross-validation (the model never trains on the slice it's graded on), with bootstrap confidence intervals and a Youden's-J bust threshold · 2018-2023 train + held-out fold. Run SHA 34162549. Full validation write-up on the methods page.
How the model has improved
Each row is a model iteration shipped over the past year, graded on how well it separates eventual busts from non-busts on held-out folds. One caveat matters: the rows below the divider grade a different, easier task— they use NFL-career signals that don't exist before draft night. The pre-draft number that matters is the one above: 0.686 on bust-vs-not, knowing only what everyone knew before the draft.
Baseline
0.564
pre-draft
+ nflverse contracts
0.596
pre-draft
+ class weighting
0.604
pre-draft
+ hard-cut bust definition
0.616
pre-draft
+ full-career proxy
0.588
pre-draft
+ Wikipedia awards label
0.589
pre-draft
A different task from here down.These rows grade players using signals from their NFL careers — play quality once they're in the league, second contracts. Grading a player after you've watched him play pro football is a far easier job than projecting him before the draft, so these numbers are not comparable to the pre-draft accuracy above.
+ Charting-proxy NFL grade
0.890
NFL-side
+ PFR / NGS advanced
0.902
NFL-side
+ Snap consistency + injuries
0.883
NFL-side
+ Continuous AV regression proxy
0.935
NFL-side
+ Second-deal AAV predictor
0.933
NFL-side
Bar scaled 0.50–0.95 for visual contrast. The two dips above the divider (full-career proxy + awards label) were label changes — they made the positive class harder, which is honest progress even though the accuracy number dipped on its own.
AAV predictor: how close is it?
The model predicts a player's second-deal AAV — the average annual value of his second contract, in $M per year — from every NFL-side signal it has. Below: accuracy on the 1,333 prospects who've signed second deals 2018-2024, with every prediction scored by a model that never trained on that player. The typical miss is $0.61M (median); the average miss, dragged up by the big-money outliers, is $2.72M.
within $3M
75%
999 / 1,333
within $5M
82%
1,098 / 1,333
within $8M
89%
1,182 / 1,333
within $12M
94%
1,257 / 1,333
5-fold out-of-fold cross-validation — each prediction comes from a model that never saw that player. The 75% within $3M number is the headline — most second-deal AAVs are predicted to within rookie-deal money of the actual signed value, which is hard to do from public data alone. The 6% past $12M is the QB tail (Mahomes, Allen, Burrow) where small predictive errors translate to big dollar differences.
Model trust by position
How reliably the model calls bust risk, position by position. STRONG means its bust calls at that position have real teeth; WEAK or LOW means the model has thin signal there — take its bust numbers for those positions with more salt.
Offensive lineMODERATE
0.75
160 prospects graded
WRMODERATE
0.74
199 prospects graded
SMODERATE
0.72
119 prospects graded
EDGE / DL / LBMODERATE
0.69
478 prospects graded
RB / WR / TEMODERATE
0.68
429 prospects graded
CB / SWEAK
0.64
319 prospects graded
CBWEAK
0.58
200 prospects graded
RBWEAK
0.58
74 prospects graded
LBLOW
0.43
210 prospects graded
Calibration
Does P(bust) mean what it says?
2018-2024 · n=1,798
The verdict: Reliable up to about 30% bust risk; above that the model undershoots — when it says 75%, about 39% actually bust (a thin bin — only 28 calls that extreme).Dot color: within 5 points · within 15 · off by more
Most overconfident in the 70-80% band: the model predicts 75% bust there but the actual rate is 39% (36pp off, n=28). The middle of the curve is well-calibrated; trust the tails less.
Each bin shows: of all predictions Draftanomics made in that bust-probability range, what fraction actually became busts — produced less than 5 career Approximate Value. Y=X is perfect calibration.
Where the board diverged from draft slot
The sample: 11 draft classes graded, 5 with complete four-year windows— off-slot outcomes are scored on the complete windows only. The number: hits + saves − stretches − misses, so positive means the board's disagreements with the draft aged well on balance.