I can’t possibly have the answer to how vulnerable the NPI is right now, but the following list shares some concerns I do have and the reasons for them. These are similar to items I have addressed to varying degrees in producing the T100 “ratings system,” while remaining mindful the NPI is essentially a “rewards system.” One difference is that rating systems bake in some subjective residual data related to the end of the previous season and pre-season rankings. This provides an origin point before any games have been played with something other than a “clean slate.” I recognize why most might think a rewards system shouldn’t do that to be fair, even if the practice has been proven to have predictive value, the reason the best analytics in the world typically do it. There’s the rub. The NPI doesn’t purport to be able to predict who would defeat whom, nor will it likely compete well with any system that does, even just for this one reason alone. (Also, adding to why you won’t see it published for at least another 3 weeks.) However, should any item on the list below, or those not conceived as of yet, turn out to be conspicuous, it would erode credibility of a rewards system sooner, and to a higher degree than any ratings system for that very reason. Not unlike how a virus would affect a person with a weakened immune system more than one whose isn’t. This a consequence of the smallish sample size even a season’s worth of matches produces.
Over-Rated Win%?
I know the quintessential “average” team in America in 2024, ranked 62nd out of 124 teams, went 18-7. (Inside Hitter & Massey had them at 54th & 65th so it wasn’t just the T100 making this claim.) The answer to a single question I asked and was answered, “It came in at NPI #48 in 2024.” i.e. Ranked 20% better than the mean of three credible models. Its signature win was against a team ranked in the high 60’s. Who has a signature win against a team ranked lesser than itself across 25 matches played? The NPI’s #48 in 2024, apparently. However, the NPI doesn’t need to be accurate in rating the middle of the landscape, though it sure would provide some confidence if it did. It needs to be most accurate among the top 10% for getting at-large bids right, and maybe just decent in the top 25% for reasonable seeding accuracy. So does this unreliable #48 really matter? Consider this.
When looking at the bubble of Fall Sports’ NPI metrics I notice a preponderance of those who got in on the NPI bubble, compared to the one’s Massey would have indicated as more worthy, held better W-L records. This fact, together with NPI metrics of these bubble teams not being remotely far enough apart to be statistically significant, though understandable, is concerning because Massey’s algorithm has long been one of the gold standards in rating algorithms related to sports team strength valuation. It begs the question if the NPI SOS component has as much “correction power” as we are led to believe by the max setting on its dial? Certainly, if its SOS is driven by the rankings across the whole landscape, and too many of those in the middle 50% turn out to be unreliable, like #48 described above, then it would be left trying to make lemonade from lemons with its SOS construct. The argument goes something like this:
If win-rate propagates errors in the ratings of a large enough subset of teams throughout the middle of the landscape, then the teams who played them would then propagate that error even more via their metrics, and then the ones who played them would do the same again and again until a compounding effect might create some unintended outcomes. An economist would call this a “multiplier effect,” but outcomes put in place by them are usually intentional.
Credibility for a close call too often falling to the win-rate as the difference maker while the dial is pegged to the limit of 20%W/80%SOS would give me pause for concern. It suggests impotency to skillfully differentiate the bubble, something I never believed afflicted our committee in year’s past. EVER!
The above represented a litmus test for the middle of the landscape. A team with a 72%-win-rate who was the most average team in America. The best litmus test I can think to demonstrate if win rate is overrated by the NPI at the high end of the spectrum would be back-testing to see where Loras would have been ranked had it lost the rubber match with Carthage in last year’s CCIW Final at home. We know based on protocol Carthage would have earned an automatic bid, forcing a two-loss Loras team to probably be one of the at-large bids offered to Juniata and Vassar. Knowing the retrospective on its back-testing last year “Would not have altered the committee’s choices” then, it follows that Carthage’s 2nd loss in 3 games against Loras did in fact keep their retrospective NPI metric below those of Juniata’s and Vassar’s. If Carthage wins that rubber, instead, does the NPI have the juice to rank higher, a 5-loss Carthage CCIW Champion with 2 out of 3 wins against a 2-loss Loras team, both whose losses would have been to Carthage? That might be all any of us would need to know about NPI’s ability to differentiate teams among the top 10? i.e. Would the NPI have prioritized a conference champion winning 2 of 3 head-to-head during the season with 5 losses on its resume to the one it defeated having just two, both to that same 5 loss team?
If the day ever comes when the NPI rates the 2nd best team in conference standings, who lost to the winner head-to-head, claiming it as the more deserving at-large candidate when neither won their tournament, nor played each other in it, I do wonder how that will be received by those paying attention.
Minimal Quality Win(s) plus Inflated Win% Synergy?
I have concern just 1 “Quality Win” for a team who competes in an 8 to 12 team conference with only two superiors to the rest of a mediocre field could become inflated. This might force two undesired outcomes. A less deserving at-large bid could be awarded and those good teams on the bubble playing it will have garnered an artificial boost toward that end. (Playing both would offer two times that boost!)
Case Study: A & B each go 17-1 in a 10-team conference home/away, and then each win one of just 3 opportunities playing a top 12. Kudos to them! Throw in another half dozen no-threat wins against teams nearer #40, possibly like the one rated #48 indicated above which really wasn’t, and the loser of their conference tournament is 24-4 with one signature win against a top 12.
I fear this would be an NPI at-large candidate when the T100, Inside Hitter, Massey, and/or any human poll might not rate the conference championship loser among the best 20. They’d probably be correct about that. Even should it end up just on the outside looking in, those top 12 teams who defeated them will have been provided an artificial boost to their metric. The margins on the bubble of the NPI in all sports have been so small to date, it could be the difference maker more often than not. This is troublesome if it were to happen.
Double or Triple Win Fallacy?
Is it really new information when Team A defeats Team B a second time in the same season? Whatever leverage obtained for the first of these is then doubled as if it were 2 wins over different teams ranked identical in every way? Shouldn’t there be a law of diminishing returns so the second time it counts proportionally less, and then the 3rd time even less than that, if at all. (Those who believe the third ought to not count at all would be hard pressed to say the 2nd should count as much because of contradictory transitive logic!)
Sometimes competition goes beyond the strength of an opponent relative to everybody else in the landscape and falls to familiarity amongst coaches and tendencies between teams. Two wins by an inferior ranked team against a better ranked opponent isn’t the same as having defeated two different opponents each ranked better and of the same ilk.
For Darwinian Natural Selection to properly take place, you shouldn’t have arranged marriages between the same royal bloodlines. Something Europe even realized about a quarter century after he published “On the Origin of Species.” It can create unintended outcomes more often than you’d like. Not exactly the same, but I couldn’t resist! LOL
Strength of Wins & the 3-2 Dilemma?
Strength of Wins is key to the NPI, and this is a good thing. Last year alone there were teams who would win a few games against the later “teens” only to go 0-9 against the rest of the top 20. To be honest, that is what the 17th best team in the country typically does – Should do. The last thing anybody wants to see is that team in the top 10 of the former SOS (RIP-RPI) list at the end of the year because they played so many good teams yet performed exactly as they should have under the circumstances. SOS is slippery for this reason. More so than SOW (Strength of Wins). The good news is Wins are King in the NPI model. It will continue to count them past 13 as long as they help a team’s metric. They stop counting the moment any win doesn’t help it! This makes it impossible for a win to hurt a team when it comes against a far inferior opponent. Something SOS failed to do in the past when a far inferior opponent would artificially skew its SOS to a lesser value! (However, don’t forget the quality of those wins and the 13 already baked in is largely determined by their opponents’ metrics, which as discussed above might already be of concern for other reasons.)
Coach-speak will tell you a win is a win, and in most ways, like in the standings, that is true. Is a set to 15 exactly like any other? The chances of a 5th set differing by 2 is 50% more likely than when it’s a game to 25. Is that because of the condition the teams have already won two each, showing they are well matched, or is it purely numerical because it is a game to 15? Should a team win all 5 of its Cinco matches, would it be just the steely eyed grit that allowed it to happen every time?
I can live with a team being the first one out having lost all five 5th sets. There is something to be said about perseverance & grit. I can certainly live with a team having lost all five 5th sets, being the last one in, too. Hell, give them a medal for that! What I’d prefer not to live with is a team being the last one in the tournament because they won all five 5th sets. That’s not just perseverance no matter what “coach speak” says!
I think an 80%-20% split for Cinco’s makes sense. They only happen at a rate about 1 in 7 matches and I would be willing to concede a team going 5-0 in Cinco’s is probably on par with another going 4-1 having matches against the same five all end up with a 3-1 score. (In the T100 model a 3-1 victory is scaled to 90% leverage, not the proportion of 3-1. i.e. 75%. In a 3-2 match I slice it back a little more to 75% rather than its 60% set ratio. – My original thought was to not differentiate 3-0 & 3-1 because too often they’re not significantly different in their point scoring, but I never believed a 3-2 wasn’t deserving to be treated differently, and I still don’t.)
I will stop right here with 4 items on the list so far. There are still a couple more that need some refinement. I will offer those in the next post. Look for that a little later.
