Bracketology 3/7

I’ve added a new post on my method which, if you have the patience to read it, will help shed some light on how I arrived at the lists below.  To clarify, the list on the left is what my model says the committee SHOULD do; the list on the right is what my adjusted model says the committee WILL do.

Should Be:

#1 seeds: Virginia, Duke, UNC, Michigan

#2 seeds: Tennessee, Michigan State, Gonzaga, Kentucky

#3 seeds: LSU, Kansas, Texas Tech, Purdue

#4 seeds: Houston, Florida State, Kansas State, Buffalo

#5 seeds: Virginia Tech, Nevada, Wisconsin, Cincinnati

#6 seeds: Wofford, Maryland, Marquette, Washington

#7 seeds: Iowa, Mississippi State, Villanova, Auburn

#8 seeds: Oklahoma, VCU, Belmont, UCF

#9 seeds: Iowa State, Utah State, Louisville, New Mexico State

#10 seeds: Minnesota, UC Irvine, Toledo, NC State

#11 seeds: Furman, Murray State, Ohio State, Lipscomb

#12 seeds: Temple, Syracuse, St. John’s, Baylor, Clemson, Ole Miss

#13 seeds: Arizona State, Vermont, Liberty, Hofstra

#14 seeds: Old Dominion, Drake, Georgia State, Harvard

#15 seeds: Montana, South Dakota State, Northern Kentucky, Radford

#16 seeds: Colgate, Sam Houston State, Prairie View A&M,
Norfolk State, Fairleigh Dickinson, Iona

Last Four Byes: Temple, Syracuse, St. John’s, Baylor

Last Four In: Clemson, Ole Miss, Arizona State, Liberty

First Four Out: Florida, Seton Hall, East Tennessee State, Indiana

Next Four Out: TCU, Davidson, Texas, Alabama

Will Be:

#1 seeds: Virginia, Duke, UNC, Michigan St.

#2 seeds: Michigan, Kentucky, Tennessee, LSU

#3 seeds: Gonzaga, Kansas, Purdue, Kansas St.

#4 seeds: Texas Tech, Florida State, Wisconsin, Houston

#5 seeds: Buffalo, Virginia Tech, Marquette, Nevada

#6 seeds: Maryland, Cincinnati, Mississippi State, Iowa

#7 seeds: Washington, Villanova, Oklahoma, Wofford

#8 seeds: Iowa State, Auburn, Louisville, VCU

#9 seeds: Minnesota, UCF, Utah State, Baylor

#10 seeds: Syracuse, Ohio State, Belmont, New Mexico State

#11 seeds: St. John’s, Toledo, Arizona State, Ole Miss

#12 seeds: NC State, Seton Hall, Temple, Clemson, Florida, Texas

#13 seeds: UC Irvine, Lipscomb, Vermont, Old Dominion

#14 seeds: Georgia State, Drake, Hofstra, Harvard

#15 seeds: Radford, Montana, South Dakota State, Northern Kentucky

#16 seeds: Colgate, Sam Houston State, Prairie View A&M,
Norfolk State, Fairleigh Dickinson, Iona

Last Four Byes: Arizona State, Ole Miss, NC State, Seton Hall

Last Four In: Temple, Clemson, Florida, Texas

First Four Out: Furman, Indiana, Alabama, TCU

Next Four Out: Creighton, Murray State, Davidson, Liberty

Bracketology – More on Method

I started on this Bracketology project last year. My initial goal, and the goal that still interests me most, was to come up with a purely quantitative, objective model for selecting at-large teams that completely removes subjectivity from the process. Such a model would be based on the following principles that are (to me, at least) self-evident:

  1. The goal is not to select the “best” teams; the goal is to select the teams that have earned spots in the tournament based on their record of wins and losses
  2. Wins always help, and losses always hurt
  3. Wins against good teams help more than wins against bad teams, and losses against bad teams hurt more than losses against good teams
  4. Wins and losses are all that matter (margin of victory is irrelevant)
  5. Timing of games is irrelevant (the first game of the season is just as important as the last)

If these principles are not self-evident to you, it’s probably because you’re not understanding or agreeing with #1 above. To illustrate what I mean, let’s take an example. I think Penn State is a better team than Seton Hall this year. If they were to play tomorrow on a neutral court, I would pick Penn State to win. Penn State is ranked #41 at kenpom.com, while Seton Hall is ranked #60. But Penn State is 13-17. They haven’t earned a spot in the tournament by winning enough games. Now, you might say, if they’re so good, why are they 13-17? Good question. For one thing, they played the second toughest schedule in the country. On top of that, they had some bad luck. In games decided by 5 points or less or in overtime, they went 2-8. Perhaps it’s more than bad luck; I’m sure Penn State fans have their theories. But whatever the reason, the bottom line is that they didn’t win enough games to make the tournament, regardless of how “good” they are.

So, building on that foundation, I created a model with the following characteristics:

  • Every game yields a score from 0-100. Wins always yield positive scores, while losses always yield negative scores.
  • Wins against really good teams yield scores close to 100. Wins against really bad teams yield scores close to 0.
  • Losses against really good teams yield negative scores close to 0. Losses against really bad teams yield negative scores close to -100.
  • A team then becomes the sum of its game scores.

That’s it. Total up the game scores for each team, and select and seed the teams according to their scores. The only tricky part is, how exactly to come up with the game score. After thinking about it for a while, the solution hit me: the game score should be the probability that an NCAA Tournament team would win (or lose) that game.

For an illustration, let’s take Alabama. On November 6, they beat Southern at home. Southern is ranked #341 in kenpom. NCAA tournament teams beat #341 at home 99.9% of the time. Therefore, Alabama gets only (100-99.9)=0.1 points for that win.

On December 4, Alabama lost to Georgia State at home. Georgia State is ranked #131 in kenpom. NCAA Tournament teams beat #131 at home 90.3% of the time. So for losing that game, Alabama gets -90.3 points.

Moving to the good side of the ledger, Alabama won at Missouri on January 16. Missouri is ranked #88 in kenpom. NCAA Tournament teams, playing on the road at #88, win 54.1% of the time. Therefore Alabama gets (100-54.1) = 45.9 points for that win.

So using a probabilistic model solves the problem of rewarding teams more for beating good teams, and punishing them more for losing to bad teams. It also provides a way to adjust for home vs. away games – a factor which, incidentally, is not considered enough in evaluating teams’ resumes.

So I built this whole spreadsheet-based model to do exactly what I’ve just explained for all the teams worthy of tournament consideration last year. And once I finished it, it was immediately clear that the logic used by my model is not the logic used by the selection committee. While the model does an fine job of selecting the top seeds in the tournament, when it comes to the bubble, it consistently selects small conference teams with good records over big conference teams with relatively poor records. Is that a flaw in my logic, or a flaw in the selection committee’s criteria? That’s for you to decide; but it did lead me to a few conclusions about the committee’s thinking that seem evident from comparing my model to their actual selection.

  1. Wins help more than losses hurt.
  2. Teams are given very little credit for winning games, especially road games, against teams ranked 75-150.

To elaborate on #1. According to my model, if you lose to #92 on a neutral court (-70 points) and then beat #27 on the road (70 points), those 2 games basically cancel each other out, and would have the same net effect as if you beat #349 and #350 (each of which would yield ~0 points). But I think in the committee’s eyes, the first team has more merit. There is some credit given for playing better teams, even when you lose.

To elaborate on #2. This hurts a few small conference teams every year. Let’s take Toledo this year. They’re not even being talked about for an at-large bid by Lunardi or Palm, presumably because they “didn’t beat anybody”. But actually, they beat #122, #123, #104, #147, #135, and #179 on the road, and #91 on a neutral court. Are those gimme games? Hardly. The win probabilities for those games for NCAA Tournament teams are 63%, 64%, 59%, 69%, 67%, 76%, and 70%. Those teams are basically Georgia Tech this year. Is it easy to go win at Georgia Tech? 7 times?

So, I’ve started trying to tweak my model to account for the biases of the selection committee. I want to see if I can adjust the model over time to get closer and closer to the committee’s logic, but still in a purely quantitative model with no subjectivity. The current version of the model weights wins twice as heavily as losses. I found that by doing this, I come much closer to the actual selections in past years.

So in my future updates, I’m going to give you 2 lists. One is what I think the committee should do, based on the guiding principles described above; the other is what my adjusted model says they will do. I hope to learn enough to continue improving my model to conform more to the committee’s method.

Bracketology Intro

I’ve decided to try my hand at Bracketology – my way.

I’ve been interested in the subject for years. Last year, I created an Excel-driven method for figuring out what the committee should do, not what they would do. And I still have that formula, but I’ve tweaked it a bit to account for the committee’s actual demonstrated behavior, and I think I’ve created a pretty good formula-driven replica of how they assess the teams. Now it’s time to put her in the water and see if she floats.

What’s unique, I guess, about my approach is that no human judgment is applied to my picks. It’s purely formula-driven. Of course it’s risky to try to replicate human behavior with a formula, because it assumes that human beings behave predictably and rationally, which is demonstrably false. But I still am curious as to how close I can come to replicating the logic of the committee.

I’ll try to post an update every day and provide some commentary on what changed and why, and explain some of my more idiosyncratic picks.

When all is said and done, I’ll compare my results to Joe Lunardi and Jerry Palm, and we’ll see who wins when spreadsheet faces human. Although for all I know, those guys are using a spreadsheet too.