The Ombudsman Goes First: Grading A College Basketball Model That Has Not Made A Pick Yet
Starting in November this site will publish a predictive model for college basketball, betting spreads, totals and the occasional moneyline on a paper bankroll, with every pick posted before tip-off and graded in public. Before any of that happens, the case against it.
Somebody has to be the villain, so it may as well be this column.
Starting in November, this section will publish a college basketball model. It will project a margin and a total for every game it can see, compare those numbers against what the market is offering, and post picks on the spread, the total, and occasionally the moneyline. Every pick goes up before tip-off with the line and the price we actually had. Every pick gets graded. The record and the return live on the site whether they flatter anybody or not.
And every week, this column will follow the recap and argue that none of it means anything.
Ombudsman columns normally run second, because normally there is something to review. This one runs first, and the reason is not modesty. It is that the case against this project can be made in full right now, before a single number has been generated, using nothing but arithmetic that has been sitting in public for decades.
So: the prosecution's opening statement.
Exhibit one, the tax#
A standard college basketball bet is priced at -110. You risk $110 to win $100. That is the house's cut, and it is the first fact anybody building one of these should have tattooed somewhere visible.
To break even at -110 you have to win 52.38 percent of your bets. Not 50. A coin flipper, someone with genuinely zero predictive skill who picks sides at random, does not go home even. He loses 4.55 cents on every dollar he risks, forever, with perfect consistency, until he runs out of dollars.
Now look at the top of the range. A model that hits 54 percent, which would be a genuinely good model and better than most things sold on the internet under the word "expert," returns about 3.1 cents on the dollar. Fifty-five percent gets you a nickel.
That is the whole business. The distance between a coin and a legitimate edge is a bit over three percentage points of win rate, and the tax sits right in the middle of it. There is no version of this where the model is a little bit good and it works out fine.
Exhibit two, small samples lie, and they lie in the flattering direction#
Here is the part that should worry a reader more than the vig, because the vig is at least honest about what it is doing.
Suppose the model has no skill at all. Pure coin. Run it for 100 bets, which is about three weeks of a full college schedule, and there is roughly a one in three chance it finishes that stretch above the breakeven line, looking for all the world like it has found something.
How often a coin flipper looks like a genius
Probability that a bettor with zero skill finishes a stretch above the 52.38 percent breakeven line at -110, by number of bets graded. Small samples do not just fail to prove an edge. They actively manufacture the appearance of one.
Read that chart in both directions, because it cuts both ways and the second direction is the one that gets people fired.
Going down the chart, the noise is the danger. A hot November means nothing. A hot December means nearly nothing. Anybody who tells you about a 60 percent start over 80 picks is telling you about weather.
Going the other way, a real edge is fragile too. A model that is truly 54 percent, a legitimately good one, still finishes a full 1,000-bet season below breakeven about 15 percent of the time. One winter in seven, the good model looks broken. That is the exact scenario in which a builder panics, retunes something that was working, and destroys it.
This column's most tedious job, and it will get tedious, is repeating those two sentences until March.
Exhibit three, the goal is a 75-game sample#
Here is the structural problem at the center of the whole enterprise, and it is worth stating plainly because it is not the kind of thing a project usually admits about itself.
The stated goal of this model is the NCAA Tournament. The regular season, more than 5,000 Division I games, exists to test and calibrate. March is the point.
The 2027 tournament is 75 games.
Seventy-five. That is a sample so small that a model with a real, genuine, hard-won 54 percent edge would expect to win about 40 of them and lose about 35, an outcome completely indistinguishable from a coin. Whatever happens in March, whether the model goes 45-30 or 30-45, will prove nothing whatsoever about whether the model works.
So the goal of this project is, by construction, a test the project cannot pass or fail.
There is only one honest response to that, and it happens to be the reason the regular season matters so much: the burden of proof has to be discharged before Selection Sunday, or it never gets discharged at all. By March we should already know what this thing is, from a thousand-plus graded picks in November, December, January and February. March is not the exam. March is the thing you spend the credibility on.
If anyone starts describing a good tournament as validation, that is the week this column gets loud.
Exhibit four, the opponent is not the bookmaker#
The last thing worth being pessimistic about in advance is the competition, which is stiffer than the framing usually suggests.
This model will be built on Ken Pomeroy's efficiency data, which is a subscription that thousands of people hold, and which is itself already an excellent predictive model. That matters more than it sounds like it does. KenPom is not a raw ingredient we transform into an edge. It is a very good competitor whose output is partly baked into the market number already. If this model turns out to be KenPom with a small twist on top, it will beat neither KenPom nor the closing line, and the honest thing will be to say so in print.
Which points at the real scoreboard, and it is not the win-loss record.
The closing line, the final number a game trades at before tip, is the best public predictor of a basketball game that exists. It contains every model, every injury report, and every dollar of sharp money. So the question that actually determines whether this project has found anything is not "did we win," it is "did we beat the close." That is closing line value, and it stabilizes far faster than profit does. A model that consistently beats the closing number and still loses money is unlucky and will come around. A model that loses to the closing number and still makes money is lucky and will not.
Both numbers get published every week. When they disagree, this column will tell you which one to believe, and it will usually be the unflattering one.
The eight extra teams#
One genuine wrinkle, scoped honestly rather than dramatically.
2027 is the first 76-team NCAA Tournament, up from 68. There will be 32 automatic bids and 44 at-large bids, and the Opening Round triples from four games to twelve, with the 12 lowest-seeded automatic qualifiers and the 12 lowest-seeded at-large teams playing their way into the main 64-team bracket.
For bracket projection, that is enormous. Every bubble model in existence was fit on a 37-at-large field. Seven more at-large bids drag the cut line meaningfully down the resume board, and anybody forecasting the 2027 field off historical bubble behavior is forecasting a tournament that will not be played.
For betting individual games, it is a smaller deal than the headlines suggest, and you should be suspicious of anyone who says otherwise. A 9-seed is a 9-seed regardless of how many teams got in behind it. What actually changes is narrower and more interesting: twelve extra Opening Round games in thin markets involving teams nobody has watched, 24 teams carrying an extra game of fatigue into the round of 64, and more mismatches in the main draw, which pushes spreads into wide ranges where any model's error profile is least tested.
That is the honest version. It is less exciting than "expansion breaks everything," and it has the advantage of being true.
The standard of proof#
The defense, such as it is, comes down to what gets published rather than what gets claimed. Six commitments, on the record, so that this column has something to hold the model against later:
Picks go up before tip-off, with the line and the price actually available to us, never the best number in the market and never after the fact. A slate that gets missed is logged as missed and never backfilled.
The full ledger is public, not the highlights. Every bet, every price, every result, including the ugly ones.
Losing weeks get the same word count as winning weeks. The failure mode of every model that has ever been published is that the recaps get shorter in February.
Model changes are versioned, dated and explained before the next pick, not quietly folded in afterward. If a parameter moves, the reason and the evidence go up with it, along with a stated condition that would make us change it back.
This is paper money. A $1,000,000 notional bankroll, one unit equal to $10,000, and not a dollar of it real. Worth saying now rather than in a footnote: a $10,000 bet is considerably more than a Tuesday night Sun Belt total would actually absorb. The unit is a scoring device, not a claim about achievable profit, and any dollar figure that ever appears on this site carries that caveat with it.
This column runs every week, and its job is to argue the model is worthless. Not to provide balance. To prosecute.
Most projects like this one fail. The arithmetic above is not pessimism, it is just the arithmetic, and it says the honest outcome is more likely to be "we could not beat the closing line" than anything anyone would want to put in a headline.
That result will get published too, in full, in this column, at the same length.
Which is the only part of any of this that is actually guaranteed to work.