First Win · Publication 001

MLB Moneyline Calibration & Returns, 2021-2025

Every moneyline in the corpus for 20212025, compared against what actually happened.

Bets analysed
23,992
event-selections
Market expected
50.00%
no-vig consensus
Actually won
50.00%
realized
Return on risk
-3.55%
-852.46u total
Vig paid
+1.87 pp
per side

Caveman Analysis

The question. Market puts a number on every game. Is the number right? And if you bet that number every time, do you finish up or down?

Finding one: the number is right. Market says 60%, team wins 60%. 9 buckets, 9 matches, worst one off by +2.06 pp. Number honest. Good.

Finding two: you still finish down. All 10 price buckets lost money. Down 852 units across 23,992 bets. Number honest, bettor still loses. Bad.

Why both are true. You do not get to bet the honest number. You bet a worse one. The gap is +1.87 pp per side, every single time. Small gap, huge volume, money gone. Fee go up, account go down.

Bottom line. Market right, good. Market still takes its cut, bad. Honest is not the same as free.

Implications

On the two pooled probabilities above: this sample contains both sides of every game, so the pooled expectation and the pooled win rate are forced to agree exactly. That agreement is an arithmetic identity, not a measurement, and its interval is zero-width. Calibration evidence comes from the probability bands, which split at 50% and so contain no mirrored pairs.

Executive summary

We took every MLB regular-season moneyline we hold from 2021 through 2025, 23,992 bets across roughly 11,996 games, and asked the two questions a serious bettor should want answered before any others. When the market says a team wins 60% of the time, does it? And what would you have earned betting that price every time?

The market is remarkably well calibrated. Across 9 probability bands, the largest gap between expectation and reality is +2.06 pp, and every single band’s confidence interval contains zero. We cannot distinguish the closing MLB moneyline from a perfectly calibrated forecast.

And every price bucket still lost money. 10 predetermined buckets; 10 returns that are negative or indistinguishable from zero. Pooled across everything: -3.55% per unit risked, or -852.46u over 23,992 bets.

Those two facts are not in tension. They are the whole point. The market’s probability estimate is right; you just don’t get to bet at it. Closing prices imply 51.87% for each side of a two-way market, so the two sides sum to about 103.7% rather than 100%. That +1.87 pp per side is the house’s fee. Calibration tells you the market is honest. It does not make it free.

What does −150 actually mean?

A price of −150 means risking 150 to win 100. Converted directly, it implies a 60.0% win probability.

But the other side of that same game is not priced at +150. At this sample’s measured margin it sits nearer +129 (43.7%). Add the two together and you get roughly 103.7%, not 100%. The excess is the vig, baked into both prices.

So there are three different numbers, and conflating them is where most public analysis goes wrong:

ConceptWhat it isAt −150
Executable oddsthe price you can actually bet−150
Raw implied probabilitythat price converted naively; includes vig60.0%
No-vig probabilitythe market’s actual estimate, margin removed57.8%

This publication keeps all three separate. Calibration is tested against the no-vig estimate, because that is the market’s genuine forecast. Returns are computed against executable odds, because that is what you would have transacted at. Using the no-vig price for both would manufacture a profit no bettor could ever collect.

Definitions

Three numbers get confused constantly in public betting analysis, and most bad conclusions come from mixing them up. These are the terms used on this page.

Executable odds (quoted price)
The American price you could actually have bet at. Risking 150 to win 100 is −150. Every return here is computed against these prices, because they are what a bettor transacts at.
Raw implied probability
An executable price converted straight to a probability with no adjustment. −150 becomes 60.0%. The bookmaker’s margin is still inside it, so the two sides of a game sum to more than 100% and this number always overstates the true chance.
Vig (also “juice”, “overround”)
The bookmaker’s margin, built into both sides of the market. Measured as how far the two sides’ raw implied probabilities exceed 100%. In this sample they sum to about 103.7%, an overround of 3.7%.
Vig paid per side
The same margin for a single bet: how far its raw implied probability sits above its no-vig probability, in percentage points. Sample average +1.87 pp. This is the fee you pay on every bet, win or lose.
No-vig probability, shown as “Market expected”
The market’s genuine estimate once the margin is stripped out. Each book’s own two-way prices are de-vigged proportionally, then eligible books are equal-weighted. Calibration is tested against this, because it is what the market actually believes. It is never used to compute returns.
Actual win rate
The share of selections in a group that won: wins divided by n.
Calibration gap (residual)
Actual win rate minus market-expected probability, in percentage points. Positive means the selection won more often than the market expected. Zero is perfect calibration.
Return on risk (ROI)
Profit or loss per unit risked, betting one unit on every selection in the group at its executable price. -3.55% means losing 3.55 units per 100 risked.
Units, total units
Cumulative profit or loss in bet units, one unit risked per selection. -852.46u means the strategy finished that far down.
Event-selection
One side of one game, and the unit of analysis here. A game contributes two, home and away. Roughly 11,996 games gives 23,992 event-selections.
Closing price
The latest complete, eligible pre-game quote no more than 30 minutes before scheduled first pitch. The market’s final and best-informed number.
Consensus
The equal-weighted average across eligible books, taken after each book’s own prices have been de-vigged individually.
Odds bucket
A fixed range of executable prices. Ranges were frozen before any result was computed.
Probability band
A fixed range of no-vig probabilities. Bands split at 50%, so a price and its complement can never land in the same band.
95% confidence interval
A range expressing sampling uncertainty. All three come from the same deterministic event-clustered bootstrap, with the game as the resampling unit so two selections from one game keep their dependence. An interval containing zero means the effect cannot be distinguished from no effect at this sample size.
“Spans 0”
A tag marking a calibration gap whose interval contains zero: not distinguishable from perfect calibration.
“Seasons +”
How many individual seasons showed a positive value, out of those reported: return for odds buckets, residual for probability bands. A check on whether a pooled result rests on one unusual season.

How this was measured

Every number derives from one definition, frozen before any result was computed. The spec hash is in the provenance panel at the foot of this page.

Calibration: the market is honest

Each point below is one probability band: where the market said it should land, against where it actually landed. Perfect calibration is the dashed diagonal.

BucketnAvg priceExpectedActualGap95% CI on gapROISeasons +
<35%1,753+21331.16%29.49%-1.67 pp-3.9% to +0.5%spans 0-8.59%1/5
35-40%2,454+15637.77%39.20%+1.43 pp-0.5% to +3.3%spans 0+0.22%3/5
40-45%3,647+12742.60%43.13%+0.53 pp-1.1% to +2.1%spans 0-2.15%3/5
45-50%4,142+10347.52%47.47%-0.05 pp-1.6% to +1.4%spans 0-3.84%2/5
50-55%4,142-12052.48%52.53%+0.05 pp-1.4% to +1.6%spans 0-3.62%3/5
55-60%3,647-14757.40%56.87%-0.53 pp-2.1% to +1.1%spans 0-4.32%2/5
60-65%2,454-18162.23%60.80%-1.43 pp-3.3% to +0.5%spans 0-5.71%2/5
65-70%1,209-23167.25%69.31%+2.06 pp-0.5% to +4.6%spans 0-0.67%4/5
70%+544-30372.37%73.16%+0.79 pp-3.1% to +4.5%spans 0-2.61%3/5

All 9 intervals contain zero. The largest apparent miss, 65-70% at +2.06 pp, is consistent with noise at this sample size. And no band’s residual holds its sign across every season: if a band were genuinely mispriced, that would show up year after year. It does not.

“Seasons +” counts seasons with a positive residual. A gap whose interval spans zero is not distinguishable from perfect calibration.

Returns by moneyline bucket

BucketnAvg priceExpectedActualGap95% CI on gapROISeasons +
+201 or longer928+23529.04%27.16%-1.89 pp-4.6% to +1.0%spans 0-9.40%1/5
+151 to +2002,318+17235.66%36.07%+0.40 pp-1.6% to +2.3%spans 0-2.56%1/5
+126 to +1502,828+13840.64%42.26%+1.61 pp-0.2% to +3.5%spans 0+0.42%2/5
+101 to +1254,098+11345.32%44.31%-1.00 pp-2.5% to +0.5%spans 0-5.69%0/5
even to -1204,198-11050.32%50.52%+0.21 pp-0.3% to +0.7%spans 0-3.29%0/5
-121 to -1402,626-13054.46%56.02%+1.56 pp-0.4% to +3.5%spans 0-0.80%2/5
-141 to -1602,297-14857.63%56.73%-0.90 pp-2.9% to +1.2%spans 0-5.01%1/5
-161 to -2002,782-17761.62%59.92%-1.70 pp-3.5% to +0.2%spans 0-6.18%0/5
-201 to -2501,211-22466.60%68.13%+1.53 pp-1.1% to +4.2%spans 0-1.44%4/5
-251 or shorter706-29271.71%72.24%+0.53 pp-2.9% to +3.7%spans 0-2.98%2/5

One bucket shows a positive pooled return: +126 to +150, at +0.42%. Its interval runs from -3.9% to +4.8%, and it was profitable in only 2 of 5 seasons. That is not an edge. That is what noise looks like when you slice 23,992 bets 10 ways.

The worst place to be was +201 or longer, at -9.40%, the one result with real economic size, and also one of the widest intervals, on only 928 bets.

Season stability: the discipline that kills most findings

This is the most important table here, and the one most public analysis omits.

Bucket20212022202320242025
+201 or longer-0.7%+5.6%-19.7%-11.0%-28.9%
+151 to +200-4.1%-11.4%+9.4%-3.2%-1.6%
+126 to +150-3.3%-2.9%+5.4%-5.1%+8.7%
+101 to +125-5.9%-9.8%-10.7%-2.0%-1.2%
even to -120-2.5%-4.3%-2.6%-3.5%-3.7%
-121 to -140-4.4%+2.9%+5.1%-0.9%-6.4%
-141 to -160+4.2%-0.9%-7.3%-12.1%-8.4%
-161 to -200-7.3%-2.6%-8.9%-2.0%-10.8%
-201 to -250+0.8%+0.7%-11.1%+1.3%+1.9%
-251 or shorter-8.5%-6.2%+3.9%-3.4%+1.0%

Look at -201 to -250. It was profitable in 4 of 5 seasons. A tout would stop reading there. But 2023 returned -11.1%, and the pooled result is -1.44%. One season erased the rest.

Only 3 buckets hold a consistent sign across every season, and all 3 are consistently negative: +101 to +125, even to -120, -161 to -200. Not one bucket is reliably profitable; every apparently positive one flips sign at least once.

Return per unit risked, by bucket and season. Shading tracks magnitude; the value is always printed, so nothing is conveyed by colour alone.

Why a calibrated market still loses you money

This is the part worth internalising. Across the whole sample the average no-vig expectation is 50.00% and the average realized win rate is 50.00%. Those two agree by construction rather than by measurement: this sample holds both sides of every game, and a home residual is exactly minus the away residual, so the pooled figures cancel to an arithmetic identity whose interval has zero width. The calibration evidence is the band table above, where nothing cancels.

But you cannot bet at that number. You bet at prices implying 51.87%. The difference, +1.87 pp per side, a 3.7% overround on the two-way market, is the fee. Our pooled return of -3.55% is close to what you would expect from paying it on every bet. The house edge, not bad forecasting, takes the money.

The practical implication is uncomfortable and important: beating a market like this requires being more accurate than it by more than the vig. Being merely as accurate as the market is a slow, arithmetically certain loss. Any strategy that does not clear that bar is a fee-paying exercise, however sophisticated it looks.

Appendix: 2020

Reported separately and excluded from the headline sample: a shortened season on a thinner book consensus is a different market.

Bets
1,556
Expected
50.00%
Actual
50.00%
Return
-2.93%

Limitations

Ask your own question of this data

The Historical Explorer runs the same engine behind these figures over six seasons of closing prices. Pick a price range, a side and a season, and see what actually happened.

Open the Historical Explorer

Get new sports-market research when we publish

Occasional email, only when a new research note goes out. No picks, no promotions. Everything on this site stays free to read without an account.

First Win Pro

Go beyond the published finding

The public research answers what did we find. Pro answers what else the research showed, where the result breaks down, and how to interrogate the work yourself. It is the deeper layer behind the public studies, for people who want to use the research rather than only read the conclusion.

The public work stays complete. Nothing is held back from an article to manufacture a reason to subscribe.

What members get today

Research Packs
The full workings behind each major study: every supporting table, the robustness and sensitivity work, and the cuts too detailed for a public article.
Pro-only research notes
Focused quantitative investigations into questions that do not warrant a full publication but are still worth answering properly.
Extended analysis and robustness work
Where a result holds, where it weakens, and what happens to it under a different estimator or a different cut of the same data.
Downloadable research assets
Derived tables from the sealed research packages, where the underlying licence permits redistribution.
Early access
Completed studies reach members before they are published publicly.
Research queue input
A say in which questions get investigated next. Research topics, not requests for selections.

What we are building

The current membership is a research membership. The full First Win Pro research terminal is still being built, and founding membership is a claim on it rather than access to something that exists today.

Founding membership opens shortly and is not yet on sale, so there is nothing to buy today. The newsletter above will carry the announcement. No picks, no tips, and no claim that any of this will make you money betting.

Methodology, version and provenance
Research methodology0.2.0
Publication versionv1.0.1
Uncertainty methodevent-clustered bootstrap (game = resampling unit)
Publication spec hash4bdad3003d9fb6b9ffc9c581a1ce3c95b7e4b9f3fbc63b0b12a28652fdeab5a9
Parent spec versionv1
Parent spec hashd24e88a7de499c21b298fb46230c4e9f88f919fbbd59ebd26c6f2fe44369e07d
Fact build id5
Generated2026-08-24T03:27:26.726081+00:00
Primary seasons2021, 2022, 2023, 2024, 2025
Appendix seasons2020
Event-selections23,992

Intervals are 95% percentile intervals from a deterministic event-clustered bootstrap: 2,000 iterations, fixed seed, the game as resampling unit, so multiple selections from one game keep their observed dependence. Events are ordered canonically before the random stream is consumed, so an interval does not depend on input row order. The parent hash above is the original v1 freeze, retained as provenance; v1.0.1 changes only the uncertainty method, never a point estimate. These figures are a sealed artifact produced once against the fact build above. They are served read-only and are not recomputed on request, so a published number cannot drift when the corpus is rebuilt.