MLB Moneyline Calibration & Returns, 2021-2025
Every moneyline in the corpus for 2021–2025, compared against what actually happened.
Caveman Analysis
The question. Market puts a number on every game. Is the number right? And if you bet that number every time, do you finish up or down?
Finding one: the number is right. Market says 60%, team wins 60%. 9 buckets, 9 matches, worst one off by +2.06 pp. Number honest. Good.
Finding two: you still finish down. All 10 price buckets lost money. Down 852 units across 23,992 bets. Number honest, bettor still loses. Bad.
Why both are true. You do not get to bet the honest number. You bet a worse one. The gap is +1.87 pp per side, every single time. Small gap, huge volume, money gone. Fee go up, account go down.
Bottom line. Market right, good. Market still takes its cut, bad. Honest is not the same as free.
Implications
- Honest is not the same as free. A market can be telling you the truth about probability and still be a losing proposition, because you transact at a price that is worse than the truth. Accuracy and profitability are separate questions, and this data separates them cleanly.
- The bar is not “be right”. The bar is “be right by more than the fee”. Being exactly as accurate as the closing market is a guaranteed slow loss of roughly +1.87 pp per side. Any edge smaller than the vig is not an edge.
- Four good seasons out of five is not evidence. The -201 to -250 bucket won in 4 of 5 seasons and still lost 1.44% overall, because 2023 alone took 11.1%. Judging a strategy on how many seasons it won rather than on its pooled result is how people convince themselves noise is skill.
- The one profitable bucket is the trap, not the finding. If you go looking for the best row in a table of 10, you will find one. You will be finding randomness.
- This tells you where not to spend effort. If closing prices are this well calibrated, then beating the closing line by studying the closing line is not a promising direction. Whatever edge exists is somewhere this analysis does not look: earlier prices, better execution, or information the market has not yet absorbed.
On the two pooled probabilities above: this sample contains both sides of every game, so the pooled expectation and the pooled win rate are forced to agree exactly. That agreement is an arithmetic identity, not a measurement, and its interval is zero-width. Calibration evidence comes from the probability bands, which split at 50% and so contain no mirrored pairs.
Executive summary
We took every MLB regular-season moneyline we hold from 2021 through 2025, 23,992 bets across roughly 11,996 games, and asked the two questions a serious bettor should want answered before any others. When the market says a team wins 60% of the time, does it? And what would you have earned betting that price every time?
The market is remarkably well calibrated. Across 9 probability bands, the largest gap between expectation and reality is +2.06 pp, and every single band’s confidence interval contains zero. We cannot distinguish the closing MLB moneyline from a perfectly calibrated forecast.
And every price bucket still lost money. 10 predetermined buckets; 10 returns that are negative or indistinguishable from zero. Pooled across everything: -3.55% per unit risked, or -852.46u over 23,992 bets.
Those two facts are not in tension. They are the whole point. The market’s probability estimate is right; you just don’t get to bet at it. Closing prices imply 51.87% for each side of a two-way market, so the two sides sum to about 103.7% rather than 100%. That +1.87 pp per side is the house’s fee. Calibration tells you the market is honest. It does not make it free.
What does −150 actually mean?
A price of −150 means risking 150 to win 100. Converted directly, it implies a 60.0% win probability.
But the other side of that same game is not priced at +150. At this sample’s measured margin it sits nearer +129 (43.7%). Add the two together and you get roughly 103.7%, not 100%. The excess is the vig, baked into both prices.
So there are three different numbers, and conflating them is where most public analysis goes wrong:
| Concept | What it is | At −150 |
|---|---|---|
| Executable odds | the price you can actually bet | −150 |
| Raw implied probability | that price converted naively; includes vig | 60.0% |
| No-vig probability | the market’s actual estimate, margin removed | 57.8% |
This publication keeps all three separate. Calibration is tested against the no-vig estimate, because that is the market’s genuine forecast. Returns are computed against executable odds, because that is what you would have transacted at. Using the no-vig price for both would manufacture a profit no bettor could ever collect.
Definitions
Three numbers get confused constantly in public betting analysis, and most bad conclusions come from mixing them up. These are the terms used on this page.
- Executable odds (quoted price)
- The American price you could actually have bet at. Risking 150 to win 100 is −150. Every return here is computed against these prices, because they are what a bettor transacts at.
- Raw implied probability
- An executable price converted straight to a probability with no adjustment. −150 becomes 60.0%. The bookmaker’s margin is still inside it, so the two sides of a game sum to more than 100% and this number always overstates the true chance.
- Vig (also “juice”, “overround”)
- The bookmaker’s margin, built into both sides of the market. Measured as how far the two sides’ raw implied probabilities exceed 100%. In this sample they sum to about 103.7%, an overround of 3.7%.
- Vig paid per side
- The same margin for a single bet: how far its raw implied probability sits above its no-vig probability, in percentage points. Sample average +1.87 pp. This is the fee you pay on every bet, win or lose.
- No-vig probability, shown as “Market expected”
- The market’s genuine estimate once the margin is stripped out. Each book’s own two-way prices are de-vigged proportionally, then eligible books are equal-weighted. Calibration is tested against this, because it is what the market actually believes. It is never used to compute returns.
- Actual win rate
- The share of selections in a group that won: wins divided by n.
- Calibration gap (residual)
- Actual win rate minus market-expected probability, in percentage points. Positive means the selection won more often than the market expected. Zero is perfect calibration.
- Return on risk (ROI)
- Profit or loss per unit risked, betting one unit on every selection in the group at its executable price. -3.55% means losing 3.55 units per 100 risked.
- Units, total units
- Cumulative profit or loss in bet units, one unit risked per selection. -852.46u means the strategy finished that far down.
- Event-selection
- One side of one game, and the unit of analysis here. A game contributes two, home and away. Roughly 11,996 games gives 23,992 event-selections.
- Closing price
- The latest complete, eligible pre-game quote no more than 30 minutes before scheduled first pitch. The market’s final and best-informed number.
- Consensus
- The equal-weighted average across eligible books, taken after each book’s own prices have been de-vigged individually.
- Odds bucket
- A fixed range of executable prices. Ranges were frozen before any result was computed.
- Probability band
- A fixed range of no-vig probabilities. Bands split at 50%, so a price and its complement can never land in the same band.
- 95% confidence interval
- A range expressing sampling uncertainty. All three come from the same deterministic event-clustered bootstrap, with the game as the resampling unit so two selections from one game keep their dependence. An interval containing zero means the effect cannot be distinguished from no effect at this sample size.
- “Spans 0”
- A tag marking a calibration gap whose interval contains zero: not distinguishable from perfect calibration.
- “Seasons +”
- How many individual seasons showed a positive value, out of those reported: return for odds buckets, residual for probability bands. A check on whether a pooled result rests on one unusual season.
How this was measured
Every number derives from one definition, frozen before any result was computed. The spec hash is in the provenance panel at the foot of this page.
- Closing price. The latest complete, eligible pre-game bookmaker observation no more than 30 minutes before scheduled first pitch. A book’s stale quote is not a closing price.
- Consensus. Each book’s own two-way prices are de-vigged first, then eligible books are equal-weighted. Order matters: averaging vigged prices and de-vigging afterwards bakes in the spread of book margins.
- Sample. 2021, 2022, 2023, 2024, 2025 regular season, all eligible books, typically 8–10 per game. 2020 is reported separately below.
- Returns. One unit risked per selection, each game weighted equally regardless of how many books priced it.
- Buckets were fixed in advance. No edge was moved after seeing a result, and no bucket was merged or split.
Calibration: the market is honest
Each point below is one probability band: where the market said it should land, against where it actually landed. Perfect calibration is the dashed diagonal.
| Bucket | n | Avg price | Expected | Actual | Gap | 95% CI on gap | ROI | Seasons + |
|---|---|---|---|---|---|---|---|---|
| <35% | 1,753 | +213 | 31.16% | 29.49% | -1.67 pp | -3.9% to +0.5%spans 0 | -8.59% | 1/5 |
| 35-40% | 2,454 | +156 | 37.77% | 39.20% | +1.43 pp | -0.5% to +3.3%spans 0 | +0.22% | 3/5 |
| 40-45% | 3,647 | +127 | 42.60% | 43.13% | +0.53 pp | -1.1% to +2.1%spans 0 | -2.15% | 3/5 |
| 45-50% | 4,142 | +103 | 47.52% | 47.47% | -0.05 pp | -1.6% to +1.4%spans 0 | -3.84% | 2/5 |
| 50-55% | 4,142 | -120 | 52.48% | 52.53% | +0.05 pp | -1.4% to +1.6%spans 0 | -3.62% | 3/5 |
| 55-60% | 3,647 | -147 | 57.40% | 56.87% | -0.53 pp | -2.1% to +1.1%spans 0 | -4.32% | 2/5 |
| 60-65% | 2,454 | -181 | 62.23% | 60.80% | -1.43 pp | -3.3% to +0.5%spans 0 | -5.71% | 2/5 |
| 65-70% | 1,209 | -231 | 67.25% | 69.31% | +2.06 pp | -0.5% to +4.6%spans 0 | -0.67% | 4/5 |
| 70%+ | 544 | -303 | 72.37% | 73.16% | +0.79 pp | -3.1% to +4.5%spans 0 | -2.61% | 3/5 |
All 9 intervals contain zero. The largest apparent miss, 65-70% at +2.06 pp, is consistent with noise at this sample size. And no band’s residual holds its sign across every season: if a band were genuinely mispriced, that would show up year after year. It does not.
“Seasons +” counts seasons with a positive residual. A gap whose interval spans zero is not distinguishable from perfect calibration.
Returns by moneyline bucket
| Bucket | n | Avg price | Expected | Actual | Gap | 95% CI on gap | ROI | Seasons + |
|---|---|---|---|---|---|---|---|---|
| +201 or longer | 928 | +235 | 29.04% | 27.16% | -1.89 pp | -4.6% to +1.0%spans 0 | -9.40% | 1/5 |
| +151 to +200 | 2,318 | +172 | 35.66% | 36.07% | +0.40 pp | -1.6% to +2.3%spans 0 | -2.56% | 1/5 |
| +126 to +150 | 2,828 | +138 | 40.64% | 42.26% | +1.61 pp | -0.2% to +3.5%spans 0 | +0.42% | 2/5 |
| +101 to +125 | 4,098 | +113 | 45.32% | 44.31% | -1.00 pp | -2.5% to +0.5%spans 0 | -5.69% | 0/5 |
| even to -120 | 4,198 | -110 | 50.32% | 50.52% | +0.21 pp | -0.3% to +0.7%spans 0 | -3.29% | 0/5 |
| -121 to -140 | 2,626 | -130 | 54.46% | 56.02% | +1.56 pp | -0.4% to +3.5%spans 0 | -0.80% | 2/5 |
| -141 to -160 | 2,297 | -148 | 57.63% | 56.73% | -0.90 pp | -2.9% to +1.2%spans 0 | -5.01% | 1/5 |
| -161 to -200 | 2,782 | -177 | 61.62% | 59.92% | -1.70 pp | -3.5% to +0.2%spans 0 | -6.18% | 0/5 |
| -201 to -250 | 1,211 | -224 | 66.60% | 68.13% | +1.53 pp | -1.1% to +4.2%spans 0 | -1.44% | 4/5 |
| -251 or shorter | 706 | -292 | 71.71% | 72.24% | +0.53 pp | -2.9% to +3.7%spans 0 | -2.98% | 2/5 |
One bucket shows a positive pooled return: +126 to +150, at +0.42%. Its interval runs from -3.9% to +4.8%, and it was profitable in only 2 of 5 seasons. That is not an edge. That is what noise looks like when you slice 23,992 bets 10 ways.
The worst place to be was +201 or longer, at -9.40%, the one result with real economic size, and also one of the widest intervals, on only 928 bets.
Season stability: the discipline that kills most findings
This is the most important table here, and the one most public analysis omits.
| Bucket | 2021 | 2022 | 2023 | 2024 | 2025 |
|---|---|---|---|---|---|
| +201 or longer | -0.7% | +5.6% | -19.7% | -11.0% | -28.9% |
| +151 to +200 | -4.1% | -11.4% | +9.4% | -3.2% | -1.6% |
| +126 to +150 | -3.3% | -2.9% | +5.4% | -5.1% | +8.7% |
| +101 to +125 | -5.9% | -9.8% | -10.7% | -2.0% | -1.2% |
| even to -120 | -2.5% | -4.3% | -2.6% | -3.5% | -3.7% |
| -121 to -140 | -4.4% | +2.9% | +5.1% | -0.9% | -6.4% |
| -141 to -160 | +4.2% | -0.9% | -7.3% | -12.1% | -8.4% |
| -161 to -200 | -7.3% | -2.6% | -8.9% | -2.0% | -10.8% |
| -201 to -250 | +0.8% | +0.7% | -11.1% | +1.3% | +1.9% |
| -251 or shorter | -8.5% | -6.2% | +3.9% | -3.4% | +1.0% |
Look at -201 to -250. It was profitable in 4 of 5 seasons. A tout would stop reading there. But 2023 returned -11.1%, and the pooled result is -1.44%. One season erased the rest.
Only 3 buckets hold a consistent sign across every season, and all 3 are consistently negative: +101 to +125, even to -120, -161 to -200. Not one bucket is reliably profitable; every apparently positive one flips sign at least once.
Return per unit risked, by bucket and season. Shading tracks magnitude; the value is always printed, so nothing is conveyed by colour alone.
Why a calibrated market still loses you money
This is the part worth internalising. Across the whole sample the average no-vig expectation is 50.00% and the average realized win rate is 50.00%. Those two agree by construction rather than by measurement: this sample holds both sides of every game, and a home residual is exactly minus the away residual, so the pooled figures cancel to an arithmetic identity whose interval has zero width. The calibration evidence is the band table above, where nothing cancels.
But you cannot bet at that number. You bet at prices implying 51.87%. The difference, +1.87 pp per side, a 3.7% overround on the two-way market, is the fee. Our pooled return of -3.55% is close to what you would expect from paying it on every bet. The house edge, not bad forecasting, takes the money.
The practical implication is uncomfortable and important: beating a market like this requires being more accurate than it by more than the vig. Being merely as accurate as the market is a slow, arithmetically certain loss. Any strategy that does not clear that bar is a fee-paying exercise, however sophisticated it looks.
Appendix: 2020
Reported separately and excluded from the headline sample: a shortened season on a thinner book consensus is a different market.
Limitations
- Closing prices only. We measure the market at its most efficient moment, and a bettor cannot know the closing line in advance. Nothing here says whether earlier prices are beatable.
- Mean-across-books execution. We assume the average eligible closing price. A disciplined line-shopper would do better; someone stuck at one book, worse.
- Uncertainty is clustered on the game. Where one game puts both of its sides into a cohort, those selections are not independent draws, so the resampling unit is the game. Mirrored sides are negatively dependent, so treating them as independent would have reported a wider interval than the data support for the near-pick’em bucket, not a narrower one. The probability bands contain no mirrored pairs at all: they split at 50%, so a price and its complement can never share a band.
- Scope. No live or in-game markets, no derivatives, no player props. Results come from Retrosheet, which issues periodic corrections; figures are reproducible against the pinned build recorded below.
- We did not search for profitable subsets. A determined search over odds ranges, teams, months and parks would certainly surface something profitable-looking. It would almost certainly be noise, and finding it is not research.
- This is description, not prediction. No forecasting model is fitted or evaluated anywhere here, and no team, player or situational variable is used. Nothing should be read as a claim about why the market is calibrated, only that it was.
Ask your own question of this data
The Historical Explorer runs the same engine behind these figures over six seasons of closing prices. Pick a price range, a side and a season, and see what actually happened.
Open the Historical ExplorerGet new sports-market research when we publish
Occasional email, only when a new research note goes out. No picks, no promotions. Everything on this site stays free to read without an account.
First Win Pro
Go beyond the published finding
The public research answers what did we find. Pro answers what else the research showed, where the result breaks down, and how to interrogate the work yourself. It is the deeper layer behind the public studies, for people who want to use the research rather than only read the conclusion.
The public work stays complete. Nothing is held back from an article to manufacture a reason to subscribe.
What members get today
- Research Packs
- The full workings behind each major study: every supporting table, the robustness and sensitivity work, and the cuts too detailed for a public article.
- Pro-only research notes
- Focused quantitative investigations into questions that do not warrant a full publication but are still worth answering properly.
- Extended analysis and robustness work
- Where a result holds, where it weakens, and what happens to it under a different estimator or a different cut of the same data.
- Downloadable research assets
- Derived tables from the sealed research packages, where the underlying licence permits redistribution.
- Early access
- Completed studies reach members before they are published publicly.
- Research queue input
- A say in which questions get investigated next. Research topics, not requests for selections.
What we are building
The current membership is a research membership. The full First Win Pro research terminal is still being built, and founding membership is a claim on it rather than access to something that exists today.
- A professional research terminal, with richer historical analysis than the public Explorer
- Comparison workflows across cohorts, seasons and markets
- Saved research, saved screens and saved queries
- Advanced statistical views over the same canonical research layer
- Additional markets and sports as the corpus expands
- Alerts and watchlists, where they can be built without becoming a tipping service
Founding membership opens shortly and is not yet on sale, so there is nothing to buy today. The newsletter above will carry the announcement. No picks, no tips, and no claim that any of this will make you money betting.
Methodology, version and provenance
Intervals are 95% percentile intervals from a deterministic event-clustered bootstrap: 2,000 iterations, fixed seed, the game as resampling unit, so multiple selections from one game keep their observed dependence. Events are ordered canonically before the random stream is consumed, so an interval does not depend on input row order. The parent hash above is the original v1 freeze, retained as provenance; v1.0.1 changes only the uncertainty method, never a point estimate. These figures are a sealed artifact produced once against the fact build above. They are served read-only and are not recomputed on request, so a published number cannot drift when the corpus is rebuilt.