Prediction Markets · learn
Prediction Market Accuracy Report 2026: Calibration Data
By Odds Reference Published March 4, 2026 Updated July 18, 2026 Editorial Policy
Prediction markets are well-calibrated on liquid events: prices track close to the calibration diagonal, with only modest, qualitatively-flagged deviations concentrated in thinly-traded bins. This report breaks that claim down by probability bin, category, and platform across Polymarket, Kalshi, and Metaculus — using honest, qualified language wherever our resolved-market sample is still thin.
How Do We Measure Prediction Market Accuracy?
Calibration analysis compares predicted probabilities against actual outcome frequencies across resolved markets. A well-calibrated source assigns probabilities that match reality: events priced at 70% should resolve positively about 70% of the time, and events at 30% should resolve positively about 30% of the time.
We use two primary metrics:
Calibration error measures the gap between predicted probability and observed frequency within each probability bin. A market that prices events at 80% but sees them occur meaningfully less often than that is overconfident in that bin. Perfect calibration means zero gap across all bins.
Brier score captures both calibration and sharpness in a single number. It computes the mean squared difference between the predicted probability and the binary outcome (0 or 1). The scale runs from 0 (perfect) to 1 (maximally wrong). A naive forecaster who always predicts 50% scores 0.25 on balanced binary events — this is the baseline to beat.
| Metric | Formula | Range | Interpretation |
|---|---|---|---|
| Calibration error | |predicted % - observed %| per bin | 0-100% | Lower is better; 0 = perfect alignment |
| Brier score | Mean of (predicted - outcome)^2 | 0-1 | Lower is better; <0.25 = better than coin flip |
| Sharpness | Distribution of predicted probabilities | 0-100% | Higher concentration near 0% and 100% = more informative |
A forecaster can be well-calibrated but uninformative (predicting 50% for everything) or sharp but poorly calibrated (predicting 90% on events that happen 60% of the time). The best prediction markets are both calibrated and sharp.
For a deeper explanation of calibration concepts, see our calibration and forecasting guide.
What Does the Calibration Data Show?
Across probability bins with adequate volume, resolved prices track close to the calibration diagonal — the implied probability roughly matches how often the event actually happened. Bins at the extremes carry thinner samples, since fewer markets trade at extreme prices for extended periods before resolving.
| Probability Bin | Bin Midpoint | Observed Pattern | Sample Depth |
|---|---|---|---|
| 90-100% | ~95% | Tracks closely to the bin midpoint | Strong |
| 80-89% | ~85% | Modest overconfidence — observed rates trend a few points below the midpoint | Strong |
| 70-79% | ~75% | Tracks closely | Strong |
| 60-69% | ~65% | Tracks closely | Moderate |
| 50-59% | ~55% | Tracks closely; highest inherent uncertainty at this bin | Moderate |
| 40-49% | ~45% | Tracks closely | Growing |
| 30-39% | ~35% | Tracks closely | Growing |
| 20-29% | ~25% | Tracks closely | Growing |
| 10-19% | ~15% | Tracks closely, wider variance at current sample depth | Limited |
| 0-9% | ~5% | Tracks closely, wider variance at current sample depth | Limited |
We are deliberately not publishing point-estimate observed win rates or precise calibration-error percentages in this table. Doing so would imply a level of statistical confidence our current resolved-market sample doesn’t yet support, particularly in the “Growing” and “Limited” rows. The pattern that does hold up: markets track the diagonal closely, with the clearest deviation being modest overconfidence in the 80-89% bin — a pattern consistent with the favorite-longshot bias documented in academic forecasting literature.
Bins below 30% and above 90% carry thinner samples because fewer markets trade at extreme probabilities for extended periods. As our resolution dataset grows, we plan to publish tighter, number-backed estimates in place of the qualitative labels above.
The Odds Reference dashboard tracks active market prices in real time. Calibration analysis updates as markets resolve.
How Does Accuracy Vary by Category?
Accuracy is not uniform across topics. Political and economic markets calibrate tightly thanks to deep liquidity and unambiguous resolution criteria; crypto markets show wider errors driven by speculative momentum; and long-range science and geopolitics questions — Metaculus’s specialty — accumulate calibration data more slowly but show promising early patterns.
Politics and elections. Political markets attract the deepest liquidity on both Polymarket and Kalshi. High-profile elections draw thousands of traders and produce strong calibration results. Political markets tend to be well-calibrated outside of the final 24 hours before resolution, when last-minute information can cause sharp price swings that do not always align with realized probabilities.
Economics and policy. Markets on Federal Reserve decisions, inflation prints, and GDP releases perform well when tied to specific, verifiable data releases. Resolution criteria are typically unambiguous (did the Fed cut or not?), which reduces post-resolution disputes and keeps calibration clean. Kalshi carries particularly deep liquidity in this category.
Crypto and technology. Crypto markets show higher volatility and wider calibration errors. This reflects both genuine uncertainty and the influence of speculative momentum on prices. Markets on token prices or protocol events attract crypto-native traders whose risk preferences may differ from the broader forecasting population.
Science and geopolitics. Metaculus dominates these categories. Long-range questions about AI milestones, pandemic risks, and geopolitical events benefit from Metaculus’s deep forecaster community, many of whom have domain expertise. However, these markets often have multi-year horizons, meaning calibration data accumulates slowly.
Sports. Sports prediction markets benefit from abundant statistical data and frequent resolution. Game-level markets tend to be well-calibrated, with most of the signal concentrated in closing prices. Our analysis of prediction markets vs. sports betting covers the structural differences in more detail.
For the broader picture on how liquidity drives calibration across every category, see our guide on whether prediction markets are accurate.
How Does Accuracy Vary by Platform?
Each platform’s calibration profile reflects its liquidity, regulatory status, and incentive structure. Polymarket and Kalshi both show strong calibration on high-volume markets, with results converging as volume increases, while Metaculus — despite using no real money — produces calibration that holds up well against real-money markets on long-range questions.
| Platform | Strengths | Calibration Profile | Key Factor |
|---|---|---|---|
| Polymarket | Politics, crypto, major global events | Strong on high-liquidity markets; weaker on thin markets | CLOB depth drives accuracy |
| Kalshi | Economics, policy, regulated US events | Consistent across category; clean resolution criteria | Regulatory clarity reduces ambiguity |
| Metaculus | Science, AI, long-range geopolitics | Strong track record in academic evaluations | Reputation incentives, expert community |
Polymarket benefits from being one of the largest prediction markets by volume. Deeper order books mean more information gets incorporated into prices. Our data shows Polymarket calibration improves as total market volume increases; thinly-traded markets produce noisier, less reliable prices than deep ones.
Kalshi operates as a CFTC-regulated exchange with standardized contract specifications. The regulatory framework forces clear resolution criteria, which reduces one source of calibration error: ambiguous outcomes. Kalshi’s accuracy data is most robust on economic indicators where the resolution source (BLS data, Fed announcements) is unambiguous.
Metaculus presents a distinct case. No real money changes hands — forecasters earn reputation points. Despite this, Metaculus has produced calibration results that compete with real-money markets, particularly on scientific and long-range questions. The platform’s forecaster community includes researchers and domain experts who bring specialized knowledge that money alone does not attract.
Cross-platform consensus — where Polymarket, Kalshi, and Metaculus agree on a probability — tends to produce the strongest calibration signal. When platforms diverge, the disagreement itself carries information. Our platform comparison tracks these divergences in real time. Once you have a calibrated probability, turning it into a position size is a separate question — our EV calculator handles the math of converting price and probability into expected value.
What Are the Limitations of This Report?
This report’s biggest constraint is sample maturity: many tracked markets, especially long-range ones, have not yet resolved, and our qualitative sample-depth labels exist precisely because we don’t yet have grounds to publish precise point estimates in every bin. Read the specific caveats below before weighting any single data point heavily.
Sample size is growing, not final. Many markets listed on our tracked platforms have not yet resolved. Long-range markets on 2027 or 2028 events will not contribute calibration data for years. Current results are weighted toward short-duration markets that resolve within weeks or months.
Resolution timing matters. We snapshot prices at standardized intervals before resolution, but the choice of snapshot time affects calibration results. A market that sits at 80% for three months but drops to 50% in the final hour looks very different depending on which price you use. We use a pre-resolution snapshot that balances stability against recency.
Platform differences complicate comparison. Polymarket and Kalshi trade real money in different regulatory environments with different user bases. Metaculus uses no money at all. Comparing calibration across these platforms is informative but not apples-to-apples. Incentive structures, market microstructure, and trader demographics all influence results.
Thin markets distort statistics. Markets with few active traders can produce extreme prices that do not reflect genuine probability estimates. We apply minimum liquidity filters, but the threshold is a judgment call. Our methodology page details these filters and how we collect and validate the underlying data.
We are still building resolution data. This report updates as our dataset grows. Current numbers should be read as early indicators, not definitive measurements. The academic literature on prediction market calibration — drawing on decades of data from sources like the Iowa Electronic Markets — provides additional context that our dataset alone cannot yet match.
Prediction market trading carries real financial risk regardless of how well-calibrated the prices are. If you or someone you know needs support around gambling or trading behavior, see our responsible gambling resources.
Key Takeaways
- Prediction markets are well-calibrated on liquid events. Across bins with adequate sample size, resolved prices track close to the calibration diagonal — we report qualitative sample-depth labels rather than invented precision where the underlying sample is still thin.
- Accuracy correlates with liquidity, not platform identity. High-volume markets on Polymarket, Kalshi, or Metaculus all show strong calibration. Thin markets on any platform show wider, less reliable pricing.
- Category matters. Political and economic markets calibrate tightly. Crypto markets show more noise. Long-range scientific questions accumulate calibration data slowly but show promising early results on Metaculus.
- Cross-platform consensus is the strongest signal. When multiple platforms agree on a probability, the combined forecast tends to outperform any single source.
- This is a living report. Our dataset grows with every resolved market, and we tighten qualitative labels into precise figures as sample depth allows. Updated calibration data is available on the Odds Reference dashboard.