Prediction Markets · learn
Are Prediction Markets Accurate? Calibration Data
By Odds Reference Published March 4, 2026 Updated July 19, 2026 Editorial Policy
Our dataset of resolved markets across Polymarket, Kalshi, and Metaculus shows that prediction markets are well-calibrated on liquid events. Contracts priced at 70% resolve positively roughly 70% of the time. This calibration holds across categories, platforms, and time periods — but only when sufficient trading volume exists.
What Does “Calibration” Mean for Prediction Markets?
Calibration measures whether predicted probabilities match actual outcome frequencies over many predictions. A prediction source is perfectly calibrated if events it assigns a 60% probability occur exactly 60% of the time, events at 80% occur 80% of the time, and so on.
This is distinct from resolution accuracy (did any single prediction come true?) and from sharpness (how close to 0% or 100% are the predictions?). A forecaster who predicts 50% for every event would be perfectly calibrated on a large sample but utterly uninformative. Good prediction markets are both well-calibrated and sharp — they assign extreme probabilities when the evidence supports it, and those extreme probabilities prove correct.
The standard quantitative measure is the Brier score, which ranges from 0 (perfect) to 1 (maximally wrong). A Brier score of 0.25 corresponds to always predicting 50% — the baseline for a binary event with no information. Prediction markets on major events consistently score well below this baseline. For the full formula and how it decomposes into calibration and resolution components, see our calibration and forecasting guide.
What Does the Calibration Data Show?
Our resolved-market dataset tracks close to the calibration diagonal across every probability bin, with the clearest deviation a modest overconfidence in the 70-89% range — a pattern consistent with the favorite-longshot bias. Bins near 0% and 100% carry thinner samples, so we report qualitative sample-depth labels rather than invented precision.
| Probability Bin | Bin Midpoint | Observed Pattern | Sample Depth |
|---|---|---|---|
| 90-100% | ~95% | Tracks closely to the midpoint | Strong |
| 70-89% | ~80% | Modest overconfidence — observed rates trend a few points below the midpoint | Strong |
| 50-69% | ~60% | Tracks closely; highest inherent uncertainty at this range | Moderate |
| 30-49% | ~40% | Tracks closely | Growing |
| 10-29% | ~20% | Tracks closely, wider variance at current sample depth | Growing |
| 0-9% | ~5% | Tracks closely, wider variance at current sample depth | Limited |
This is a methodology choice, not a data gap: we are not publishing point-estimate observed win rates or exact sample counts in this table, because doing so would imply more statistical confidence than our current resolved-market sample supports, particularly in the “Growing” and “Limited” rows — markets don’t sit at extreme prices for long before resolving, so those bins fill slowly. We will publish exact counts and precise percentages once each bin holds enough resolved markets to make the number meaningful rather than noisy (methodology and reasoning last verified July 2026). For the full bin-by-bin breakdown with Brier scores and category splits, see our 2026 accuracy report, which explains this same qualitative-label approach in its own FAQ and uses it for the same reason.
The Odds Reference dashboard tracks active market prices across platforms in real time, and we update calibration analysis as resolution data accumulates.
How Do Prediction Markets Compare to Polls and Expert Forecasts?
The academic record on this question spans decades. The Iowa Electronic Markets, which operated continuously from 1988 through multiple election cycles, provided the foundational dataset. Researchers found that market prices outperformed major polls in predicting election outcomes 74% of the time in head-to-head comparisons.
Several structural reasons explain this advantage:
Information aggregation speed. Markets incorporate new information within minutes. A debate performance, an economic data release, or a policy announcement moves contract prices almost immediately. Polls take days to field and report, creating a structural lag.
Incentive alignment. Market participants risk real money on their beliefs, creating a direct incentive to be accurate rather than to signal social desirability. Polls capture what respondents say they believe, which research shows can diverge from their actual expectations, particularly on socially charged topics.
Diverse information sources. A single market price reflects the combined knowledge of political analysts, quantitative modelers, local observers, and general-interest traders. No individual poll methodology captures this breadth.
Continuous updating. Markets produce a real-time probability estimate that adjusts constantly. Polls produce periodic snapshots that may already be outdated by publication.
The advantage is not absolute. Polls provide demographic and geographic breakdowns that markets cannot replicate. And on events where market participation is thin, polls with large sample sizes may outperform. The strongest forecasting approach, as demonstrated by research at the Good Judgment Project, combines market signals with structured polling and expert assessment.
Does Liquidity Affect Accuracy?
Yes, and it is one of the strongest patterns in our dataset. Markets with sustained daily volume above roughly $50,000 track the calibration diagonal closely; markets trading under $500 a day produce noisy, unreliable prices driven by a handful of participants rather than genuine information aggregation (thresholds last verified July 2026).
| Liquidity Level | Daily Volume | Calibration Quality | Spread |
|---|---|---|---|
| High | >$50,000 | Strong (close to diagonal) | 1-2 cents |
| Medium | $5,000-$50,000 | Good (slight deviations) | 2-5 cents |
| Low | $500-$5,000 | Moderate (visible bias) | 5-15 cents |
| Minimal | <$500 | Unreliable | 10-30+ cents |
Markets with daily volume above $50,000 — typically major political events, high-profile economic indicators, and viral cultural moments on Polymarket — demonstrate calibration that matches or exceeds the best academic forecasting benchmarks.
Below $5,000 in daily volume, calibration degrades noticeably. The prices in these markets reflect a small number of opinions rather than genuine information aggregation. A contract at $0.65 in a low-liquidity market might represent two traders rather than a robust probability estimate. Our market liquidity guide breaks down the specific order-book and spread metrics behind these buckets.
This is why cross-platform comparison matters. When the same event trades on both Kalshi and Polymarket, price convergence between the two platforms signals stronger reliability than either price alone. Price divergence signals either different information sets or insufficient liquidity on one or both platforms.
What Are the Known Limitations of Prediction Market Accuracy?
Several well-documented failure modes limit prediction-market accuracy: manipulation on thin markets, information cascades where traders herd on a single source, reduced participation from regulatory restrictions, wider deviations on long-duration contracts, and systematic underpricing of rare tail events. None of these make markets useless — they define where to apply extra skepticism.
Manipulation on thin markets. A single large trade can move a thin market by 10-20 percentage points. While research suggests manipulation effects are temporary on liquid markets (other traders arbitrage the mispricing away), thin markets may remain distorted for extended periods.
Correlated information cascades. When most participants rely on the same information sources — a single poll, a viral social media post, a dominant media narrative — the market price reflects that shared source rather than genuinely diverse information. This is most common on politically polarized events.
Regulatory uncertainty affecting participation. US restrictions on Polymarket trading reduce the participant pool, potentially excluding knowledgeable traders and degrading accuracy. Kalshi operates as a CFTC-regulated exchange, which increases trust but limits the types of events it can list — and its state-by-state availability is itself contested in court, tracked on our legal tracker.
Long-duration markets. Contracts that resolve months or years in the future tend to show wider calibration deviations than short-duration markets. The discount rate, opportunity cost of capital, and information uncertainty all increase with time horizon.
Tail events. Markets systematically underestimate the probability of extreme outcomes. Events priced at 2-5% occur more frequently than the price implies, a pattern consistent across financial markets and prediction markets alike. Our data in the 0-9% bin, while still limited, shows early signs of this longshot bias.
Prediction market trading carries real financial risk regardless of how well-calibrated the prices are. If you or someone you know needs support around gambling or trading behavior, see our responsible gambling resources.
How Should You Interpret Prediction Market Prices?
A contract price is a probability estimate, not a forecast of what will happen — confusing the two is the most common misreading of prediction market data. Three ideas separate correct interpretation from folk intuition: what a single price means, why prices near 50% carry the most uncertainty, and why the trend matters more than the level.
A $0.70 contract is not a prediction that the event will happen. It means the market estimates a 70% chance. Three out of ten times, the event should fail to occur. If you observe a $0.70 contract resolve to $0 and conclude the market was “wrong,” you are misunderstanding probability.
Prices near 50% carry the most uncertainty. A $0.50 contract is the market’s way of saying it has no strong lean. These markets are the hardest to trade profitably and the most likely to surprise.
Price movement matters as much as price level. A contract moving from $0.40 to $0.65 in a week signals meaningful new information. A contract sitting at $0.65 for months signals stability in the market’s assessment. The trajectory tells a story the snapshot cannot.
Once you have a view on where the true probability sits relative to the market price, turning that gap into a position size is a separate question — our EV calculator handles the math of converting price and probability into expected value. For deeper context on reading these signals, see our guide on how to read prediction market data and the fundamentals of prediction markets.
How Does Calibration Vary by Category?
Category matters as much as liquidity. Politics calibrates tightest thanks to abundant polling and binary outcomes; economics performs well on markets tied to clear numerical thresholds; science and technology are mixed depending on specialist participation; and sports trades thinner than dedicated sportsbooks, weakening the signal.
Politics and elections represent the strongest calibration category. High public interest drives deep liquidity, abundant polling data provides external anchoring, and binary outcomes (win/lose) simplify the resolution criteria. Major US elections on Polymarket have demonstrated near-perfect calibration in the final 48 hours before resolution.
Economics and monetary policy show strong calibration on events with clear numerical thresholds (Fed rate decisions, jobs reports above/below consensus). Markets that require longer time horizons — annual GDP growth, recession probability — show wider deviations.
Science and technology exhibit more variance. Markets on FDA approvals and AI benchmarks attract informed specialist traders and calibrate well. Markets on speculative technology timelines (fusion energy, AGI) tend to be thinner and less reliable.
Sports on prediction markets are less liquid than dedicated sportsbooks, and the calibration data reflects this. Our platform comparison breaks down which platforms offer meaningful sports coverage and where sportsbook odds provide a stronger signal, and our analysis of prediction markets vs. sports betting covers the structural differences in more detail.
Our methodology page explains how we collect, match, and validate the resolved-market data behind every table on this page.
Key Takeaways
- Prediction markets are well-calibrated on liquid events: 70% contracts resolve positively approximately 70% of the time, confirmed across our multi-platform dataset
- Accuracy depends heavily on liquidity — markets with over $50,000 daily volume match or exceed academic forecasting benchmarks, while sub-$500 markets produce unreliable signals
- Academic research on markets like the Iowa Electronic Markets shows prediction markets outperform polls in head-to-head election forecasting, with the advantage growing as the event approaches
- Known failure modes include manipulation on thin markets, correlated information cascades, and systematic underpricing of tail events
- Cross-platform price convergence is a stronger reliability signal than any single platform’s price — the dashboard tracks these spreads automatically