Prediction Markets · learn
What Is Calibration in Forecasting? A Data-Driven Explainer
By Odds Reference Published March 4, 2026 Updated July 19, 2026 Editorial Policy
A forecaster is calibrated when their predicted probabilities match observed frequencies. If a forecaster says “70% likely” across one hundred different predictions, roughly seventy of those events should actually occur. Calibration is the most rigorous standard for evaluating whether probability estimates are meaningful, and it separates useful forecasts from confident guesses.
How Is Calibration Measured?
Calibration is assessed by grouping predictions into probability buckets and comparing the predicted frequency against the actual outcome frequency. A forecaster who assigns 90% probability to fifty events is well-calibrated if approximately forty-five of those events occur. Analysts plot this relationship on a reliability diagram and measure the gap between the diagonal and observed curve to quantify over- or under-confidence.
The standard visualization is a calibration curve (also called a reliability diagram). The x-axis shows predicted probability, the y-axis shows observed frequency of outcomes, and a perfectly calibrated forecaster falls along the diagonal:
| Predicted Probability | Ideal Outcome Rate | Overconfident Example | Underconfident Example |
|---|---|---|---|
| 10% | 10% | 5% (too few happen) | 18% (too many happen) |
| 30% | 30% | 15% | 42% |
| 50% | 50% | 35% | 60% |
| 70% | 70% | 55% | 78% |
| 90% | 90% | 75% | 95% |
Overconfidence means extreme probabilities are assigned too frequently — the forecaster says 90% but the event only happens 75% of the time. Underconfidence means probabilities cluster toward the center — the forecaster says 60% when the true rate is 80%. Most human forecasters exhibit overconfidence, assigning extreme probabilities more often than their accuracy justifies.
Calibration requires volume. A single prediction cannot be evaluated for calibration. You need hundreds or thousands of forecasts across different probability levels to produce a meaningful calibration curve. This is one reason why platform-level data is more informative than any individual forecaster’s track record.
What Is the Brier Score and Why Does It Matter?
The Brier score is the standard metric for scoring probabilistic forecasts: the mean squared error between each predicted probability and its actual binary outcome. It matters because it rewards both accuracy and confidence — a forecaster who hedges every call to 50% cannot post a strong score, and neither can one who is boldly wrong.
Brier Score = (1/N) x SUM(forecast - outcome)^2
Where forecast is the predicted probability (0 to 1) and outcome is 1 if the event occurred and 0 if it did not. The sum runs across all N predictions.
Key properties of the Brier score:
- Range: 0 to 1. Lower is better.
- Perfect score: 0 (every prediction was 100% for events that happened and 0% for events that did not).
- Worst score: 1 (every prediction was completely wrong).
- Random baseline: 0.25 (assigning 50% to everything). Any Brier score below 0.25 beats coin-flipping.
- Climatological baseline: Using the historical base rate for every prediction. Beating this baseline means the forecaster adds value beyond simple frequency knowledge.
The Brier score can be decomposed into three components: calibration (do probabilities match frequencies?), resolution (do predictions vary meaningfully, or does the forecaster just say 50% for everything?), and uncertainty (how inherently unpredictable are the events?). A forecaster with good calibration and good resolution — meaning they assign diverse, accurate probabilities — will have a low Brier score.
For a broader look at how forecasting accuracy plays out across platforms, see our prediction market accuracy analysis.
How Calibrated Are Prediction Markets?
Prediction markets aggregate the beliefs of many participants into a single price, and the resulting calibration is generally strong — particularly on liquid markets. Decades of research on markets like the Iowa Electronic Markets, plus recent tracking of Polymarket and Kalshi, show contracts priced at 70% resolving positively close to 70% of the time once enough volume is trading.
The Iowa Electronic Markets, run continuously by the University of Iowa since 1988, produced the foundational dataset here: market prices closely tracked actual election outcomes, outperforming major polls in the majority of head-to-head comparisons across multiple cycles. More recent tracking of Polymarket and Kalshi confirms the same pattern — high-volume markets produce well-calibrated prices, thin ones do not.
Liquidity is the strongest predictor of calibration quality in our data. A contract trading at $0.70 on a market with deep, sustained volume reflects genuine information aggregation across many independent participants; the same price on a market with only a handful of trades may be driven by one or two people and carries far less predictive weight. Our 2026 accuracy report breaks this down by probability bin and sample depth rather than a single point estimate — current sample sizes in the thinnest bins don’t yet support a precise dollar-volume cutoff (last verified July 2026).
The Odds Reference dashboard tracks prices across multiple platforms. When the same event trades on Polymarket, Kalshi, and community platforms like Metaculus, cross-platform price convergence is itself a calibration signal — agreement across independent participant pools strengthens confidence in the implied probability.
Community forecasting platforms offer an interesting calibration comparison. Metaculus, which uses no real money, has demonstrated calibration comparable to financial prediction markets on overlapping question sets. This challenges the assumption that financial incentives are strictly necessary for good calibration, though the topic remains actively debated in the research literature.
What Are Superforecasters and Why Do They Matter?
The term “superforecaster” comes from Philip Tetlock’s research, particularly the Good Judgment Project, a multi-year study funded by IARPA (Intelligence Advanced Research Projects Activity). The project identified individuals who consistently outperformed both chance and professional intelligence analysts on geopolitical forecasting questions.
Superforecasters share several measurable traits:
- Granularity. They assign precise probabilities (67% rather than “likely”) and adjust them frequently as new information arrives.
- Intellectual humility. They treat their own estimates as hypotheses to be updated, not positions to be defended.
- Base-rate awareness. They start with the historical frequency of similar events and adjust from there, rather than reasoning from narratives alone.
- Active updating. They monitor evidence continuously and revise forecasts in small increments.
Tetlock’s research demonstrated that structured probabilistic thinking — the kind embodied by calibration-focused forecasting — produces measurably better results than traditional expert judgment, which tends toward overconfidence and narrative-driven reasoning.
Prediction markets incorporate superforecaster-like behavior structurally: participants with better information or better models make more money, and their trades move prices proportionally to their capital and conviction. The market price, in theory, reflects the beliefs of the best-informed participants weighted by their willingness to back those beliefs with money.
How Can You Improve Your Own Calibration?
Calibration is a skill that improves with deliberate practice, not innate talent. Superforecaster research identifies four concrete habits that build it: tracking every prediction against its outcome, anchoring to reference-class base rates, actively hunting for disconfirming evidence, and drilling confidence-mapping on structured trainers. Each habit is measurable and improvable over time.
Track your predictions. Record every probability estimate you make and the eventual outcome. Over time, group your predictions into buckets and check whether your 70% predictions happen about 70% of the time. Without tracking, self-assessment is unreliable.
Use reference classes. Before estimating the probability of a specific event, ask: “How often do events like this happen historically?” Starting from the base rate and adjusting based on specific evidence consistently produces better-calibrated estimates than starting from intuition.
Seek disconfirming evidence. Actively look for reasons your estimate might be wrong. Overconfidence typically stems from anchoring on confirming evidence and underweighting contradictory signals.
Practice on calibration trainers. Several online tools present trivia questions and ask you to assign confidence levels, then score your calibration in real time. These build the habit of mapping internal confidence to accurate probability estimates.
Once you have a genuinely calibrated probability for an event, the next step is comparing it to the market’s price. If your estimate is 60% and a contract trades at $0.50, that 10-point gap is your edge — our EV calculator turns that gap into an expected-value figure so you can judge whether the difference is large enough to act on.
The prediction market glossary defines key terms used throughout calibration and forecasting analysis.
Prediction market trading carries real financial risk regardless of how well-calibrated the prices are. If you or someone you know needs support around gambling or trading behavior, see our responsible gambling resources.
Key Takeaways
- Calibration means your probability estimates match reality: 70% predictions should come true approximately 70% of the time
- The Brier score (mean squared error of forecasts) is the standard metric — scores below 0.25 beat random guessing, and decomposition reveals calibration versus resolution
- High-liquidity prediction markets demonstrate strong calibration; thin markets are noticeably less reliable
- Superforecasters, identified through Tetlock’s research, outperform experts by using precise probabilities, base rates, and frequent updating
- Track your own predictions systematically to build calibration as a measurable, improvable skill