Methodology · Ratings

OR Score Methodology

The OR Score is our framework for rating prediction-market exchanges from reproducible data we measure ourselves: market breadth, category coverage, cost, liquidity, and settlement reliability. In v1 we publish the metrics we can measure as a transparent, side-by-side comparison — not a single ranked number, and not an opinion. Affiliate relationships have no weight in any of it.

Framework v1.0.0 · Metrics or-score-metrics-v1 · Metrics as of · Last Updated:

Why There Is No Single Score Yet

v1 deliberately does not publish a single ranked OR Score. A composite would mislead today for two concrete reasons: category coverage is a flat 6/6 across every venue, so it cannot tell them apart; and we do not observe Polymarket’s bid-ask spread (we ingest its CLOB midpoint, not its order book), so its all-in cost is unknown — a blended score would structurally flatter the one venue whose true cost we cannot see. So we publish the measured metrics below, honestly, and defer the composite to v2.

Advertising disclosure: This page may contain links to third-party platforms. Where OddsReference has an active, disclosed partnership with a platform, we may earn a commission if you click a link and complete a qualifying action — at no extra cost to you. Affiliate status has no bearing on our data, rankings, or editorial conclusions. See our Affiliate Disclosure and Editorial Policy for details.

01 — MEASURED METRICS

How Do The Exchanges Compare Today?

The metrics we can measure per venue, side by side. No composite score — the numbers speak for themselves.

Every value below is measured directly from our ingestion pipeline and published fee schedules. Where a metric is not observable for a venue, the cell says so plainly rather than showing a zero or a guess.

Venue Active markets Category coverage Published fee Observed median spread All-in cost
Kalshi 6,511 6/6 1.75% 5% 6.75%
Polymarket 1,528 6/6 0% Not observed Incomplete — spread not observed
Gemini 339 6/6 0% 2% 2%

For Polymarket we ingest the CLOB midpoint price, not its full order book, so we do not capture its bid-ask spread. Its published trading fee is 0%, but its true all-in cost depends on a spread we cannot yet observe — so we report the spread as “Not observed” and its all-in cost as incomplete, rather than implying a 0% cost that would flatter it against venues whose spread we do measure.

Data Quality & What We Omit

This v1 payload measures only 3 factors. LIQUIDITY (per-venue 30d volume) and SETTLEMENT (dispute rate) are intentionally excluded because the underlying data is unreliable or absent: per-venue volume in source_prices is known-bad (Kalshi reports ~$812K all-time, a data bug), and the resolution_disputes table has 0 rows. Metaculus is excluded entirely — it is a forecasting aggregator, not a tradeable exchange, so cost/spread are undefined for it.

Polymarket spread is unobservable from our data: we ingest only the CLOB midpoint, so Polymarket source_prices rows carry no bid/ask. Its medianSpreadPct is null and spreadObserved is false. Kalshi and Gemini spreads are the median (percentile_cont 0.5) of (ask - bid) over quotes observed in the last 3 days where bid>0, ask>0, ask>=bid.

Measured in v1

  • • active_markets
  • • category_coverage
  • • effective_cost

Omitted (deferred to v2)

  • • liquidity
  • • settlement
02 — RUBRIC & ROADMAP

What Will The Composite OR Score Measure?

Five reproducible factors. Three are measured in v1; two arrive with the v2 composite.

The composite OR Score will be a weighted blend of five factors, each measured from data. Three of them already drive the v1 comparison above (market breadth, category coverage, and cost where the spread is observable). The remaining two — liquidity depth and settlement reliability — need new per-venue measurement before any blended number is honest. The weights and how we measure each factor are fixed and published here.

Factor Weight v1 status How we measure it
Cost efficiency lower-is-better · all-in % of notional 25% Measured where observable (v1) Effective all-in trading cost: explicit per-contract/commission fees plus the typical bid-ask spread we observe on that venue, expressed as a percent of notional. Measured from published fee schedules and our own order-book snapshots.
Liquidity depth higher-is-better · trailing-30d USD volume 25% Coming in v2 Venue-level trailing-30-day traded volume in USD, aggregated across the markets we ingest, then log-scaled. Deeper liquidity means tighter fills and more trustworthy prices.
Market breadth higher-is-better · active markets 20% Measured (v1) Count of currently active markets on the venue, log-scaled. Measures how much a trader can actually do on the venue, independent of how much each market trades.
Category coverage higher-is-better · of 8 categories 15% Measured (v1) Number of the 8 canonical categories (politics, economics, sports, crypto, tech, science, climate, entertainment) in which the venue has real market activity, as classified by our ingestion pipeline.
Settlement reliability higher-is-better · 0-1 reliability 15% Coming in v2 A 0-1 reliability score combining resolution transparency (are criteria pre-defined and disclosed?) with our measured settlement track record (dispute rate and time-to-settle on resolved markets). Regulated venues with pre-defined criteria and low dispute rates score highest.
Total 100% Weights sum to 100% by construction and are checked in our test suite. They apply to the deferred v2 composite, not to any v1 number.

When the composite launches, each subscore will be computed against a fixed measured range so results are reproducible: cost maps a 0% all-in cost to full marks and a 6%+ all-in cost to zero; liquidity depth and market breadth are log-scaled; category coverage counts activity across our 8 canonical categories; settlement reliability blends resolution transparency with our measured dispute and settlement-time record. The composite stays deferred until all five factors are measurable across every venue.

03 — RATIONALE

Why These Five Factors?

Each factor is something a trader can feel, and something we can measure.

We chose factors that are both decision-relevant and reproducible. If a factor could not be measured from our ingestion pipeline or a published fee schedule, it did not make the rubric — that constraint is what keeps the OR Score a data product rather than an editorial ranking, and it is exactly why two factors are held back from v1 until we can measure them for every venue.

Cost efficiency · 25%

Fees plus typical spread are the clearest drag on a trader's return. We measure the fee for every venue today; the spread component is only measured where we observe an order book, which is why Polymarket's all-in cost is still incomplete.

Liquidity depth · 25%

Deep liquidity means tighter fills and prices you can trust as probability estimates. Measured as trailing-30-day venue volume, log-scaled. Deferred to v2 pending a per-venue volume rollup.

Market breadth · 20%

How much you can actually do on the venue. A count of active markets, independent of how much each one trades. Measured live in the v1 table above.

Category coverage · 15%

Whether the venue serves one niche or the full spread of politics, economics, sports, crypto, and more. Measured live in the v1 table — though today it is a flat maximum across venues, so it cannot yet separate them.

Settlement reliability · 15%

Do markets resolve on pre-defined, disclosed criteria, and do they settle cleanly and on time? We blend resolution transparency with our measured dispute and settlement-time record. Deferred to v2 pending computation from resolved-market history.

04 — TRUST

Do Affiliate Relationships Affect The OR Score?

No. The method has no term for commercial relationships.

No. The OR Score is built purely from the measured factors above. Whether OddsReference has an affiliate partnership, an advertising deal, or no relationship at all with a venue carries zero weight. A venue we earn nothing from and a venue we partner with are measured by the identical method, from the identical data.

Trust Commitments

  • • Commercial relationships are never an input to any metric or score.
  • • The rubric, weights, and normalization anchors are published on this page and versioned in code.
  • • We never publish a number built from data we do not actually have — where a metric is unobservable, the cell says “Not observed,” and the composite score stays deferred until every factor is measurable.
  • • Any active partnership is disclosed on our Affiliate Disclosure, and our editorial standards live in our Editorial Policy.
05 — ROADMAP TO V2

What Still Needs Sourcing For The Composite?

The v1 comparison is live. The v2 composite waits on four data items.

The single ranked OR Score publishes once every factor is measurable across every venue. Two factors already feed the v1 comparison. The rest — and one gap that specifically blocks a fair cost comparison — need wiring before a blended number is honest.

Factor input Data status Source to wire
Active markets Live in v1 Per-venue count of active markets, already in the comparison table above.
Category coverage Live in v1 Per-venue distinct categories with activity. Live, but a flat maximum today, so it cannot yet discriminate.
Cost — Polymarket spread Blocking Polymarket bid-ask ingestion, so its spread is observed like Kalshi's and Gemini's and its all-in cost is comparable.
Liquidity depth Partial Per-market volume exists; needs a trailing-30d per-venue rollup.
Settlement reliability To build Transparency is a structural fact; dispute rate and time-to-settle need computing from resolved-market history.

Until every row above reads “Live,” we publish the measured metrics and no composite. This is the same standard we hold our content to: see our Editorial Policy.

What The OR Score Rates (And What It Doesn't)

The OR Score applies to prediction-market exchanges — Kalshi, Polymarket, and the other venues we ingest — because those are the venues we have reproducible data for. We do not issue an OR Score for casinos, daily fantasy sites, or poker rooms. We have no live performance feed for those products, so scoring them would mean inventing numbers. Those verticals stay a structural feature comparison (see casino, DFS, and poker) until real data exists.

Frequently Asked Questions

What is the OR Score?
The OR Score is our framework for rating prediction-market exchanges (Kalshi, Polymarket, Gemini, and the other venues OddsReference ingests) from reproducible data — market breadth, category coverage, cost, liquidity, and settlement reliability. In v1 we publish the measured metrics as a transparent comparison. We do not yet publish a single ranked composite number.
Why is there no single OR Score number yet?
A single composite would mislead right now. Category coverage is a flat 6 of 6 across every venue we score, so it cannot discriminate between them. And we do not observe Polymarket’s bid-ask spread (we ingest its CLOB midpoint, not its order book), so its all-in cost is unknown — a composite would structurally flatter the venue whose true cost we cannot see. So v1 publishes the metrics we can measure, honestly labeled, and defers the composite to v2.
When does the composite OR Score launch?
The composite launches in v2, once cost is measurable across every venue (which needs Polymarket bid-ask ingestion so its spread is observed like Kalshi’s and Gemini’s) and per-venue liquidity and settlement-reliability data exist. Until every factor is measurable for every venue, ranking them by a blended number would be a guess dressed as data.
Do affiliate or advertising relationships affect the OR Score?
No. The OR Score is built only from measured venue metrics. Whether OddsReference has a commercial relationship with a venue has zero weight and never adjusts a metric or a future score up or down. A venue we earn nothing from and a venue we partner with are measured by the identical method, from the identical data.
Does OddsReference score casinos, DFS sites, or poker rooms with the OR Score?
No. The OR Score covers prediction-market exchanges only, because those are the venues we have reproducible performance data for. Casino, DFS, and poker pages stay feature-comparison-only — structural facts cited from our reviews, never a numeric score — until we have real performance data for those products.
06 — RELATED

Related Reading

Some links on this page may be affiliate links. If a partnership is active, OddsReference may earn a commission at no extra cost to you — see our Affiliate Disclosure. This never affects our data or rankings; see our Editorial Policy.