Graded against reality.
Every probability we publish comes from a sensor we can score: when a prediction market resolves, reality hands us an answer key we did not write. This page is the referee’s ledger — the crowd consensus we cite, and the deterministic engine we run on top of it, graded with proper scoring rules. As of 2026-08-08.
We grade to keep the probabilities we publish honest, not to outguess the crowd. Markets are a sensor we cite, never an opponent we score against. Matching the crowd here is corroboration: an independent check on a price that is sometimes built on thin volume, not a failed edge test. That is also why these truth surfaces are free. We do not sell an edge; we sell maintained, verified currency, and this ledger is how we prove the maintenance works. Case-selection methodology →
The living benchmark
Forecasts are logged while each market is still open, then graded only after reality resolves it — the record cannot be back-filled, because the forecast exists before the answer does. So far 51,372 open-market forecasts have been captured and 31,187 have matured and been scored.
| forecaster | Brier (95% CI) | log loss | skill vs base rate | null (same markets) | markets |
|---|---|---|---|---|---|
| Crowd consensus (the sensor we cite) | 0.0643 [0.0626, 0.066] | 0.207 | 0.4651 | 0.1201 | 31,187 |
| Our engine (crowd + deterministic correction) | 0.0642 [0.0625, 0.066] | 0.2071 | 0.4654 | 0.1201 | 31,187 |
Read each Brier against its null. In this set 13.1% of markets resolved yes, so a low absolute Brier is partly the rarity of yes outcomes — skill is the honest comparison. The null column is the base rate scored on that row’s own markets, the exact figure its skill divides by (skill = 1 - Brier / null), so the arithmetic checks out row by row. The out-of-sample table below grades the base rate as its own row.
Calibration curve, as a table
A calibrated forecaster’s 30% calls should come true about 30% of the time. Each row groups the crowd’s matured forecasts by stated probability and shows what reality delivered — this is the resolved-outcome table, bin by bin.
| stated probability | resolved markets | mean prediction | came true | read |
|---|---|---|---|---|
| 0–10% | 19,998 | 1.6% | 1.1% | calibrated |
| 10–20% | 2,656 | 14.4% | 11.2% | overconfident by 3.2% |
| 20–30% | 2,063 | 24.8% | 21% | overconfident by 3.8% |
| 30–40% | 1,678 | 34.7% | 30.4% | overconfident by 4.3% |
| 40–50% | 1,545 | 44.8% | 31.5% | overconfident by 13.3% |
| 50–60% | 929 | 53.4% | 46.1% | overconfident by 7.4% |
| 60–70% | 363 | 64.5% | 60.6% | overconfident by 3.9% |
| 70–80% | 351 | 74.9% | 68.4% | overconfident by 6.5% |
| 80–90% | 372 | 85% | 84.4% | calibrated |
| 90–100% | 1,232 | 96.6% | 98.1% | calibrated |
The out-of-sample check
A second benchmark guards against fitting the past: 34,325 resolved markets split by resolution date, the engine fitted only on the earlier half and scored only on the later half, against a base-rate null. This is a fixed benchmark, generated 2026-06-23 over a fixed window — it re-runs when the engine version changes, not daily.
| forecaster | Brier (95% CI) | log loss | skill vs base rate | null (same markets) | markets |
|---|---|---|---|---|---|
| Base rate (the null) | 0.1063 [0.103, 0.1097] | 0.3699 | 0 | 0.1063 | 17,163 |
| Crowd consensus | 0.0705 [0.0658, 0.0756] | 0.2339 | 0.5555 | 0.1586 | 4,069 |
| Our engine | 0.07 [0.0652, 0.0751] | 0.2318 | 0.5589 | 0.1587 | 4,069 |
Reconciling the numbers: each forecaster is graded only on the markets it actually priced — the base rate covers every resolved market, the crowd only those it quoted — so each row’s skill is measured against the base rate on that row’s own markets (the null column), not against the standalone base-rate row above it. That is why skill = 1 - Brier / null checks out row by row while the base-rate row’s own Brier, scored over a larger market set, differs. No LLM sits in any forecaster here: this table grades the crowd we cite and the deterministic engine we run on it. The Answer Engine — the paid product — is a separate, younger cohort graded per answer, below; nothing here claims to measure it.
The honest reading: the crowd is a strong, well-calibrated sensor, which is exactly why we cite it, and our engine tracks it at parity. Parity is the point. It means the probabilities we attach to maintained records are as good as the best public sensor, graded in the open. We publish the grade either way; the referee runs whether or not we like the score.
Answer receipts — the per-answer contract
The tables above grade probabilities in public. Answer quality is graded differently — on the answer itself, for the person who asked it. Every answer the desk gives carries its own verification result (verified), its own honesty state (abstained when no market, receipt, or citation grounds it), and the source receipts it rests on — so a caller, human or machine, can accept or quarantine each answer on its own flags, at the moment it arrives. We also run a standing internal quality program over these signals — it drives the engineering queue daily — and what it changes shows up here the only way that counts: in the flags on the next answer.
Reproducibility
The full dataset behind this page — every score, confidence interval, and calibration bin — ships as machine-readable JSON. The forward slice regenerates daily as new markets mature (this build: 2026-08-08); the out-of-sample slice is a fixed benchmark (2026-06-23) that re-runs on engine changes. Each slice carries its own generatedAt in the JSON — freshness is per-component, never one blanket date.
Informational only — not financial, legal, or investment advice. Prediction-market prices are shown as a signal of what the crowd believes.