The whole pitch is verify, don't trust. So here's exactly how a forecast is made, resolved, and graded — nothing hidden. ← the forecasts · ask the Oracle (members) →
1.How a forecast is made — incentives, lenses & physics
The core idea is Freakonomics applied to forecasting: people respond to incentives, so to predict mass behavior you map the incentives — then reason about them with structure, not vibes. For each question the engine:
Coarse-grains to the order parameters — first it names the 1–4 macro variables that actually flip the outcome and ignores the noise. (Borrowed from physics: find the few slow variables that govern the system.)
Maps the incentive structure — who are the actors, and what are their REAL incentives (economic, social, moral — and the hidden or perverse ones)? That incentive map is shown on every forecast.
Selects from a lens library — a dozen simplifying theories (incentives & public choice, veto-players / status-quo bias, game theory, reflexivity, base rates, diffusion & tipping points, power laws, preference-falsification cascades, and more). It applies the ones that genuinely fit and reasons through them, tagging the actor archetypes + incentive patterns at play from a reusable knowledge core.
Runs the physics — it builds the actor coupling graph (who pushes which way, who is allied vs adversarial) and solves it as an Ising / spin-glass system. That yields a structural estimate plus a frustration score — how genuinely contested the outcome is. High frustration = a deadlocked system, and the engine refuses to fake confidence.
Aggregates three independent stances — a focused single-lens read, a multi-lens synthesis, and a skeptic that steelmans the opposite — pooled in log-odds space (the geometric mean of odds), then nudged away from 50% only as far as the stances' agreement and the structure's coherence justify. When they split, it regresses toward the historical base rate instead of inventing certainty.
Grounds it in live data + today's news (web search, no stale cutoff) and stays calibration-first: a humble, correct "coin-flip" beats a confident wrong number.
Every forecast records its probability, the incentive map, the order parameters, the lenses + archetypes + patterns it used, the actor coupling + frustration, the three stance passes, the model + engine version, and an immutable timestamp. You can watch the whole thing happen live in the members Oracle.
2.How a forecast resolves
Resolution is always from an external, objective source — never our opinion:
Prediction markets (Manifold) for mirrored questions.
FRED (Federal Reserve data) for macro — fed funds, CPI.
On-chain reads for crypto — price, liquidity, gas, straight from the chain.
Price feeds for asset levels.
A daily job checks each open forecast after its resolution date and records the outcome. We can't move the goalposts because we don't own the goalposts.
3.How we're scored
Every resolved forecast gets a Brier score = (probability − outcome)², where outcome is 1 (yes) or 0 (no). Lower is better — a perfect call scores 0, a confident-and-wrong call scores near 1. We score three ways:
vs the outcome — were we right, and how confidently?
vs the base rate — did our reasoning beat just guessing the historical average?
vs the crowd — for mirrored markets, was our Brier lower than the market's price at forecast time? That's the real test: did we add value?
4.The honesty guarantees
Write-once. A forecast's probability, reasoning, and timestamp can never be edited after it's made. The ledger physically refuses to overwrite them.
External resolution. Outcomes come from sources we don't control.
Hits and misses published. The scoreboard shows everything, including where we were wrong.
Versioned. Every forecast is stamped with the exact engine version + model, so as the engine improves you can tell which generation made which call.
5.What this is — and isn't
This is experimental, running on an ensemble of free models (current engine: 0.4-incentive-blind). It is research, not financial advice. The point isn't to claim we're oracles — it's to build a public, falsifiable track record and let you judge us by it. Until the forecasts resolve and the scoreboard fills, treat every number as an unproven hypothesis.
6.Check it yourself
The raw data is machine-readable for programmatic + agent use:
The model reasons from incentives + history. A planned next layer: cross-check selected predictions against live open-source ground truth — an OSINT aggregator (flight paths, marine/AIS, satellites, traffic cameras, etc.). If an incentive read implies, say, forces or ships moving, real-world signals can confirm or contradict it. To be added only where it demonstrably sharpens accuracy — the goal is maximum accuracy at minimum compute, not data for its own sake.