Calibration & honesty
When we said 65%, were we right 65% of the time?
The idea
A prediction tool that never checks its own answers is an entertainment product. The defining feature of a serious one is a feedback loop: every forecast is stored, resolved against what the market actually did, and scored — automatically, including the embarrassing ones. Public track record pages exist because the numbers exist; the numbers exist because resolution is built into the engine, not bolted on.
How a prediction gets scored
When a forecast's horizon elapses, the engine fetches the realised price path and computes the realised return. Direction calls are marked hit or miss — with a dead-band around zero, because calling a coin-flip on a move of a few hundredths of a percent teaches the system nothing and inflates both hit and miss counts with noise. Probability quality is measured with Brier scores, which punish confident wrongness far more than honest uncertainty, and the forecast cone is checked for coverage — a 90% interval that contains reality only 70% of the time is overconfident and gets widened.
Auditing each channel separately
The ensemble's four channels are audited individually with rank information coefficients — a statistic quantifying whether a channel's scores actually rank future returns better than chance. A channel whose coefficient decays toward zero is quietly losing its edge and its ensemble weight; one that strengthens earns more. This is how the system adapts to changing markets without anyone hand-tuning knobs on vibes.
Segmented, because averages hide sins
A single global accuracy number can conceal a system that is excellent on daily equity charts and useless on one-minute memecoins. Calibration is therefore tracked per asset class, per timeframe and per regime — the same segmentation the live engine uses — so a weakness in one cell can't hide behind strength in another.
What this buys you
Not certainty — nothing buys that. It buys meaningful probabilities: when the engine shows a wide cone and modest confidence, that reflects measured disagreement in the evidence, and when it shows conviction, that conviction has a track record behind it. Start with the methodology overview or see the ensemble's inputs: pattern matching, technical indicators and sentiment.