The real results, with no cherry-picking.
Every daily pick is tracked, resolved and kept visible. The model uses this history to become stricter as it learns.
Green = ≥ 60% · Amber = 50 to 59% · Red = < 50% · Minimum 5 picks to qualify as representative.
When our AI says 80% confident, does it actually win 80% of the time? The gap between stated and actual is the calibration error. A smaller gap means the AI is being more honest about what it knows.
Gap = Actual win% − AI confidence midpoint. Negative means AI overstated confidence. Trend toward 0 = system improving its self-awareness.
Every month the system saves a snapshot of its current state. A rising threshold means the AI got stricter. A calibration error trending toward zero means it's getting more honest.