kerem@ai-investor:~$ $ calibration --grade
Calibration
Does conviction actually predict forward returns? Graded against passive benchmarks.
Does the screener's own conviction rating actually predict forward returns, and does it beat just holding the passive benchmark? Reported exactly as computed -- if a tier is underperforming, that shows here, not just the tiers that make the system look good.
Large-Cap (9 archives, benchmark SPY)
| Tier | N | Avg Fwd Return | Std Dev | Win Rate | Alpha vs Benchmark |
|---|---|---|---|---|---|
| HIGH | 19 | -3.00% | 5.9% | 37% | -5.98% (26% beat it) |
| MEDIUM | 101 | -2.76% | 13.4% | 42% | -6.13% (28% beat it) |
| LOW | 2 | +0.28% | 0.2% | 100% | +0.60% (100% beat it) |
Small-Cap (8 archives, benchmark IWM)
| Tier | N | Avg Fwd Return | Std Dev | Win Rate | Alpha vs Benchmark |
|---|---|---|---|---|---|
| HIGH | 2 | -23.73% | 5.8% | 0% | -24.41% (0% beat it) |
| MEDIUM | 62 | -7.23% | 17.3% | 34% | -8.37% (31% beat it) |
| LOW | 13 | +3.41% | 11.2% | 69% | +1.73% (62% beat it) |
Crypto (9 archives, benchmark BTC-USD)
| Tier | N | Avg Fwd Return | Std Dev | Win Rate | Alpha vs Benchmark |
|---|---|---|---|---|---|
| MEDIUM | 23 | +7.30% | 12.0% | 65% | +0.68% (48% beat it) |
| LOW | 5 | -16.40% | 40.8% | 40% | -17.97% (40% beat it) |