Results · Signals · validation record
Breakout & regression signal — the validation record
What the signal measures, what it is worth against comparable players, and what it does not claim. Every call we have published is listed on the call tracker — this page is the measurement behind it.
The effect scales with the size of the gap.
Three independent constructions, all held out, all clearing 90%. Every figure is measured against comparable players at the same production level — never against the league. A low-production player rises about 104 OPS points in 60 days with no signal at all, so an unmatched comparison overstates the effect by roughly a third.
The ladder is the finding. A model that only knew whether a player was diverging would produce one flat number. These three rise with the size of the gap — the signal carries magnitude, not just direction.
Two records. They answer different questions.
The historical replay is resolved and closed. The forward record is open and accruing. They are never pooled, averaged, or shown as one number.
Now reading: the historical replay
The historical replay — 83 scored out-of-sample weeks across the four resolved seasons (2022–2025), each week's signal produced by a model trained only on outcomes that closed before that week. Reconstructed, but leak-safe by construction. The promotion rests on this.
A season whose interval crosses zero is not distinguishable from no effect. It is shown, kept separate from the resolved seasons, and never averaged into them.
4 of 4 resolved seasons clear zero · 83 scored out-of-sample weeks
Calibration — not just whether, but how much.
Predicted move against realized move, on published calls only. Measured in OPS points, the unit the product publishes in.
Ten conviction deciles, 3,057 closed calls. Most buckets land near the diagonal, but the two highest-conviction deciles sit above it — there the model predicts about +28 points and realizes about +16. That is over-prediction where conviction is strongest, and it is shown rather than smoothed. Pooled coverage and bias are beside it.
30-day · 4,294 closed windows
80%
of outcomes fell inside the published 80% interval
Mean absolute error 131 points · bias −1 points
60-day · 3,057 closed windows
81%
of outcomes fell inside the published 80% interval
Mean absolute error 101 points · bias −3 points
Record computed Invalid Date
Coverage lands on nominal at both horizons and bias is about zero: the intervals are the width they claim to be, and the point estimate is not systematically high or low.
Tier separation · reached +75 OPS points over the 60-day window
62.8% of true-breakout calls (n=669), against 29.8% for the flagged-but-no-edge tier (n=2,388).
Improvement candidate is NOT a positive call. Measured against the unflagged base it sits at or below it in all five replayed seasons; it appears here as the contrast that shows the tiers separate, not as a claim of its own.
This is the metric the four-season record above is stated in, which is the only reason a threshold rate appears anywhere. It is an aggregate bridge, never a verdict on a single call.
What this signal does not claim.
Five of them, stated at the same size as the results above. If any of these changes, it changes here first.
The effect scales with the size of the gap.
Three independent constructions, all held out, all clearing 90%. Every figure is measured against comparable players at the same production level — never against the league. A low-production player rises about 104 OPS points in 60 days with no signal at all, so an unmatched comparison overstates the effect by roughly a third.
The ladder is the finding. A model that only knew whether a player was diverging would produce one flat number. These three rise with the size of the gap — the signal carries magnitude, not just direction.
Two records. They answer different questions.
The historical replay is resolved and closed. The forward record is open and accruing. They are never pooled, averaged, or shown as one number.
Now reading: the historical replay
The historical replay — 83 scored out-of-sample weeks across the four resolved seasons (2022–2025), each week's signal produced by a model trained only on outcomes that closed before that week. Reconstructed, but leak-safe by construction. The promotion rests on this.
A season whose interval crosses zero is not distinguishable from no effect. It is shown, kept separate from the resolved seasons, and never averaged into them.
4 of 4 resolved seasons clear zero · 83 scored out-of-sample weeks
What this signal does not claim.
Five of them, stated at the same size as the results above. If any of these changes, it changes here first.
Calibration — not just whether, but how much.
Predicted move against realized move, on published calls only. Measured in OPS points, the unit the product publishes in.
Ten conviction deciles, 3,057 closed calls. Most buckets land near the diagonal, but the two highest-conviction deciles sit above it — there the model predicts about +28 points and realizes about +16. That is over-prediction where conviction is strongest, and it is shown rather than smoothed. Pooled coverage and bias are beside it.
30-day · 4,294 closed windows
80%
of outcomes fell inside the published 80% interval
Mean absolute error 131 points · bias −1 points
60-day · 3,057 closed windows
81%
of outcomes fell inside the published 80% interval
Mean absolute error 101 points · bias −3 points
Record computed Invalid Date
Coverage lands on nominal at both horizons and bias is about zero: the intervals are the width they claim to be, and the point estimate is not systematically high or low.
Tier separation · reached +75 OPS points over the 60-day window
62.8% of true-breakout calls (n=669), against 29.8% for the flagged-but-no-edge tier (n=2,388).
Improvement candidate is NOT a positive call. Measured against the unflagged base it sits at or below it in all five replayed seasons; it appears here as the contrast that shows the tiers separate, not as a claim of its own.
This is the metric the four-season record above is stated in, which is the only reason a threshold rate appears anywhere. It is an aggregate bridge, never a verdict on a single call.