PROPHET 9
We measure whether a hitter’s underlying skill has diverged from his box score — and we publish every call.Four-season public record →

Results · Signals · validation record

Breakout & regression signal — the validation record

What the signal measures, what it is worth against comparable players, and what it does not claim. Every call we have published is listed on the call tracker — this page is the measurement behind it.

The effect scales with the size of the gap.

Three independent constructions, all held out, all clearing 90%. Every figure is measured against comparable players at the same production level — never against the league. A low-production player rises about 104 OPS points in 60 days with no signal at all, so an unmatched comparison overstates the effect by roughly a third.

wOBA lift vs matched control, held-outConviction →
Highest-conviction signals
the extremes of the gap distribution
+.065
Breakout vs regression
flagged up minus flagged down
+.035
All flagged players
any divergence, either direction
+.028

The ladder is the finding. A model that only knew whether a player was diverging would produce one flat number. These three rise with the size of the gap — the signal carries magnitude, not just direction.

Two records. They answer different questions.

The historical replay is resolved and closed. The forward record is open and accruing. They are never pooled, averaged, or shown as one number.

Now reading: the historical replay

The historical replay — 83 scored out-of-sample weeks across the four resolved seasons (2022–2025), each week's signal produced by a model trained only on outcomes that closed before that week. Reconstructed, but leak-safe by construction. The promotion rests on this.

SeasonLift over matched controlLift
2022+19.5ppCI excludes zero [+15.6pp, +23.4pp]
2023+21.1ppCI excludes zero [+8.9pp, +33.2pp]
2024+32.5ppCI excludes zero [+21.8pp, +43.0pp]
2025+22.8ppCI excludes zero [+11.4pp, +33.7pp]
2026+8.9ppCI crosses zero [−5.9pp, +24.2pp]

A season whose interval crosses zero is not distinguishable from no effect. It is shown, kept separate from the resolved seasons, and never averaged into them.

4 of 4 resolved seasons clear zero · 83 scored out-of-sample weeks

Stated limits

What this signal does not claim.

Five of them, stated at the same size as the results above. If any of these changes, it changes here first.

Last reviewed 2026-08-02 · next review October 2026
01 · Unresolved
The forward record is not evidence yet.
It has not reached its pre-registered read threshold. Until it does, the only validated numbers on this page are the historical replay's. No interim figure is published — a progress count invites the reader to extrapolate a record from it.
02 · Under watch
2026 reads +8.9pp, and its interval crosses zero.
Weaker than every prior season and not distinguishable from no effect, on 31 de-replicated calls across 9 closed weeks — most of the season's windows have not matured. It is not called a decline and not called noise: it is flagged, kept out of the four resolved seasons, and re-read in October 2026.
03 · Publishing nothing yet
Daily calls carry no figures while the rebuild is scored.
The old engine was defective, so its record was withdrawn; the replacement is mid walk-forward. If the new record does not clear its pre-registered gates, nothing is published at all — the design does not assume success.
04 · Out of scope
The pitcher side is an unvalidated axis.
Everything here is hitters — expected contact quality against surface results. Pitcher divergence carries no validation record, so it carries no claim and no tiers.
05 · Where it is silent
Small gaps get no reading at all.
The edge is concentrated in the extremes. Where a player's skill and surface lines agree inside the noise band the model returns no divergence rather than a small number — an absence, with its reason, never a neutral-looking midpoint.

Calibration — not just whether, but how much.

Predicted move against realized move, on published calls only. Measured in OPS points, the unit the product publishes in.

Predicted vs realizedPERFECT CALIBRATIONdecile 1 · n=305 · predicted 46p · realized 41pdecile 2 · n=306 · predicted 34p · realized 37pdecile 3 · n=306 · predicted 25p · realized 30pdecile 4 · n=305 · predicted 34p · realized 40pdecile 5 · n=306 · predicted 31p · realized 34pdecile 6 · n=306 · predicted 26p · realized 14pdecile 7 · n=305 · predicted 29p · realized 29pdecile 8 · n=306 · predicted 25p · realized 23pdecile 9 · n=306 · predicted 29p · realized 17pdecile 10 · n=306 · predicted 27p · realized 15p

Ten conviction deciles, 3,057 closed calls. Most buckets land near the diagonal, but the two highest-conviction deciles sit above it — there the model predicts about +28 points and realizes about +16. That is over-prediction where conviction is strongest, and it is shown rather than smoothed. Pooled coverage and bias are beside it.

30-day · 4,294 closed windows

80%

of outcomes fell inside the published 80% interval

Mean absolute error 131 points · bias 1 points

60-day · 3,057 closed windows

81%

of outcomes fell inside the published 80% interval

Mean absolute error 101 points · bias 3 points

Record computed Invalid Date

Coverage lands on nominal at both horizons and bias is about zero: the intervals are the width they claim to be, and the point estimate is not systematically high or low.

Tier separation · reached +75 OPS points over the 60-day window

62.8% of true-breakout calls (n=669), against 29.8% for the flagged-but-no-edge tier (n=2,388).

Improvement candidate is NOT a positive call. Measured against the unflagged base it sits at or below it in all five replayed seasons; it appears here as the contrast that shows the tiers separate, not as a claim of its own.

This is the metric the four-season record above is stated in, which is the only reason a threshold rate appears anywhere. It is an aggregate bridge, never a verdict on a single call.