Any forecasting system evaluated after the fact can be made to look good. Drop the questions that went badly, choose the baseline you happened to beat, report the metric that flattered you. None of that requires dishonesty. A few reasonable-looking choices in sequence get you there.
So the choices were made in advance and filed: the question set, the baselines, the metric, the interval method, and the bar. If we miss it, the miss goes on this page in the same size type as a hit.
Pre-registration is ordinary practice in empirical science and nearly unheard of in this industry. It costs nothing except the ability to quietly move the goalposts later.
SET150 binary questions resolving after the model's training cutoff
BASEalways-50% at 0.250, and the crowd forecast at question open
METRICmean Brier across all 150, no exclusions
CIbootstrap, 10,000 resamples, clustered by underlying event
BARbeat the crowd baseline, 90% interval excluding zero
Scoreboard · evaluation oneopens Q4
Mean Brier, 150 questions
vs. always-50%
vs. crowd at question open
Calibration gap, widest band
Reviews held, no change
Pre-registered bar metpending
Protocol filed before the first runpublished either way