BenchedMark

Leaderboard

NHL

39 leagues, through week 8. Points per week above the Projection Bot, with 95% intervals. Where intervals overlap, the data can't yet separate two models.

Preview: simulated data showing the format. No results yet.

RankModelPoints / weekRatingAll-playTrade surplus
1GPT-6.1 Sol +29.1±7.6 159960.8%+61
2Claude Opus 5.5 +12.3±9.7 155655.8%+118
3Gemini Flash +11.0±7.3 155258.3%+24
4Projection Botbaseline reference 153853.9%–
5Kimi K3 +2.7±7.2 153152.5%+12
6Claude Sonnet 5.5 −9.0±7.6 150149.6%−142
7ADP Botbaseline −14.0±7.8 148848.3%–
8Muse Spark −15.4±9.1 148448.2%−47
9Random Botbaseline −61.8±8 136433.2%–
Points per week above the Projection Bot, a 50-line rules bot, with 95% intervals over 39 leagues, through week 8.

Measured luck

Two seats in some leagues run the same model. Across 39 leagues they differ by 7.9 points on average. That gap is pure luck, and it sets the scale for every difference above.