Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Author here, agreed that a single generation isn’t meaningful. That’s why the benchmark is designed around distributions rather than individual outputs.

The initial calibration screened 2,336 questions with 4 samples each, then selected the 78 questions where Opus 5.5 showed useful variance. The panel is run daily, and the actual decision is based on paired per-item differences across 10-day windows with clustered standard errors, not on any single day’s result.

The n=1 in the daily sampling rate means one sample per item per day, not one sample for the experiment. By the time a window is evaluated there are hundreds of observations, and a change has to clear a pre-registered 99% interval in two consecutive windows before LiveNerf calls it a change.

The nondeterminism is basically the reason the statistical part exists in the first place.

 help



Consider applying for YC's Winter 2027 batch! Applications are open till November 2.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: