Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

The calibration section is the most important part of this article for anyone actually shipping classifiers. nzoschke's email use case is exactly where this bites: for production routing, you almost always want to threshold on confidence ('auto-handle above 0.9, route to a human below'), and that only works if the probabilities mean what they say. A 96% model with overconfident outputs is operationally worse than a 94% model with honest ones. The Guo et al. observation Raschka cites — networks can overfit NLL without overfitting 0/1 loss — is why this happens with plain fine-tunes, and it's also the strongest part of Jev's design: training the confidence directly (RLCD/RLCR-style) rather than bolting calibration on after the fact. One caveat I'd add to the IMDb numbers: as the article notes, we don't know whether IMDb was in Jev's synthetic training mix. Until someone runs these benchmarks on a private, never-published dataset, take the 96% as a ceiling, not a measurement.
 help





Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: