Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

i feel that the proposed approach is close to becoming "demo prawn"* territory.

similar to cpu benchmarks (e.g. geekbench, which also had a recent update), the goal is to distinguish model in some arbitrary thing. however, it is not quantifiable but subjectively even the latest show distinctive differences regardless.

large context window is still relatively new and not "feasible" even commercially, so making that a requirement for a benchmark would not make it accessible especially for open-source ones.

* yes, i misspelled it intentionally



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: