Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Tokenization is done on the CPU. Models never see the raw characters. That's why you get trick questions like the number of r's in strawberry.

There are many research papers on models using characters directly. One challenge is that effective context length is smaller.



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: