Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

they are cleverly arranged / configured next-most-likely-token predictors, possibly with some clever procedures / attachments on top.
 help



Nope. This isn't right.

> "it's not just a next-token predictor because a bunch of the training isn't about predicting the next token."

clever procedures on top of the base transformer architecture.

i used simplified words/phrases to summarise the same thing you two were saying (the intent being: here's a version that may be digestible when discussing with others).

apparently that means i'm wrong though, no idea why because it seems you've decided to be dismissive rather than constructively elaborate on why this simplified and digestible version might be wrong :shrug:


They aren't predicting the next token. It's quite literally not a prediction.

They're estimating a probability distribution over the next token, from which a sample is taken. Close enough.

It's not an estimation of something. It's a policy.

Sure, you're right.

But it's a policy learned from a next-token prediction task. You could also call it an inferrer or generator or whatever. The point is that it takes as input a sequence of preceding tokens and emits one more token to continue the sequence.


The policy is not learned token by token during RLHF and RLVR. The reward model doesn't score token by token.

We're talking to a RLVR bot.

Bad bot.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: