The way he proposed to deal with "How urgent is this customer support message?" is "a score over ordered levels".
I'm not sure that's a good way to use Jev for this use case. If I had this problem, I would ask Jev to answer several different yes/no questions about each message, and then use the probabilities as inputs into a logistic regression model that predicts urgency.
If you ask the model specific yes/no questions which can be answered reasonably objectively from the input, I think the answers are going to be more stable over successive generations of models.
e.g. if you ask 'Is the customer angry?' I'd expect that answer to have high agreement between models and between models and humans. But directly answering the 'is it urgent' question is much harder. (Although I suppose you can try to put the rules in the prompt.)
I'm not sure that's a good way to use Jev for this use case. If I had this problem, I would ask Jev to answer several different yes/no questions about each message, and then use the probabilities as inputs into a logistic regression model that predicts urgency.
If you ask the model specific yes/no questions which can be answered reasonably objectively from the input, I think the answers are going to be more stable over successive generations of models.
e.g. if you ask 'Is the customer angry?' I'd expect that answer to have high agreement between models and between models and humans. But directly answering the 'is it urgent' question is much harder. (Although I suppose you can try to put the rules in the prompt.)