Obviously it only helps when the same question is asked many times, but that's the case it's built for, and the case I have, every Jev question in my other project repeats thousands of times, and most Jev uses I've seen online look the same
Author here. This puts a proxy in front of repeated Jev classification calls. At first everything goes to Jev; from Jev's answers it trains a small head on frozen sentence embeddings, picks a confidence threshold with an exact finite-sample bound so that at most 2% of all requests get an answer Jev wouldn't have given, and then answers the confident share locally at ~15 ms on a CPU. A permanent 2% audit keeps checking; if agreement breaks, everything falls back to Jev and it retrains.
Known limits: agreement is not accuracy (if Jev is wrong, so is the local model); coverage tracks how consistent Jev itself is (22% on noisy tweet tasks, 80% on news); it speaks Jev's API only, an OpenAI-compatible front is on the roadmap. Since 0.4.0 the guarantee can also cover "would Jev have been unsure", which matters if your code routes low-confidence answers to review. Apache 2.0.
Why would it be against the rules to use the previous answers that the AI model gave you within your own project? That's like saying you can't use your Claude code convo history to answer questions within your codebase.
Not today, but Interesting idea. The main motivation was a drop-in for an existing Jev setup, so the only teacher right now is Jev and the audit measures agreement with Jev. A correction would have to become a second label source that overrides Jev's for that input.
The head is a multinomial logistic regression: one linear layer plus softmax on top of a frozen sentence-embedding model (bge-small by default, swappable). That head is the entire local model, the encoder is off the shelf and never changes.
reply