Quite impressed by the energy people are putting into making OSS Jev-like models.
I understand the hype but I wonder: what are the use cases for this kind of model? Could it be used in the context of coding agents, or is it more relevant in totally different situations?
Consider every situation where you "force" an LLM to output only a choice / category, or a set of them. If you have workflows like that, you're now being promised significant cost- and latency reduction.
For coding agents it'd only be useful in a subset of situations. E.g. you could imagine using one to classify bash tool calls into safe and unsafe for example.
To develop a smart ai system for my 2d roguelike platformer?
game has way too many moving system for classic state-machine ai + i cant spare the time to develop it.
its low latency entices me.
for the tactical side kev and the decision-model route are the right call, a state machine there isnt where a local llm helps. where a small local model earns its slot in a roguelike is the text: dialogue, item flavor, npc barks, the stuff that blows up your content budget. running it on the players box means no api round-trip and no cost per line, which matters once a run spits out thousands of them. nobodywho wraps llama.cpp as a godot node for that, though if your moving parts are all combat ai and no text it wont buy you much. what engine are you on?
> what are the use cases for this kind of model? Could it be used in the context of coding agents
Yeah, it could. The most obvious usage would be to have local fast cheap "feedback" / "control" over a slower more expensive agent (i.e. cc / codex / opencode). Things like "goals" could now be split from a long prompt into "actions" and "verifiers". Where for each action you also produce a verifier. Then after each action you run the verifier w/ this kind of "universal classifier" and decide if the step was done correctly, if it needs follow-up and so on.
Example: implement auth in this repo -> llm_plan() -> for item in plan generate_verifier() -> for item in plan implement() ; verify() ; accept() / followup().
Verifiers could be something like this. take a plan item as input, generate classification questions that might verify the task "is this following project conventions?" | "is this touching files from other tasks?", etc.
You can do that with LLMs, but some things might become cheaper / faster. And you can pretty much use it to check against an ever growing list of conventions. Yours or project specific.
The git repo linked at top has some good examples for email classification for automatic email forwarding to specific departments, along with judging email tone and severity / priority, like for customer service emails.
Edit: The Flipper One is planning to have an LLM acceleration co-processor, and be able to host up to a 4GB VRAM size LLM. One use case they envision in their planning is using the microphone along with text to speech to be able to say, "Create an .ini file for this system with these specs" and the small LLM can do that on-device (its a handheld device) and then the user can use/send/upload that file.
Second Edit: I would love a mini LLM in KiCad or Altium that could take a component datasheet and produce a good footprint and schematic symbol for it.
The big problem is that Jev is only useful when you both:
- need fast response
- can tolerate Jev's mediocrity compared to real frontier models
(I explicitly ignore cost, because if you desperately need a cost-optimized classifier you would just build one)
Outside of fun demos these two rarely come together: if it's critical enough to require sub second speed, then it can't be mediocre.
The reason so many OSS models are being built is that Typesafe team made a ton bombastic claims about Jev being a huge breakthrough, and ML folks are realizing they can build Jev-like model in 2 days instead of 2 years.
Classifiers can be very useful for guiding the agentic loop which is basically a state machine. You have the agent propose a task, write some tests, write code, run tests, tests fail, write more code, go to acceptance, etc. So, a classifier can judge state transitions and decide what the agent should do next for example.
There is an enormous amount of time that goes into ticket and other triaging admin across industries. Its use cases go beyond that, but that alone would save a huge number of person-hours.
I understand the hype but I wonder: what are the use cases for this kind of model? Could it be used in the context of coding agents, or is it more relevant in totally different situations?