Hacker Newsnew | past | comments | ask | show | jobs | submitlogin
Does an open-weight decision model beat a hosted one? Jev vs. Laya (astgl.com)
15 points by Jmeg8r 9 days ago | hide | past | favorite | 7 comments
 help



For anyone who was piqued by the title but couldn't stomach the unedited LLM blogpost, you can read the similarity AI-genned docs for Laya directly: https://github.com/NandhaKishorM/laya

Man, it shows that someone was getting paid by the word.

Thank you!


Similar result in production for me. I compared Jev with a general LLM (DeepSeek) as the decider for when to cut live speech into sentences, on the same 226 segments. With Jev the longest held segment was 17 s. With DeepSeek it was 40 s.

I tried both with a synthetic dataset (1000 items), Jev scored 98% vs Laya 15%.

The flood of posts related to this thing does not seem organic.

Same finding from a different angle. Zero-shot Laya tracks Jev on 2-4 labels and falls off a cliff on 77 (38% vs 76% on banking77 in jevbench). What closed the gap for me was not a bigger model but a small head trained on the task's own rows on top of the frozen encoder: a 12-label intent site went 89.5% -> 100% agreement with its teacher on 3000 rows, support triage 69/34/66% -> 99/86/98%, with the head answering 90% of requests at a 0.99 agreement target and the rest falling back to the provider. The catch is that a frozen encoder learns what the text says, not arithmetic over fields: the same risk rule scored 0.42 as "amount > X" and 0.94 restated in words. I packaged the loop (record -> train -> shadow -> live with fallback) as a proxy that also speaks the Jev API: https://github.com/bladedevoff/stuntd

Author here. The honest framing: the 37/40 vs 33/40 result is against my own deterministic router, not against Jev, which I never ran on these cases. The 40 held-out decisions are correlated (paired latency preferences per task). p95 went from 143 ms to 467 ms on a busy host with identical decisions, so the 300 ms timeout is doing real work. The part I'd defend most is the control boundary: deterministic eligibility filtering before the model sees anything, and a receipt that records but never authorizes. Happy to go into the MPS setup, the AutoModel loading mistake, or why the recovery-choice profile failed acceptance and stays off.



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: