Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Well I tested it on 7900XTX with the same prompts and their draft model gave me about 30t/s. Their own snippet of code with regular MTP model gave me 60t/s.

Also model with their draft answered incorrectly. With MTP it answered correctly.

Question was: "Does MikroTik CRS312-4C+8XG-RM have combo ports?". The answer is Yes.

 help



That's not how you're supposed to use LLMs. You shouldn't expect a tiny little local model to know random facts about every obscure consumer product on earth. That's the job of tool calling. At best a model of this size is just giving you a random guess.

You're basically saying "I tried rolling these dice one time, the green dice rolled a 6 and the blue dice rolled a 1, so green dice are better"


agreed. a niche knowledge callout is about the worst benchmark one can give a smaller model.

smaller models are attempting to distill the useful methodologies, not the license plate number of an obscure extras car on Magnum PI.

that said I wonder if there is a small 'trivia' model out there. Seems like the kinda thing Google would tackle.


Which was not he point because I was testing their solution for MPT and it was just funny addition. But of course in internet you always will find some 'well akchually' person straight from the meme.

When I changed the number of draft tokens to 3 in both, it helped and they Draft is actually performing a bit better:

- draft: 67.17

- MTP: 64.18

Why they used those examples? Seems strange.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: