Hacker Newsnew | past | comments | ask | show | jobs | submit | devyy's commentslogin

In my case, it was pretty fast i would say, using S24 Fe, on Gemma3n E2B int 4, it took around 20 seconds to answer "Describe this image". And the result was pretty amazing.

Stats -

CPU -

first token - 4.52 sec

prefill speed - 57.50 sec tokens/s

decode speed - 10.59 tokens/s

Latency - 20.66 sec

GPU -

first token - 1.92 sec

prefill speed - 135.35 sec tokens/s

decode speed - 11.92 tokens/s

Latency - 9.98 sec


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: