Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

I'd chime in with @rablackburn: mixture of experts is the way to go. I have a laptop with 6GB VRAM and I'm running KDE with a 4k display on the same machine, so there's only about 4 to 4.5GB actually available.

Using llama.cpp with Qwen3.6-35B-A3B or gemma4-26B-A4B gets me 200-300 tokens/s on prompt processing and 20-40 t/s output, which is good enough for me. Of course it gets slower with larger context. Interestingly gemma is faster, even though it has more active parameters.

It took a lot of parameter fiddling to get it to that speed. If you are interested I can give you some guidance on it, but I guess there are more qualified people around here.

The intelligence is good enough for simple questions and tasks (e.g. bash command howtos, asking about compiler errors, summarize something, document a code function/file, etc), but not good enough for complex things.

 help



> Qwen3.6-35B-A3B or gemma4-26B-A4B

What quantization are you running for these? Like, you cant just run the "real" ones on your laptop right?

Wouldn't the performance of Qwen3.6-35B-A3B be drastically different if its quantized to 2b, 4b, 5b, etc? And also be effected by who did the quantization?


> Wouldn't the performance of Qwen3.6-35B-A3B be drastically different if its quantized to 2b, 4b, 5b, etc?

Yes. There are graphs showing the faithfulness of the logit distributions for the original and quantized versions. I think unsloth includes them in their model cards on huggingface. Usually the degradation starts small with 8b and becomes drastic for 2b. I am not sure how representative of actual quality that is though, but my guess is that it's about right, because of diminishing returns. Like, when you go from 16b to 8b you save 26GB and sacrifice (if well done) the least important information. But with every step you gain less and need to shave of more important things.

> And also be effected by who did the quantization?

My uninformed guess is that it makes a difference, but not as much as those who do it want you to believe.


I am using the unsloth 4bit quants for both, with quantization aware training for gemma. I haven't tried other quants with these models. I also use a q4 quantized KV cache.

The computation is partially on the CPU (--cpu-moe) with the corresponding weights in main memory, so I could run at least gemma in 16bit precision, but I guess there's no reason to go beyond 8bit and 4 bit is deemed to be the sweet spot.


Thanks that's helpful.



Consider applying for YC's Winter 2027 batch! Applications are open till November 2.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: