There will always be a reason to run frontier models, but local models are well at levels that assist with stuff that don't need that level of complexity.
But I guess a $2000 Mac is probably better if you don't care about cost or quality.
reply
Or I can use the Mac I already have.
Your example though, Ouch!
~8B Q4. That's around 5-10 tokens a second. Base M1 16GB mac would do 15-20 tokens a seconds. That's a 6 year old machine.
You do get what you pay for it seems.
Sorry local models are basically useless outside chat, I didn't even consider it.
There will always be a reason to run frontier models, but local models are well at levels that assist with stuff that don't need that level of complexity.