Or I can use the Mac I already have.
Your example though, Ouch!
~8B Q4. That's around 5-10 tokens a second. Base M1 16GB mac would do 15-20 tokens a seconds. That's a 6 year old machine.
You do get what you pay for it seems.
Sorry local models are basically useless outside chat, I didn't even consider it.
Or I can use the Mac I already have.
Your example though, Ouch!
~8B Q4. That's around 5-10 tokens a second. Base M1 16GB mac would do 15-20 tokens a seconds. That's a 6 year old machine.
You do get what you pay for it seems.