Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

I feel like it will certainly need to be a hybrid approach for the moment. They’ll get a long way running small highly specific and tuned models locally on the neural engine for a lot of stuff. But, at some point, people are going to expect to ask questions and have something in the league of a Mistral/Claude/ChatGPT talk back and that’s just not possible with today’s hardware. Over time I expect more and more of that will get moved locally, and less will hit the “escape hatch” of throwing things to a giant LLM.

If they end up renting rather than building that capability, it says to me that they’re betting on being able to move most of that locally to the phone on a relatively short time horizon, which is interesting.



>Over time I expect more and more of that will get moved locally

I'm not so sure whether this is the direction of travel. The economics are pretty harsh for local, general purpose AI. And battery will always be a limiting factor.

If you're sending a few hundered requests per day to a very powerful AI then sharing the cost of the inference machinery with others has overwhelming economic advantages. And it will need access to current data anyway, so it can't be completely local.

There will be tasks such as keyboard autocomplete and a lot of other specialised tasks where latency or privacy is more important than quality. So yes I do believe in hybrid. But I think the cloud will always do the heavy lifting when it comes to more general tasks.

I can imagine an alternative scenario where a lot of AI processing happens on Mac and PC and mobile devices benefit from that. But many people don't have powerful desktop devices. So I'm not sure Apple can rely on this.




Consider applying for YC's Fall 2026 batch! Applications are open till July 27.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: