None of the other items Intel has let go of are an existential must have. GPUs are a must have in any computer big or small. Raw number crunching is core to the very profitable HPC market. Players like Nvidia have fully custom interlinks that threaten to dry up even the CPU sales Intel still gets in this market. GPU is a focus on mobile where expectations have been much raised.
This group is supposedly 6 years old. But reciprocally, the other players have been pouring money in over decades. GPUs are fantastically complicated systems. Getting started here is enormously challenging, with vast demands. Just shipping is a huge accomplishment. Time to grow into it & adjust is necessary. It's such a huge challenge, and I really hope AXG is given the time, resources, iterations, & access to fabs it'll take to get up to speed.
>None of the other items Intel has let go of are an existential must have.
Except mobile ARM chips and mobile LTE modems, both of which Intel sold off, and those are some of the most desirable things to make right now. Just ask Quallcomm.
Yep. Competent mobile ARM and LTE modem cores would be very nice to have in their tech stack right about now. I think a credible (if not fully competent), GPU stack is pretty essential for strategic relevance going forward. This seems like something they have to make work.
Intel should double down on its core business which is selling x86 chips for Windows PCs. If it loses to AMD+TSMC here, there's nowhere left to hide. IFS is very long term and a new business opportunity in a market where Intel is not the leader. x86 for Windows is their bread and butter. They need to throw everything else out and downsize to just excelling at what they've always done. My guess is AXG is a separate division than the division that makes the graphics for their core microprocessors.
> Intel should double down on its core business which is selling x86 chips for Windows PCs. If it loses to AMD+TSMC here, there's nowhere left to hide. IFS is very long term and a new business opportunity in a market where Intel is not the leader.
The very second someone else but Apple brings a competitive ARM desktop CPU to the market, it's game over for Intel. x86_64 literally cannot compete with modern ARM designs simply because of how much utter garbage from about thirty years worth of history it has accumulated and absolutely needs to support in the future because even the boot process still requires all that crap, whereas ARM was never shy about cutting out stuff and breaking backwards compatibility to stay performant.
The only luck that Intel has at the moment is that Samsung and Qualcomm are dumpster fires - Samsung has enough problems getting a phone to run at a halfway decent performance with their Exynos line and Qualcomm managed to completely botch their exclusive deal with Microsoft [1] (hardly surprising to anyone who has ever had the misfortune to have to work with their crap). A small startup that is not bound by ages of legacy and corporate red tape should be able to complete such a project - Annapurna Labs have proven it's possible to break into the server ARM CPU market well enough to get acquired by Amazon.
The fact you can run old software on any PC still is a pretty tremendous feature. Apple kicked of 32 bit apps from Mac and iPhones, its user accept that, business PC users might not. I have a number of IOS games I enjoyed that are no longer usable (yeah I could have stopped upgrading my i device, but the security problems make that not a viable option)
Conventional wisdom in the mid 90s was powerpc (RISC) would eventually be better the x86, but it never happened. They worked around the issues. And Microsoft eventually made an OS that wasn't a crash fest (I'm looking at you windows ME)
Also no one can afford to be on TSMCs best node when Apple buys all the production. Apple's been exclusive on the best node for at least a couple years now. Even AMD isn't using TSMCs best node yet.
> Apple kicked of 32 bit apps from Mac and iPhones, its user accept that, business PC users might not.
Yeah, but that was (at least on the Mac) not a technical requirement, they just didn't want to carry around the kernel-side support any more. IIRC it didn't take long until WINE/Crossover figured out a workaround to run old 32-bit Windows apps on modern Macs.
> Also no one can afford to be on TSMCs best node when Apple buys all the production. Apple's been exclusive on the best node for at least a couple years now. Even AMD isn't using TSMCs best node yet.
Samsung has their own competitive fab process, but they still have yield issues [1]. It's not like TSMC has a monopoly by default.
Apple will always be niche because they are too expensive.
ARM based PC client processors accounted for about 9% of the total market.
But I don't expect that number will grow a lot higher due to cost of Macs.
Also, you can't bring out a desktop PC based on Linux because that a very slowly growing market segment and MacOS is proprietary. Don't think Apple will start another PowerComputing type scenario where they open their platform to others.
It's not a matter of performance for this reason (price).
So, on Windows, assuming Macs don't take away any more market share from Windows,
they need to beat AMD+TSMC.
> So, on Windows, assuming Macs don't take away any more market share from Windows, they need to beat AMD+TSMC.
No. All it needs is
- Microsoft and Qualcomm breaking their unholy and IMHO questionably legal alliance
- an ARM CPU vendor willing to do the same as Apple did and add support for accelerating translation of x86 code (IIRC, memory access models/barriers are done differently between x86 and ARM, and Apple simply extended their cores to be able to use the same memory access/barrier model as x86 on translated-x86 threads)
- an ARM CPU vendor willing to implement basic functionality like PCIe actually according to spec - even the Raspberry Pi which is the closest you can get to a mass market general-purpose ARM computer has that broken [1]
- someone (tm) willing to define a common standard of bootup sequence/standard feature set. Might be possible that UEFI fills the role; the current ARM bootloaders are a hot mess compared to the old and tried BIOS/boot sector x86 approach, and most (!) ARM CPUs/BSPs aren't exactly built with "the hardware attached to the chips may change at will" in mind.
Rosetta isn't patented to my knowledge, absolutely nothing is stopping Microsoft from doing the same as part of Windows.
If you translate, you do face performance issues.
Apple or some other vendor cannot make an ARM chip so fast at a competitive cost
that beats an x86 chip in emulation mode.
The underlying acceleration techniques for both ARM and x86 are the same.
> Apple or some other vendor cannot make an ARM chip so fast at a competitive cost that beats an x86 chip in emulation mode.
This reminds me of Iron Man 1... "Tony Stark was able to build this in a cave! With a box of scraps! - Well, I'm sorry. I'm not Tony Stark."
Apple has managed to pull it off so well that the M1 blasted an i9 to pieces [1]. The M1 is just so damn well more performant than an Intel i9 that the 20% performance loss compared to native code didn't matter.
And for 99.9999% of users the M1 performance benchmarks vs. i9 _don't matter one bit._
The use case for the vast majority of laptops include I/O- and memory-bound applications. Very few CPU-bound applications are run on consumer laptops, or even corporate laptops, for the most part. CPU-bound applications should be getting run on ARM or GPU clusters in the cloud.
The use case for an M1 in laptops is the power benchmarks vs. an i9.
> The use case for the vast majority of laptops include I/O- and memory-bound applications.
Where the M1 just blows anything desktop-Intel out of the water, partially because they integrate a lot of stuff directly on the SoC, partially because they place stuff like RAM or persistent storage extremely close to the SoC whereas on desktop-Intel RAM, storage and peripheral controllers are all dedicated chips.
The downside is obviously that you can't get more than 16GB RAM with an M1 and 24GB RAM with the new M2's and you cannot upgrade either memory at all without a high-risk soldering job [1]... but given that Apple has the persistent storage so closely attached to the SoC to swap around, it doesn't matter all that much.
No.
The architecture is the biggest factor. Damn, even Jim Keller talks about how most programs use a very small subset of instructions.
It isn't like RISC makes miracles, but sure helps them when your power budget is small.
The "x86 translation support" bits are part of ARM64 ISA, they just weren't standardised in time - apple effectively implemented a WIP version of it AFAIK, though main changes are in interface for OS.
There is just one serious problem with the x86 architecture and that is the difficulty of decoding instructions which could be anywhere between 1 and 14 bytes. It's not hard to design a quick instruction decoder but hard to design a quick instruction decoder that is power efficient.
... Oh yeah, and there are the junkware features like SGX and TSX and also the long legacy of pernicious segmentation that means Intel is always playing with a hand tied behind its back, for instance, the new laptop chips that should support AVX512 but don't because they just had to add additional low performance cores.
I may be showing ignorance of some important use case here, but I'd put AVX512 in the "junkware feature" category. It can run AI/HPC code faster than the main processor, but still so much slower than a GPU as to be pointless; when it does run, it lowers the frequency and grinds the rest of the processor to a halt until it eventually overheats anyway. Maybe, maybe it has a place in their big desktop cores, but why would anyone want it in a laptop? You can't even use it in a modern ultrabook-style casing without a thermal shutdown and probably first-degree burns, without even mentioning the battery drain.
If you have all the cores running hard the power consumption goes up a lot but if it is just one core it won't go up too much. If worse come to worse you can throttle the clock.
In general it's a big problem with SIMD instructions that they aren't compatible across generations of microprocessors. It's not a problem for a company like Facebook that buys 20,000 of the same server but normal firms avoid using SIMD entirely or they use SIMD that is many years out of date. You see strange things like Safari not supporting WebP images on a 2013 Mac while Firefox supports them just fine because Apple wants to use SIMD acceleration and they'd be happier if you replaced you 2013 Mac with a new one.
I worked on a semantic search engine that used an autoencoder neural network that was made just before GPU neural networks hit it big and we wrote the core of our implementation in assembly language using one particular version of AVX. We had to do all the derivatives by hand and code them up in assembly language.
By the time the product shipped we bought new servers that supported a new version of AVX that might have run twice as fast but we had no intention of rewriting that code and testing it.
Most organizations don't want to go through the hassle of keeping up with the latest SIMD flavor of the month so a lot of performance is just left on the table. Intel is happy because their marketing materials can tell you how awesome the processor is but people in real life don't experience that performance.
Slowdown/overheating are not inherit to AVX512, they are specific to older implementations. And no the GPU is not a replacement, the latency difference between GPU and AVX512 is off the charts.
Okay, there's a use case I didn't know about. Do you know any SIMD projects that require the lower latency?
I'm glad to hear some newer AVX512 implementations are better. I haven't used an Intel one for years, but I have used the current-gen Ryzen version. Unfortunately it's still suffering the same problems-last time I stress-tested it my CPU hit 100C and some absurd power number before I was too scared to continue. Granted, that was with Prime95, which I believe is close to peak utilization.
No one in their right mind will touch anything like Loongson or whatever else the Chinese government puts out. Maybe the Chinese military, but that's it.
This group is supposedly 6 years old. But reciprocally, the other players have been pouring money in over decades. GPUs are fantastically complicated systems. Getting started here is enormously challenging, with vast demands. Just shipping is a huge accomplishment. Time to grow into it & adjust is necessary. It's such a huge challenge, and I really hope AXG is given the time, resources, iterations, & access to fabs it'll take to get up to speed.