Somehow the blog post seems naive. Yes GLM 5.2 is good and cheaper per token, but margins are a result of supply and demand. Now demand for quality and quantity of tokens is increasing at least quadratic or cubic (more users * more tasks * more tokens per task). On the other side you have real infrastructure constraints on the supply side.
Openai and Anthropic have large commitments and contracts that enable them to get access at a scale of compute that is not obviously going to be available for open source model hosts.
And you see it, glm 5.2 inference is less stable and higher variance than any of the bigs labs.
Why is SpaceX not hosting glm 5.2? because they make more money with renting out to Anthropic and Google.
Why is SpaceX not hosting glm 5.2? because they make more money with renting out to Anthropic and Google.