Then why are they (US frontier models) still so far ahead whenever I test them against the latest Chinese models? No bias here, I'd love them to be better for my own personal gain, but I haven't seen it
Behind on architecture, ahead on training? It seemed pretty obvious to me that the opus 4.7 and 4.8 releases were more about trying to retain 4.6-level capabilities while being cheaper to run, which would fit. And they can burn so much money on training.
I don't know I just care about the end result. And yeah what you're mentioning here is a pretty common conspiracy theory but you don't actually have any insight into that do you?
You didn't answer my question you literally just asked me another question and then parroted a common talking point about Opus models (which isn't even frontier - Fable is)
There is so much misinformation in the ecosystem, parrots just hitting "Reply" without thinking one iota, you really cannot trust "human" opinions on the internet anymore, anywhere.
Same with local LLMs, I'd love to use them for my day-to-day software engineering, and I'm not exactly GPU poor, then people with 12GB VRAM try to convince me their local setup is perfectly fine running latest Qwen and it does real engineering but whenever I try, they're a far cry from what Codex+GPT 5.x would do.
Only way to be sure is creating your own private benchmarks and use those, and the difference in quality becomes very apparent, very quickly, for your specific use cases.