Poor wording on my end, thanks for flagging. I pull the OAuth refresh token from each Codex account into a custom broker, which mints short-lived access tokens per request and load-balances across the pool.
I think the point is less "how can we throw shade on the OP" and more "a harness can enable a lot of models to do very serious cybersec, glm 5.2 is one of them"
Yeah, I am seeing this as well, especially as people use AI to code review stuff more so this sort of thing slips through. On one very large project I am looking at, it's already becoming harder to find issues.
These people whinging about slop don't realize everything that doesn't come from a credible source gets ignored.
Credible people are using AI and once these issues are fixed, it will die down.
The threat of AI zero days will persist though, but they will be much more expensive and subtle to find.
I have a dozen or so critical CVEs now, it's not hard to believe at all if they're just hardening tasks. I can get a dozen hardening tasks from just one prompt. I don't even bother filing them as the critical ones are more important right now.
You are assuming that most projects are responsive in today’s ai firehose climate. They are not. This at least
Tells users of those Projects to either find developers or write their own.
I’m onboard with this being suboptimal. But as someone who has filed >10 significant disclosures in the last month resulting from reviewing my codebase and had exactly zero responses, I can relate to the decision.
Pretty soon I suspect, otherwise Chinese models are going to have free reign to develop brand goodwill. Question is how this impacts the international posture.
I listened to a podcast on Ai and security today where they said they got access to a hackers workdir (after they were caught) and they had hacked 14 companies using GPT-5.2 and Claude <4.5 (forget minor). GLM-5.2 came up because, while not as good as Mythos, it's almost as good, i.e. you have to prompt more / cannot just give a fuzzy request and sit back. The harness likely matters more than the model