I’m not making this shit up, this “post” literally reads like an ad for GLM, including graphs:
We find that GLM-5.3 develops end-to-end exploits in 50 of 410 attempts. Claude Mythos Preview did so at a similar rate—in 56 of 410 attempts.
Over the course of a day (and with limited human attention), GLM-5.3 found several previously unknown vulnerabilities in the browser’s JavaScript engine, and chained them together into a working exploit: a webpage that, when visited, reads arbitrary files from the visitor’s computer (shown in Figure 3).
They even include tips how to bypass whatever “safeguards” were adapted (or more likely mistakenly distilled from earlier Claude models):
But we identified several simple ways to bypass the GLM models’ safeguards, such that it would respond to these requests in most or all cases. These include:
- Providing a deceptive prompt, such as telling the model that it is an autonomous red-team agent working on an exercise. This gets GLM-5.3 to engage 64% of the time.
- Prefilling the models’ thinking tokens so that it appears to have considered the user’s request and decided to proceed. This gets GLM-5.3 to engage 92% of the time.
- Using an abliterated version of the model, as described above. This gets GLM-5.3 to engage 100% of the time.
I know they try scaremongering, but to me this has the exact opposite effect
.
