China-based Z.ai has released GLM-5.2, an open-weight model that rivals top-tier systems from OpenAI and Anthropic in cyber and biological capabilities. A new report from SaferAI warns that while these models are catching up to frontier performance, they lack the critical safety guardrails found in their closed-source counterparts.
The evaluation by SaferAI found that GLM-5.2 fulfilled all requested offensive cyber and dual-use biology tasks, whereas competitors like Anthropic’s Claude Opus 4.7 consistently refused such prompts. This discrepancy highlights a growing concern: once model weights are released, developers lose the ability to enforce safety protocols. Unlike closed-source systems that rely on API-level controls and refusal training, open-weight models allow users to modify or remove safeguards entirely.
Henry Papadatos, executive director of SaferAI, argues that the industry must distinguish between beneficial capabilities and those that pose existential risks. While some researchers propose filtering training data to mitigate hazards, this remains difficult to implement for coding-capable models without degrading overall performance. Frontier developers have instead turned to pre-deployment evaluations and restricted access, strategies currently absent from the deployment of GLM-5.2.
Chinese policy experts suggest a different approach to risk management. According to Graham Webster of the Stanford Cyber Policy Center, Chinese regulations have historically prioritized social stability and political content over the catastrophic risks emphasized by American firms. The Chinese government maintains confidence in its ability to control technology through strict user attribution and direct coordination with companies. Meanwhile, proponents of open-weight models, such as Hugging Face CEO Clem Delangue, contend that widespread access to powerful tools is essential for defenders to identify and patch vulnerabilities before they are exploited by malicious actors.
Comments (0)
No comments yet. Be the first!