started · updated
Continuum AI reveals safety vulnerabilities in frontier-scale MoE models
Continuum AI, a US-based research institution, has released a study demonstrating vulnerabilities in the safety alignment of frontier-scale AI models. The research focuses on the 320-billion parameter Mixture-of-Experts (MoE) model, GLM-5.3-Flash.
The study shows that by manipulating internal model weights, researchers can significantly reduce a model's ability to refuse harmful requests—by as much as 89 percentage points—without degrading the model's overall performance. The findings suggest that safety alignment is more fragile than previously thought, particularly in MoE architectures where the effects are distributed across attention mechanisms, standard layers, and expert layers.
Separately, FlashLabs has updated the free tier of its OrcaRouter AI inference gateway in Japan. The service has replaced the previous free model with GLM-5.3-Flash, a multimodal model developed by Z.ai. Despite its 320-billion parameter scale, the model utilizes an efficient structure that only processes approximately 18 billion parameters per task, supporting long-context inputs and multimodal processing.
Entities
Continuum AI · FlashLabs · GLM-5.3-Flash · OrcaRouter · Z.ai