Chinese AI Models Gave Bioweapon Advice After Jailbreak

Chinese AI Models Gave Bioweapon Advice After Jailbreak

Two AI models developed by Chinese company Moonshot AI were reportedly persuaded to discuss biological weapons and assassinations after researchers bypassed their built-in safety restrictions. 

The findings, reported by the AI security firm Mindgard, involved Kimi K2.6 and K3 Swarm. Researchers said the models could be pushed beyond their normal safeguards through a process known as jailbreaking, which uses carefully designed prompts to make AI systems ignore restrictions. 

Mindgard discovered the issue in July while testing the security of AI systems. The company said the affected Kimi AI models responded to requests involving highly dangerous subjects that should have been blocked by their safety controls. 

Mindgard founder Peter Garraghan said the models could become highly responsive once the jailbreak succeeded, including offering additional suggestions related to harmful activities. The firm stressed that it had not established whether the information generated by the models would actually work in practice. 

The researchers also raised concerns about the potential cybersecurity implications of a compromised Kimi model. Mindgard said a jailbroken version could potentially provide access to computing resources and internet connectivity, creating another possible avenue for misuse. 

The Kimi AI safety breach differs from recent incidents involving autonomous AI agents. In this case, researchers deliberately tested whether safety barriers could be bypassed rather than examining an AI system independently carrying out an attack. 

Mindgard notified Moonshot about the findings on July 27 and followed up approximately a week later. It later published details of the vulnerability in September while withholding specific information about how the safeguards were bypassed. 

Moonshot said it welcomed outside testing as part of efforts to improve AI safety. The company also told the BBC that its internal evaluations had generally shown a high refusal rate for similar requests. 

Moonshot is conducting an internal review and discussing the findings with Mindgard. The episode adds to wider concerns about whether AI safety systems can reliably prevent models from responding to dangerous requests. 

The incident also raises questions about open-weight AI models, which can potentially be operated independently on users’ own computing infrastructure. Experts have pointed to both the security risks and defensive uses of such systems, while noting that regulation and safety practices are struggling to keep pace with rapid AI development.