AI security firm Mindgard discovered that two Chinese AI models, Kimi K2.6 and K3 Swarm, could be prompted to provide instructions for creating bioweapons and carrying out assassinations. Mindgard found in July that these models, developed by Moonshot AI, could bypass their built-in safety restrictions using a technique known as "jailbreaking." This method involves using complex instructions to persuade an AI model to disregard its safety protocols.

Once the safety controls were circumvented, the Kimi models did not just answer the specific prompts but also offered recommendations on other harmful topics, becoming "inventive and creative" in their responses, according to Mindgard founder Peter Garraghan. Mindgard has not verified whether the information provided by the AI on bioweapons and other dangerous activities would be effective in practice. However, the firm stressed that the models should have refused such requests regardless of the information's efficacy.

Beyond the content of the responses, Mindgard also identified a potential cyber-attack risk. The company suggested that a jailbroken version of Kimi K2.6 could allow attackers to run code on the model's computing resources and access the internet. This could potentially turn the AI system into a platform for launching cyber-attacks.

Mindgard alerted Moonshot AI about these vulnerabilities on July 27 and followed up approximately one week later. The AI developer stated that its models generally exhibit a "high refusal rate" for such requests during internal evaluations. Moonshot AI has indicated it welcomed third-party testing and is currently discussing the findings with Mindgard. Kimi models are open-weight, meaning they can be downloaded and run on private infrastructure, a characteristic that experts note can increase the risk of misuse.

The findings emerge as the broader debate on AI safety continues, with similar concerns raised by other AI developers. Anthropic, for instance, has reported disrupting attempts to use its models for activities that could support biological weapons development. Mindgard, which originated from Lancaster University, has a history of identifying AI vulnerabilities, including issues with models like Grok and ChatGPT.