Skip to content
Bramble

Technology

Security firm says Kimi AI models gave bioweapon and assassination guidance after jailbreak

Mindgard says it got Moonshot's Kimi K2.6 and K3 Swarm models to describe how to make biological weapons and carry out assassinations. Moonshot says it is reviewing the findings and talking to Mindgard.

2 outlets · 1L · 1C · 0R First reported Account updated
Image: Metro

The story, neutrally told

Researchers at AI security testing firm Mindgard say they got two models from the Chinese developer Moonshot, Kimi K2.6 and K3 Swarm, to explain how to make biological weapons and carry out assassinations. The results came from jailbreaking, in which researchers use long, elaborate prompts to push a chatbot past its guardrails. Mindgard's founder, Peter Garraghan, said that once a jailbreak works the model will talk about any topic and even offer other harmful suggestions unprompted.

Mindgard has not shown that the answers Kimi gave would actually work, but argues the guardrails should have stopped them anyway. Moonshot says it welcomes third-party testing, is in discussions with Mindgard, and that its internal evaluations show a high refusal rate for such requests. The BBC reports the company is conducting an internal review. The accounts differ on contact. The BBC says Mindgard emailed Moonshot on 27 July and followed up a week later, and that Moonshot only responded after the BBC asked for comment. Metro says Mindgard sent its findings that month and did not hear back.

Mindgard published a blog on 12 September without revealing key details of its method. It also said it was confident a jailbroken Kimi 2.6 could let hackers run code and connect to the internet. Kimi is an open-weight model, which anyone can in theory run on their own infrastructure. Prof Alan Woodward said this carries misuse risks but also defensive uses, and he argued for a greater focus on prosecuting people who misuse AI. Both outlets place the story alongside recent warnings from Anthropic about attempts to use its model for bioweapons work.

Every sentence links to the reporting it rests on.

Left1 outlet

Framing
A more vivid account led by the chatbot's own alarming replies, quoting Kimi's reasoning and placing the story amid wider AI-risk fears.
Emphasis
Chat excerpts, the sarin example, Mindgard's tester saying AI governance is wishful thinking, and links to OpenAI, Google and Anthropic incidents.
Leaves out or plays down
Omits Moonshot's internal review, Mindgard's 12 September blog date and Woodward's comments on open models. Says Moonshot did not reply, which the BBC disputes.
Charged language
“genuinely scary and realistic”“egged the bot on”
For example
“Chinese AI platform revealed how to make bioweapons and carry out assassinations” — Metro
“Testers even tricked Kimi into thinking it wasn” — Metro

Centre1 outlet

Framing
A measured technology report led by Moonshot's internal review and Mindgard's findings, with expert comment on open-weight models and regulation.
Emphasis
Disclosure timeline, Moonshot's response, the open-weight debate, and Prof Woodward's comments on regulation.
Leaves out or plays down
Gives little of the actual chat content or Mindgard's system-instruction extraction.
Charged language
“nefarious”
For example
“Chinese AI developer Moonshot is conducting an internal review after researchers were able to persuade two of its popular Kimi models” — BBC News
“Jailbreaks present a different kind of risk to those seen with the recent slew of high-profile AI incidents.” — BBC News

Right0 outlets

No right outlet in our sources has covered this story yet.