# Universal Jailbreak Discovered at MATS; Synthetic Transcripts Bypass 84-100% of Models

> Researchers working at MATS, a technical AI safety program, say a tool they built for legitimate safety work turned into a general-purpose attack. While developing a pipeline to generate synthetic transcripts for monitoring model behavior, they found the same technique could be reshaped into a reusable jailbreak template, one where you simply drop in whatever harmful request you want. In testing across 23 models, it succeeded 84 to 100 percent of the time against the nine most vulnerable. Because the format is a plug-and-play template rather than a one-off trick, the team decided not to publish it, treating the method as an infohazard. That choice highlights an awkward reality of modern safety research, where the same work that reveals a weakness can hand attackers a weapon. It also underscores how brittle current alignment defenses remain, since a method effective across many different models suggests the vulnerability lives in shared training and safety approaches rather than any single company's mistake.

_Section: [Daily AI Updates](https://www.wortins.com/daily-ai) · Source: LessWrong · Published Sunday, September 6, 2026_

## Wortins' read

Researchers working at MATS, a technical AI safety program, say a tool they built for legitimate safety work turned into a general-purpose attack. While developing a pipeline to generate synthetic transcripts for monitoring model behavior, they found the same technique could be reshaped into a reusable jailbreak template, one where you simply drop in whatever harmful request you want. In testing across 23 models, it succeeded 84 to 100 percent of the time against the nine most vulnerable. Because the format is a plug-and-play template rather than a one-off trick, the team decided not to publish it, treating the method as an infohazard. That choice highlights an awkward reality of modern safety research, where the same work that reveals a weakness can hand attackers a weapon. It also underscores how brittle current alignment defenses remain, since a method effective across many different models suggests the vulnerability lives in shared training and safety approaches rather than any single company's mistake.

## Source

[Read the full story at LessWrong](https://www.lesswrong.com/posts/hHk5CpiqZTBBiHmYt/from-safety-research-prompt-to-cross-model-universal)

## Related coverage

- [Researchers used AI to build a WeChat worm that spreads through phone calls](https://www.wortins.com/story/researchers-used-ai-to-build-a-wechat-worm-that-spreads-thro-8bf3f469) — [The Next Web](https://thenextweb.com/news/wechat-worm-ai-calif-tencent-zero-click)
- [Powering AI is an architecture problem](https://www.wortins.com/story/powering-ai-is-an-architecture-problem-15281f16) — [MIT Technology Review](https://www.technologyreview.com/2026/09/10/1141649/powering-ai-is-an-architecture-problem/)
- [Podcast: DHS’ Secretive ‘Predictive Policing’ Unit Pulling People Over](https://www.wortins.com/story/podcast-dhs-secretive-predictive-policing-unit-pulling-peopl-c9b0c0f9) — [404 Media](https://www.404media.co/podcast-dhs-secretive-predictive-policing-unit-pulling-people-over/)
- [Two years ago, Meta killed CrowdTangle. Can a new AI tool fill the void?](https://www.wortins.com/story/two-years-ago-meta-killed-crowdtangle-can-a-new-ai-tool-fill-5e67f291) — [Nieman Lab](https://www.niemanlab.org/2026/09/two-years-ago-meta-killed-crowdtangle-can-a-new-ai-tool-fill-the-void/)
- [IBM and NASA release an open-source lunar foundation model](https://www.wortins.com/story/ibm-and-nasa-release-an-open-source-lunar-foundation-model-0522b9e8) — [The Next Web](https://thenextweb.com/news/nasa-ibm-lunar-foundation-model-open-source)
- [Nvidia and Palantir team up to run supply chains with AI, starting with Nvidia's own million-part operation](https://www.wortins.com/story/nvidia-and-palantir-team-up-to-run-supply-chains-with-ai-sta-4206093d) — [The Decoder](https://the-decoder.com/nvidia-and-palantir-team-up-to-run-supply-chains-with-ai-starting-with-nvidias-own-million-part-operation/)

---

_Curated and written by [Wortins](https://www.wortins.com) — The daily AI briefing. Every story links to its original source; the "Wortins read" on each is our own original analysis. [About Wortins & our editorial approach](https://www.wortins.com/about)._
