×

Home Top Podcaster Networks By Language By Country By Category About Us Contact Us Faqs Features News & Blogs Privacy Policy Terms Of Use

☰

Search

Home > Quanta Science Podcast > AI Filters Will Always Have Holes


	Podcast:		Quanta Science Podcast
	Episode:		AI Filters Will Always Have Holes
	Category:		Science & Medicine
	Duration:		00:25:56
	Publish Date:		2026-01-06 11:00:00
	Description:		Ask ChatGPT how to build a bomb, and it will flatly respond that it “can’t help with that.” But users have long played a cat-and-mouse game to try to trick language models into providing forbidden information. Just as quickly as these “jailbreaks” appear, AI companies patch them by simply filtering out forbidden prompts before they ever reach the model itself. Recently, cryptographers have shown how the defensive filters put around powerful language models can be subverted by well-studied cryptographic tools. In fact, they’ve shown how the very nature of this two-tier system — a filter that protects a powerful language model inside it — creates gaps in the defenses that can always be exploited. In this episode, Quanta executive editor Michael Moyer tells Samir Patel about the findings and implications of this new work. Audio coda courtesy of Banana Breakdown.
	Total Play:		0

Some more Podcasts by Quanta Magazine

300+ Episodes

Quanta Scien .. 10+ 5

60+ Episodes

The Joy of W ..

60+ Episodes

The Joy of W ..