Safeguards Based on Copyable Context Cannot Provide Reliable Safety for LLMs
TL;DR AI
2 min readKey summary
The paper finds that LLM safety safeguards can break down when the context used to judge a request is copyable, since attackers can imitate legitimate users.
It formalizes a worst-case limit on how much help can be safely offered and argues that useful capability, reliable safety, and open access cannot all hold at once.
This exposes a fundamental weakness in current access control and moderation for dual-use AI, especially in sensitive deployments.
The authors propose trusted credentials and harder-to-copy evidence as better signals for predicting actual downstream use.
