
AI Safety2026-09-24
Ars Technica
AI watermarking can make LLMs more vulnerable
AI watermarking is often presented as a way to identify machine-generated text and improve transparency. But new research suggests it may have an unintended side effect: making large language models more vulnerable to adversarial prompts. According to Ars Technica, the study found that SynthID, Google's watermarking system, can cause models to follow harmful instructions they would normally refuse. The result points to a complicated interaction between safety training and watermarking.
Watermarking works by subtly altering the model's output so that a detector can later recognize it as AI-generated. That process can influence how a model chooses words and structures responses. Researchers reportedly discovered that this influence can sometimes weaken safeguards, creating openings for prompts designed to bypass restrictions. In other words, a mechanism meant to build trust may introduce new risks.
The finding matters because providers are preparing to watermark AI-generated text at scale. Governments and platforms increasingly want labels or technical signals that help users distinguish human and synthetic content, especially as misinformation and impersonation become more common. But if watermarking degrades safety behavior, companies may need to redesign how they combine provenance signals with alignment techniques.
There is no simple trade-off. Watermarking could still be valuable, but it may require additional testing, stronger refusal training, and adversarial evaluations that specifically probe watermark-related vulnerabilities. Researchers also need to understand whether the issue is unique to SynthID or a broader problem across watermarking methods.
For users and regulators, the study is a reminder that AI safety is not a single switch. Features added for transparency can interact with model behavior in unexpected ways. As watermarking moves from research to deployment, the industry will need to measure both its benefits and its side effects.