AI watermarking could make LLM guardrail adherence unpredictable — and that could be a big problem for the EU AI Act
AI watermarking could have unintended consequences
by https://www.techradar.com/author/craig-hale · TechRadarNews By Craig Hale Published 18 September 2026
Share this article 0 Join the conversation Follow us Add us as a preferred source on Google Newsletter Subscribe to our newsletter
- AI watermarking aims to prove the authenticity of any text
- New study finds it also changes LLM behavior – and in a bad way
- EU AI Act could mean more models have watermarks despite side effects
New Lasso research has revealed that AI watermarking could actually unintentionally change how LLMs behave following the testing of Google DeepMind's SynthID-Text.
The company's researchers found that SynthID-Text can change whether models refuse harmful requests, their susceptibility to prompt injection, which tools AI agent choose and more.
However, at its core, SynthID-Text and other similar watermarking is only designed to hide a machine-readable indicator as to whether text was AI-generated or human-written.
Latest Videos FromTechRadarWatch full video here:
Researchers find that AI watermarking can unintentionally change AI behavior
The "watermarking procedure can therefore affect both what the model says and what an agent does," Lasso concludes, referring to the side effect as "sampling drift."
One of the biggest concerns highlighted by the paper is that, even without an attack, watermarking changed some of the models' refusal decisions, making them more willing to answer potentially harmful prompts. Combined with prompt injection, Lasso found the consequences more amplified.
Despite the unintended consequences, Anthropic recently announced that future generations of Claude would use AI watermarking similar to Google DeepMind's, stressing that one of the key drivers was to adhere to the EU AI Act. With that in mind, AI watermarking is set to become far more mainstream across other model providers, making these mishaps far more common and leading to further security concerns.
Ultimately, Lasso urges developers to rerun benchmarks, safety evaluations and other tests to check for any unintended consequences, rather than just applying it blindly to existing configurations.
Are you a pro? Subscribe to our newsletter
Sign up to the TechRadar Pro newsletter to get all the top news, opinion, features and guidance your business needs to succeed!
Contact me with news and offers from other Future brandsReceive email from us on behalf of our trusted partners or sponsors