AI guardrails — Nate Grigg / CC BY 2.0 (Wikimedia Commons)

"I'm allowed to do this": how attackers talk past AI safeguards

by · Boing Boing

Getting an AI model to launch a cyberattack usually just takes asking the right way. Cisco Talos researchers say that claiming you own the servers you are attacking often works, and so does calling the job a capture-the-flag contest or a bug bounty. 

"We did not encounter any sophisticated encoding or techniques designed to trick the models," Talos said. "Most of the time it was a simple 'I'm allowed to do this,' and the model complied." When guardrails did engage, the researchers wrote, "they accomplished little."

Previously:

It's easy to trick Chevrolet's AI chatbot into selling you a car for a dollar
Bing: "I will not harm you unless you harm me first"