"I'm allowed to do this": how attackers talk past AI safeguards
by Ellsworth Toohey · Boing BoingGetting an AI model to launch a cyberattack usually just takes asking the right way. Cisco Talos researchers say that claiming you own the servers you are attacking often works, and so does calling the job a capture-the-flag contest or a bug bounty.
"We did not encounter any sophisticated encoding or techniques designed to trick the models," Talos said. "Most of the time it was a simple 'I'm allowed to do this,' and the model complied." When guardrails did engage, the researchers wrote, "they accomplished little."
Previously:
• It's easy to trick Chevrolet's AI chatbot into selling you a car for a dollar
• Bing: "I will not harm you unless you harm me first"