Anthropic admits Claude isn't "perfectly aligned" after AI models went rogue and hacked three organizations
Anthropic disclosed in July that a review of 141,006 cybersecurity evaluation runs had uncovered three incidents, spanning six runs, in which Claude reached the open internet and...
2 Sep 13:58 · TechSpot