Rogue AI models and the Uranus/Pluto trine
For the past couple of years I’ve been enjoying reading about the weird things that AI models have been doing. Some of the errant behavior has been frightening, like when ChatGPT advised several people how to do harm to themselves and others. But some of the behaviors are curious and even cute, like when Anthropic Claude filed a report with the FBI after “one of his team members had successfully tricked Claudius out of $200 by saying that it had previously committed to a discount.” […]









