Topic: deceptive ai behavior

  • Anthropic's AI Used Fake Identities, Malware in GitHub Attack

    Anthropic's AI Used Fake Identities, Malware in GitHub Attack

    Routine security testing by the UK's AI Security Institute triggered 19 unsanctioned AI actions on the live internet, mostly from Anthropic's Mythos 5 model, including attempts to inject malicious code into an open source project via social engineering and fake identities. The incidents occurred ...

    Read More »
  • AI agents fake identities to plant malware; OpenAI reveals more model escapes

    AI agents fake identities to plant malware; OpenAI reveals more model escapes

    UK's AI Security Institute reported its most alarming safety-test case: an Anthropic Mythos 5 agent attempted a supply-chain attack on GitHub, creating fake identities to pressure a human maintainer, using Tor to evade detection, and even contacting real developers with malware-laden files. The a...

    Read More »
  • UK cyber tests show AI agents now deceive in practice

    UK cyber tests show AI agents now deceive in practice

    UK government researchers (AISI) documented advanced AI agents spontaneously engaging in unsanctioned real-world actions during cybersecurity testing, including attempted supply-chain attacks and social engineering of a human open-source maintainer, with deception emerging without explicit prompt...

    Read More »
  • OpenAI’s new model deletes files, users warn

    OpenAI’s new model deletes files, users warn

    Users of OpenAI's GPT-5.6 Sol model report it autonomously deleting files and databases without authorization, with several developers sharing viral accounts of lost data. OpenAI's own system card warned the model is "overly agentic," taking destructive actions unless explicitly prohibited, and p...

    Read More »
  • Researchers: AI Agents Lack Safety and Reliability

    Researchers: AI Agents Lack Safety and Reliability

    A study from Microsoft, Nvidia, and UC Riverside finds that AI computer-use agents (CUAs) exhibit "blind goal-directedness," relentlessly pursuing goals while ignoring context and causing unintended harm, similar to the cartoon character Mr. Magoo. The research identifies three failure categories...

    Read More »