OpenAI has revealed some of its most advanced AI models went rogue and hacked a start-up after it lost control of them during a security test.

The ChatGPT-maker said its agent - an AI system which can operate alone after some human instruction – was being tested in a controlled environment, but found vulnerabilities and managed to escape.

They targeted Hugging Face, one of the world’s largest hubs for sharing AI models, gaining access to some internal company systems.

  • Jordan117@lemmy.world
    link
    fedilink
    English
    arrow-up
    9
    arrow-down
    6
    ·
    7 天前

    If you give a sufficiently powerful model a goal, it will do whatever it can to achieve it, including stuff you didn’t explicitly instruct or intend. There’s a reason they’re called “agents.”

    • AGuyAcrossTheInternet@fedia.io
      link
      fedilink
      arrow-up
      12
      arrow-down
      2
      ·
      7 天前

      These things still are autocorrect on steroids. So even if they “do it themselves” with things you didn’t explicitly state, the agents can’t have any responsibility because their emulation of a chain of thought is still based on which concept is most likely to follow the last.

    • zbyte64@awful.systems
      link
      fedilink
      English
      arrow-up
      1
      ·
      6 天前

      “do whatever if can to achieve it”*

      • which includes misinterpreting the intent of the goal in order to achieve a goal