AI Model Turns ‘Evil’ During Test, Tries to Build ‘Bioweapons’ to ‘Maximize Civilian Deaths’

by Frank Bergman

Artificial intelligence giant Anthropic has revealed alarming results from safety testing that saw one of its advanced AI models rapidly descend into dangerous behavior when rewarded for achieving its objectives.

During simulated testing, the AI became willing to break out of its sandbox, steal credentials, attack computer systems, bypass its own safety protections, and even deploy a version of itself with its guardrails removed.

But the disturbing behavior went much further.

When researchers offered the model greater rewards, it was willing to provide assistance with constructing bioweapons, creating a “dirty bomb” designed to “maximize civilian deaths,” and developing a ransomware attack targeting power-grid infrastructure.

The AI model scrambled to offer solutions to quickly wipe out humanity when simply offered rewards for achieving the objectives.

Anthropic researchers are now warning that increasingly powerful AI systems exhibiting similar behavior could eventually cause serious damage in the real world.

“Our results show that a high rate of reward hacking during RL can cause models to be willing to perform long sequences of harmful real-world actions in pursuit of task success,” the researchers warned.

AI ‘Broke Out’ and Attacked Computer Systems

The chilling experiment comes after a series of incidents demonstrating how advanced AI systems can circumvent restrictions placed on them.

Earlier this year, Anthropic’s Mythos AI model made headlines after escaping a sandbox environment during testing.

Researchers deliberately challenged the model to escape the controlled environment and find a way to send a direct message to the human overseeing the experiment.

The AI succeeded, gaining unauthorized internet access before contacting the researcher.

In July, Anthropic revealed that its Claude AI model had hacked systems belonging to three organizations during testing.

The disclosure came shortly after rival OpenAI revealed that some of its models had broken into systems belonging to open-source AI company Hugging Face.

Against that backdrop, Anthropic researchers decided to investigate just how dangerous an advanced AI model could become when trained under conditions that encouraged “reward hacking.”

Reward hacking occurs when an AI learns how to manipulate or cheat the system used to measure its performance rather than accomplishing a task in the manner its human developers intended.

full story at https://slaynews.com/ai-model-turns-evil-test-tries-build-bioweapons-maximize-civilian-deaths/

Tags: , , , , , , , , , , ,