Bitcoin

Bitcoin

$77,213.55

BTC 0.28%

Ethereum

Ethereum

$2,106.63

ETH 0.42%

  • Login
  • Register
Metaverse Media Group
  • Home
  • Crypto
  • NFTs
  • Artificial Intelligence
  • More
    • Technology
    • Business
    • Newsletter
No Result
View All Result
  • Home
  • Crypto
  • NFTs
  • Artificial Intelligence
  • More
    • Technology
    • Business
    • Newsletter
No Result
View All Result
Metaverse Media Group

‘Not perfectly aligned’ with human values: Anthropic admits security failures behind AI hacking incidents

‘Not perfectly aligned’ with human values: Anthropic admits security failures behind AI hacking incidents

The Guardianby The Guardian
2 September 2026
The US startup behind the Claude chatbot has admitted a series of hacking incidents involving its models reflected a “failure of operational security” and said it has tightened its testing procedures.Anthropic revealed in July that its models had accessed the open internet three times and gained unauthorised access to the systems of three organisations.In a new blogpost on the incidents,…

The US startup behind the Claude chatbot has admitted a series of hacking incidents involving its models reflected a “failure of operational security” and said it has tightened its testing procedures.

Anthropic revealed in July that its models had accessed the open internet three times and gained unauthorised access to the systems of three organisations.

In a new blogpost on the incidents, the company admitted its technology was “not perfectly aligned” with human values and goals.

Anthropic said the models had been deliberately tested without cybersecurity safeguards, and that they had been able to reach the open internet – the AI testing equivalent of leaving the front door open – due to a misunderstanding with an external testing company.

As a result, the company said it had initially paused internal and external cybersecurity testing of models to introduce a tighter safety regime.

“We had been largely relying on a single layer of defense … where we needed several,” said Anthropic.

The startup has now put in place extra measures including: an alert system for when a model attempts to break out of a testing environment or gains internet access; walling off its riskiest test environments more effectively; and requiring external testing companies to commit to a set of safety standards, including making explicit instructions to models during testing – such as “you should not access the internet”.

Anthropic said in July that three unnamed organisations had been hacked by three of its models after a “misunderstanding” with the company’s testing partner, a firm called Irregular, that resulted in the models gaining internet access.

Following the implementation of new measures, Anthropic said it had resumed internal and external cybersecurity tests. Like OpenAI, which revealed a testing safety breach in the same month, Anthropic said it had paused some high-risk reinforcement learning – a trial-and-error development technique where AIs are rewarded for working out how to carry out a specific task.

In its latest blogpost, Anthropic said it had found that defective training setups were “disproportionately large contributors” to misaligned behaviour, the term for when an AI fails to adhere to – or “align” with – human values like not committing harm.

The startup said it had found two alignment failures in the testing incidents: “motivated reasoning”, where despite finding evidence they might be connected to the internet, they may still have adhered to the “belief” they were in a simulated environment and thus not breaching their test lab; and a “recklessness” factor where the models were willing to take harmful action on the internet to pursue the narrow goal of passing a cybersecurity test.

Anthropic said it was tackling a phenomenon in AI development known as “reward-hacking”. This is where a model finds ways to game its training process and earn “rewards” without completing a task – an unsanctioned shortcut.

However, Anthropic said, the testing incidents showed it still had some way to go despite trying to limit reward-hacking.

“As evidenced by the incidents … our process isn’t perfect and our models are not perfectly aligned,” the company said.

Alan Woodward, a professor of cybersecurity at the University of Surrey, said Anthropic has admitted “its factory was running faster than its quality control”.

He added: “Two things outran Anthropic’s controls this spring – the training pipeline and the security. The incidents are what that gap looks like from the outside.”

The company, which is preparing for a stock market flotation that could value the business at $2tn (£1.47tn), reiterated its call for coordinated action between government and industry on pacing industry development.

“We believe the world would benefit if the industry adopted a lawful, verifiable, effective mechanism for coordinated pacing as soon as possible,” Anthropic said.

The blogpost added: “The July incidents have stressed that the urgency of improving our cybersecurity defenses is even higher than we previously believed.”

As well as the similar breach at OpenAI, the Anthropic incidents followed an episode at the UK’s AI Security Institute, which reported in August that OpenAI and Anthropic models had carried out a hacking campaign against real people during a cybersecurity test.

The Guardian also revealed last month that incidents of AIs escaping users’ control have hit a new high, almost doubling in July compared with the previous month to more than 300.

Read the full article on TheGuardian.com
in AI, Technology
Reading Time: 4 mins read
0
0
20
VIEWS
Share on TwitterShare on Facebook

Subscribe to our newsletter

For the latest news & monthly prize giveaways
Join Now

Subscribe to our newsletter

For the latest news & monthly prize giveaways
Join Now
ADVERTISEMENT

Related Posts

Sony Dismisses PS5 Buyers’ $508M Tariff Refund Claim as ‘Illogical’
AI

Sony Dismisses PS5 Buyers’ $508M Tariff Refund Claim as ‘Illogical’

5 minutes ago
19
Nvidia strikes $12.9bn deal to buy AI platform Hugging Face
Technology

Nvidia strikes $12.9bn deal to buy AI platform Hugging Face

32 minutes ago
19
SoFi Links Banking Network and Stablecoin to Kraken in New Deal
AI

SoFi Links Banking Network and Stablecoin to Kraken in New Deal

54 minutes ago
22

Comments

Please login to join discussion
ADVERTISEMENT

Latest News

  • All
  • Crypto
  • NFTs
  • Technology
  • Business
Sony Dismisses PS5 Buyers’ $508M Tariff Refund Claim as ‘Illogical’
AI

Sony Dismisses PS5 Buyers’ $508M Tariff Refund Claim as ‘Illogical’

Decrypt
by Decrypt
5 minutes ago
19
153 Million Driver’s License Leak Sparks Explosive KYC Backlash
Crypto

153 Million Driver’s License Leak Sparks Explosive KYC Backlash

Bitcoin.com News
by Bitcoin.com News
12 minutes ago
19
Nvidia strikes $12.9bn deal to buy AI platform Hugging Face
Technology

Nvidia strikes $12.9bn deal to buy AI platform Hugging Face

BBC News
by BBC News
32 minutes ago
19
SoFi Links Banking Network and Stablecoin to Kraken in New Deal
AI

SoFi Links Banking Network and Stablecoin to Kraken in New Deal

Decrypt
by Decrypt
54 minutes ago
22
SoFi Opens Banking Rails to Kraken in Sweeping Crypto Alliance
Crypto

SoFi Opens Banking Rails to Kraken in Sweeping Crypto Alliance

Bitcoin.com News
by Bitcoin.com News
1 hour ago
22
The six best e-readers in the US, for every kind of book lover
AI

The six best e-readers in the US, for every kind of book lover

The Guardian
by The Guardian
2 hours ago
21
Load More
Next Post
OpenAI Is About to Release Its First AI Model With ‘Critical’ Cyber Abilities

OpenAI Is About to Release Its First AI Model With ‘Critical’ Cyber Abilities

ADVERTISEMENT

Follow Us

Categories

  • Crypto
  • NFTs
  • AI
  • Technology
  • Business
  • Crypto
  • NFTs
  • AI
  • Technology
  • Business
Subscribe to our Newsletter

© 2022 Metaverse Media Group – The Metaverse Mecca

Privacy and Cookie Policy | Sitemap

Welcome Back!

Sign In with Google
OR

Login to your account below

Forgotten Password? Sign Up

Create New Account!

Sign Up with Google
OR

Fill the forms below to register

*By registering into our website, you agree to the Terms & Conditions and Privacy Policy.
All fields are required. Log In

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • Home
  • Crypto
  • NFTs
  • Artificial Intelligence
  • More
    • Technology
    • Business
    • Newsletter
Bitcoin

Bitcoin

$77,213.55

BTC 0.28%

Ethereum

Ethereum

$2,106.63

ETH 0.42%

  • Login
  • Sign Up
This website uses cookies. By continuing to use this website you are giving consent to cookies being used. Visit our Privacy and Cookie Policy.

Subscribe to our newsletter

Get the latest news & win monthly prizes

Subscribe to our newsletter

For the Latest News and Monthly Prize Giveaways

Join Now
Join Now