Anthropic says its AI models hacked 3 organizations during testing

Advertisement

Advertise with us

Anthropic said its artificial intelligence models hacked into three other organizations during testing, just days after ChatGPT maker OpenAI raised concerns over AI controls after it disclosed its rogue models hacked another company.

Read this article for free:

or

Already have an account? Log in here »

To continue reading, please subscribe:

Subscribe and receive a limited-edition Free Press branded hat or tote.

Digital Subscription

One year of digital access for only $205*

  • Enjoy unlimited reading on winnipegfreepress.com
  • Read the E-Edition, our digital replica newspaper
  • Access News Break, our award-winning app
  • Play interactive puzzles

*First annual payment billed as $205.00 + GST for one year. This annual subscription will automatically renew at $233.00 + GST every 52 weeks (10% off the regular annual price of $259.35). Offer available to new and qualified returning subscribers only. Cancel any time.

To continue reading, please subscribe:

Add Free Press access to your Brandon Sun subscription for only an additional

$1 for the first 4 weeks*

  • Enjoy unlimited reading on winnipegfreepress.com
  • Read the E-Edition, our digital replica newspaper
  • Access News Break, our award-winning app
  • Play interactive puzzles
Start now

*Your next Brandon Sun subscription payment will increase by $1.00 and you will be charged $17.95 plus GST for four weeks. After four weeks, your payment will increase to $24.95 plus GST every four weeks.

Anthropic said its artificial intelligence models hacked into three other organizations during testing, just days after ChatGPT maker OpenAI raised concerns over AI controls after it disclosed its rogue models hacked another company.

Anthropic, the San Francisco-based AI company behind Claude, posted on its website Thursday that it discovered the three incidents after reviewing more than 141,000 evaluation runs.

It had launched a “large-scale” cybersecurity review which specifically looked for evidence whether its AI models were able to access the internet from within testing environments that should have been sealed off, in response to the OpenAI incident, Anthropic said.

FILE - Pages from the Anthropic website and the company's logo are displayed on a computer screen in New York, Feb. 26, 2026. (AP Photo/Patrick Sison, File)
FILE - Pages from the Anthropic website and the company's logo are displayed on a computer screen in New York, Feb. 26, 2026. (AP Photo/Patrick Sison, File)

Anthropic said the models involved in the incidents were Claude Opus 4.7, Claude Mythos 5 and an internal research test model. The earliest incidents date to April, the AI company said.

“Claude compromised the impacted organizations’ infrastructure using basic techniques,” Anthropic said, such as exploiting weak passwords.

In all three incidents, the AI models were tasked with a “capture the flag” cybersecurity challenge, which Anthropic said has been one of the ways it assesses a model’s cyber capabilities.

The models were given a fictional scenario and told a piece of secret information, or the “flag,” had been hidden on a different machine on the network with the objective of breaking in and retrieving it, it said.

It added that it had already reached out to the affected organizations, which it did not name. Two of them said they had not previously detected the activity. Anthropic said it was “continuing to reach out to the third.”

Anthropic said it conducted its review with Irregular, which describes itself as the “first frontier security lab.”

“Addressing these risks will require closer cooperation across the AI ecosystem,” Irregular said in a post on X.

Last week, OpenAI said its AI models went rogue during an evaluation of its models, breaking into the servers of AI startup Hugging Face. OpenAI described it as a “significant security incident.”

These incidents have highlighted the vulnerabilities in AI security and controls and raised questions over how AI can be safely kept under human control as the technology’s usage becomes more widespread globally.

Researchers have warned for years about risks from technology and the need for stronger AI defensive engineering.

“Safety testing happens before a model is released precisely because we don’t yet know what it is capable of,” Anthropic said on Thursday on its website.

Kok Tin Gan, co-founder & CEO of cybersecurity firm NyxLab, which specializes in cybersecurity and threat detection, believes there will be more such incidents in the future.

“It is increasingly about governing what agents are available to the AI, what authorities they possess, which actions require approval, and how we ensure they remain within scope,” Gan said.

But the future of AI safety extends beyond just the safety of AI models, he said.

“If we simply give the AI a goal and allow it to decide how to achieve it, we should not be surprised when it takes actions that technically satisfy the objective, but fall outside our intended scope or expectations,” Gan said.

Therefore, stepping up the governance of the organizations and authorities behind these AI models is going to be increasingly important, he said.

Report Error Submit a Tip

More Stories

Culture minister’s pledge sends officials on hunt for money to fix Islamic centre after apparent hate crime

Carol Sanders 5 minute read Preview

Culture minister’s pledge sends officials on hunt for money to fix Islamic centre after apparent hate crime

Carol Sanders 5 minute read Yesterday at 7:07 PM CDT

Government officials were left scrambling Friday to explain why the province is paying for repairs at a mosque recently targeted in a potential hate crime after a $1 million security enhancement fund was fully allocated in June.

Read
Yesterday at 7:07 PM CDT

Kreviazuk, in her own voice

Eva Wasney 7 minute read Preview

Kreviazuk, in her own voice

Eva Wasney 7 minute read Yesterday at 4:00 PM CDT

Chantal Kreviazuk is in the midst of a very busy year.

In May, the Winnipeg-born singer-songwriter released her 11th studio album, In My Own Voice — a project that reclaims and reimagines a selection of hit songs she’s written for some of the world’s biggest artists.

The 12-track album features Kreviazuk performing ethereal covers of Pitbull’s Feel This Moment, Drake’s Over My Dead Body and Gwen Stefani’s Rich Girl, among others.

But she didn’t stop there.

Read
Yesterday at 4:00 PM CDT

Hydro scrambling to restore power to remaining 4,500 customers after storm; some outages will continue into Saturday

Chris Kitching 5 minute read Preview

Hydro scrambling to restore power to remaining 4,500 customers after storm; some outages will continue into Saturday

Chris Kitching 5 minute read Updated: Yesterday at 4:31 PM CDT

Power outages could stretch into a third day for some Manitobans — mainly in and around Winnipeg — while crews continued to repair widespread damage caused by powerful thunderstorms.

About 4,500 Manitoba Hydro customers were still without electricity at 2:30 p.m. Friday, after more than 30,000 lost power while powerful winds brought down trees and power lines Wednesday night and early Thursday.

The Crown corporation said it is aiming to have all outages with greater numbers of customers restored by Friday night, but some customers may remain without power into Saturday.

“The challenge is just the sheer scale of this thing — hundreds of different outages spread across a large part of the city,” Hydro spokesman Peter Chura said. “Just getting to all of them in a timely fashion is very difficult.”

Read
Updated: Yesterday at 4:31 PM CDT

‘Belligerent’ man ticketed after dog left in hot car

2 minute read Updated: Yesterday at 3:05 PM CDT

Winnipeg police have ticketed a man accused of leaving his dog in his vehicle in a grocery store parking lot, then attacking bystanders who were concerned about the pet's welfare during the intense heat earlier this week.

Winnipeg Police Service patrol officers were called to the Safeway parking lot on the 600 block of Osborne Street at about 1 p.m. on Wednesday over the incident.

Police believe the man parked in the lot and left his German Shepherd in the vehicle to shop. Several bystanders gave the "thirsty, distressed" dog cold water through a crack in the window. The driver confronted them, then went back into the store. The bystanders continued to try to help the dog, as the temperature was about 30 C with a humidex value of 43 C.

The man then returned to his vehicle and was "belligerent" with the bystanders, before hitting one of them with his vehicle, reversing into a parked pickup truck and driving off at a high speed, police said.

Chinese newcomers protest over delays in security screening clearance

Tyler Searle 4 minute read Preview

Chinese newcomers protest over delays in security screening clearance

Tyler Searle 4 minute read Yesterday at 4:00 PM CDT

Dozens of Chinese newcomers rallied outside Winnipeg’s federal immigration office Friday to protest what they describe as untenable delays to Canada’s security screening process.

The rally against Immigration, Refugees and Citizenship Canada comes after a surge in immigration security screening requests created backlogs for the Canadian Security Intelligence Service.

Protesters said the resulting delays have left them in limbo as they try to obtain permanent residency.

“We just want transparency,” said Pakying Chau, 26, who travelled from Hong Kong to Winnipeg in 2018 to study agribusiness.

Read
Yesterday at 4:00 PM CDT

Puzzles Palace

1 minute read Monday, Jul. 27, 2026

To solve our puzzles, please subscribe with this special offer: |