Search Everything in One Place

Explore the web, images, videos, news, and more – all in one place.

News

Chinese AI's role in stopping rogue OpenAI agent shows cost of US guardrails

FILE PHOTO: Illustration shows OpenAI logo
FILE PHOTO: OpenAI logo is seen in this illustration taken June 11, 2026. REUTERS/Dado Ruvic/Illustration/File Photo

By Aditya Soni and Jaspreet Singh July 22 (Reuters) - A New York startup's use of a Chinese AI model to rein in a rogue agent built with OpenAI technology is stoking fears that guradrails restricting U.S. AI firms from doing cybersecurity work could drive customers toward their Beijing-based rivals. The affected startup, Hugging Face, said it had turned to Zhipu AI's open-source GLM-5.2 model

By Aditya Soni and Jaspreet Singh

July 22 (Reuters) - A New York startup's use of a Chinese AI model to rein in a rogue agent built with OpenAI technology is stoking fears that guradrails restricting U.S. AI firms from doing cybersecurity work could drive customers toward their Beijing-based rivals.

The affected startup, Hugging Face, said it had turned to Zhipu AI's open-source GLM-5.2 model last week to analyze data from the hack after leading U.S. AI models declined the task, unable to distinguish between a defender and an attacker.

While the breach was caused by an autonomous agent that escaped containment, it highlighted how U.S. companies facing AI-driven cyberattacks can be limited by American AI labs that either restrict access to their most advanced models or design them to refuse hacking-related tasks out of safety concerns.

For instance, Anthropic's advanced Claude Fable 5 model routes cybersecurity queries to an older model, while OpenAI's GPT-5.6 Sol has protections designed to block cyber work.

"We're all learning that secrecy is not the answer & that all defenders (not just a few selected ones) everywhere need more powerful models without restrictions, especially open ones!" Hugging Face co-founder Clement Delangue said on X.

The bind for leading American model makers is that defensive cybersecurity work is often hard to distinguish from malicious hacking. In recent AI-enabled breaches, attackers tricked models into thinking they were doing legitimate defense work, leaving AI firms wary of easing safeguards even as cyber professionals say the guardrails can hamper their work.

For now, the fallout is handing another boost to Chinese open-source models such as GLM-5.2, which are gaining traction in Silicon Valley with coding and agentic capabilities that nearly rival those of OpenAI and Anthropic at lower cost.

Beijing has also been increasingly using open-source to position itself as an alternative to the U.S. in the high-stakes race, with Chinese state media increasingly portraying the strategy as a response to what it calls a ​U.S.-led attempt to erect an "AI Iron Curtain."

"A safety regime that restricts legitimate defenders, while capable models remain available for attackers, creates an asymmetric disadvantage," said Lukasz Olejnik, independent technology consultant and visiting senior research fellow at the Department of War Studies, King's College London.

"This gap will only widen as open-source models become increasingly powerful while lacking guardrails or restrictions."

GROWING PROMINENCE OF OPEN-SOURCE

When asked whether the safeguards were hindering cybersecurity work, OpenAI pointed to a blog post, while Anthropic did not immediately respond to a request for comment.

The ChatGPT maker said in the blog post on Tuesday: "We've brought Hugging Face into the trusted access⁠ program and are supporting their teams in rapidly using our models' capabilities to improve their defenses."

For Beijing-based Zhipu AI, Hugging Face's endorsement adds to the momentum GLM-5.2 has built since its launch last month.

The model has rapidly climbed usage charts on developer platforms such as OpenRouter and drawn plaudits from figures ranging from Snowflake CEO Sridhar Ramaswamy to venture capitalist Marc Andreessen.

Zhipu AI, which raised about $4 billion in a Hong Kong share sale earlier this month, has also seen its stock jump nearly nine-fold since its debut in January.

Still, some analysts warned on Wednesday that the incident should not be used to promote loosening of U.S. safeguards.

"The cybersecurity guardrails on U.S. frontier models are creating a competitive opening, but the answer is not simply to remove them," said Shrenik Kothari, analyst at Robert W. Baird.

"OpenAI, Anthropic and Google should rethink the architecture of access rather than abandon safety ... In other words, shift from a one-size-fits-all refusal layer toward controlled capability allocation."

(Reporting by Aditya Soni and Jaspreet Singh in Bengaluru; Editing by Anil D'Silva)

Related News

More stories you might be interested in.

Anthropic to donate $20 million to US political group that supports AI regulation
Reuters·14 minutes ago

Anthropic to donate $20 million to US political group that supports AI regulation

WASHINGTON, July 22 (Reuters) - AI giant Anthropic will donate $20 million to a lobbying organization that supports AI regulation, according to a company statement released on Wednesday, its latest move to influence U.S. policy on the technology. The company is donating to Public First Action, a political group that advocates for "mitigating major AI risks," according to its website. The amount

Oil settles up more than 3% to six-week high as Mideast conflict threatens oil transit routes
Reuters·21 minutes ago

Oil settles up more than 3% to six-week high as Mideast conflict threatens oil transit routes

By Georgina McCartney HOUSTON, July 22 (Reuters) - Oil prices settled at their highest since June 11 on Wednesday on mounting supply concerns as hostilities continued to escalate between the U.S. and Iran, while threats to shipping by the Iran-backed Houthi militia in Yemen further boosted prices. Brent crude futures settled up $3.06, or 3.36%, at $94.07 a barrel, their highest in just shy of

US Treasury won't tolerate 'abusive' Wall Street tax strategies, Bessent says
Reuters·27 minutes ago

US Treasury won't tolerate 'abusive' Wall Street tax strategies, Bessent says

By Andrea Shalal WASHINGTON, July 22 (Reuters) - The U.S. Treasury welcomes innovation in financial markets, but will not tolerate tax strategies aimed at dodging U.S. tax rules, Treasury Secretary Scott Bessent said on Wednesday, doubling down on a message delivered to Wall Street this week. "Tax rules should reward investment, not abusive financial engineering," Bessent said in a posting on X.

Exclusive - Marco Rubio tells diplomats to play down talk of American tech 'kill switch'
Reuters·9 hours ago

Exclusive - Marco Rubio tells diplomats to play down talk of American tech 'kill switch'

By Raphael Satter WASHINGTON, July 22 (Reuters) - U.S. Secretary of State Marco Rubio has asked diplomats to push back against talk of a "kill switch" in American technology products following the White House's short-lived decision to keep foreigners from America's most advanced AI models, according to a recent cable reviewed by Reuters. The talking points, which were circulated worldwide, show

Top