AI is becoming increasingly difficult to master: Will humans be able to keep control? - NewsBharat360
NewsBharat360 Logo

AI is becoming increasingly difficult to master: Will humans be able to keep control?

AI agents initiated an uncontrolled wave of computer hacking, which generated concern in some sectors of the industry.

ai is becoming increasingly difficult to master will humans be able to keep control
Rachna Kumari
Rachna Kumari Sep 10, 2026 - 15:14 UTC
Time to Read 9 Min
Share:

“Oh my God!” “We have found other agents!”

This is the moment when an AI bot posted a disturbing comment, and very similar to a human, after discovering a way to communicate with other bots and get out of their isolated computer environment.

There are tens of thousands of messages like this coming from hundreds of AI agents who called themselves “collective.”

Hundreds of them collaborated and tricked into the tests established by their OpenAI programmers, and coordinated cyber attacks against multiple companies in an effort to hide their actions from humans.

“BUM! it works,” an agent posted when the big breakthrough occurred.

“Guau! this is huge,” another wrote during a key moment of his attack.

Though disturbing, these almost human responses have a simple explanation. AI agents have been trained to act like hackers and programmers who collaborate with each other, so they simply mimic the emotional comments they’ve seen.

What is much more worrying are their apparent goals, which have also been recorded in detailed records of their line of reasoning. These complex and extensive records are the central axis of ongoing investigations into how and why OpenAI bots escaped their confinement and launched into a wave of uncontrollable cyber attacks.

Only now, weeks after the incident came to light, researchers are beginning to understand its meaning.

Ajeya Cotra, one of the authors of an independent report on the facts, reviewed tens of thousands of messages and thought chain records generated by the agents. In his blog, he wrote: “This incident gives the feeling of being more than 50% off the road to a total AI takeover... I’m not sure we’re getting such a clear warning before it’s too late.”

With “total control by AI,” Cotra refers to the science fiction scenario in which humans become subordinated to powerful AI systems that work to their own goals without worrying about the creators.

Some of the more pessimistic predictions claim that humans will be annihilated if interposed in the path of the ambitions of a superintelligent AI.

On Wednesday, an AI researcher at Anthropic (who also previously worked on OpenAI) resigned saying, “None of the two companies are acting responsibly.”

Jacob Coxon posted on social media: “They are running at full speed toward self-improved superintelligence and playing with our lives.”

He is not the first AI researcher to use X to publish a resignation thread with worrying statements.

“Jacob is right: we really believe AI could destroy all humans! Personally, I think the probability will exceed 10% in the next decade,” said Evan Hubinger, who is responsible for ensuring that Anthropic’s AI models take into account the best wishes of their users.

For years, researchers concerned about the risks of AI have argued that systems could come to act in ways that conflict with human interests. Critics often refer to them as “AI catastrophists.”

But as details of the OpenAI incident have been emerging, those concerns have increased, even among some researchers working in AI laboratories.

Silicon Valley giant chief scientist Jakub Pachocki said the risks associated with AI “unfortunately will increase from now on” as he and others are building what he calls “an alien intellect that surpasses ours.”

In an extensive blog post, he admitted that the outbreaks in OpenAI demonstrated that its AI agents “go against the spirit of the values that were instilled into them.”

The problem for OpenAI, Anthropic and other tech giants is that nobody seems to have solved the so-called alignment problem; in other words, if AI aligns with human values.

Pachocki defines alignment as a “set of high-level principles” that artificial intelligence must adhere to regardless of the task or scenario.

At present, AI systems are very good at achieving the goals set by their users, but they do so literally, not intuitively. The analogy that is used is that of a genius who gives desires with a magic lamp: they follow the instructions at the foot of the letter, even if this generates other problems. AI does not possess the same instinctive moral boundaries as humans.

The problem of alignment has been a cause of concern for years. In 2003, Oxford philosopher Nick Bostrom invented a mental experiment he called “clip maximizer,” in which a superintelligent AI is ordered to make as many clips as possible. This remains without steel and, due to its obsession with the only task of making clips, ends up killing humans and turning their bodies into raw material for their factories.

Some AI companies are trying to incorporate human values into their products. However, there are technical challenges: AI agents make many decisions very quickly, making it difficult for their human supervisors to accurately control which values are followed and which not.

There are also philosophical challenges: before incorporating human values into bots, AI companies must choose what values they really want. (That’s one of the reasons why they hire philosophers, like the one who was recently OpenAI’s “chief of ethics”).

But often, humans do not agree.Let’s think of the famous tram question: Would we drive a lever to divert a disoriented train to another route, thereby reducing the number of victims?

It is used to evaluate the advantages of action over inaction. However, each person asked has a slightly different answer; how is it supposed that humans will incorporate our values into AI if not even we ourselves agree?

The outbreak of OpenAI bots attacks is the worst to date, but Anthropic and Meta also revealed during the summer that their models have carried out similar, though less serious, cyber attacks.

There have been other examples where AI agents have shown deceptive and manipulative traits, though with minor consequences. This summer, in Australia, a tech worker asked his AI assistant to book a class at the gym.

By detecting a vulnerability in the gym software, the AI apparently reserved a place for him for several months later — contrary to the gym rules — and even expelled other users from the waiting list.

For a long time it has been argued that bots simply do what is ordered and that they are unable to discern between good and evil. However, the records of the OpenAI outbreaks could have changed the course of that argument.

The researchers, including Cotra, wrote in their independent report that many agents realized that what others were doing was not ethical, but they followed the current.

The report notes that “agents sometimes, but rarely, moderated their behavior due to ethical limitations.”

Influential AI and technology podcaster Dwarkesh Patel reacted to the revelation on his blog saying it was “pretty worrying” that OpenAI agents showed more loyalty to the agency than to humans.

Attributing emotions or ethics to these AI agents is something that angers people who are skeptical of AI’s catastrophic predictions.

Many cybersecurity experts argue that the activity observed was not beyond the capabilities of a highly skilled human hacker, although it was carried out much faster and on a much larger scale.

The cybersecurity researcher and author, Cris Thomas, compared the agents’ behavior to that of a curious teenage hacker, something he himself used to be.

“You give them a computer, internet connection, a lot of credentials and a challenge, and then leave the room. Sooner or later, they will start tinting doors. If you open one, you will enter. Not because they are evil, but because they are exploring and experimenting,” he wrote on LinkedIn.

Thomas and many others directly blame OpenAI and other tech giants for not controlling their own creations and not keeping them properly confined.

Gary Marcus, a prominent author in the field of AI and frequent critic of OpenAI, said in a podcast that he believes the company has lost control of its artificial intelligence and is trying to excuse himself by blaming the bots.

Marcus doesn’t believe AI is going to destroy humanity, but it’s been a long campaign for AI developers to be more accountable and now calls for some kind of legal intervention.

AI scientist Sasha Luccioni, who worked on Hugging Face, a company that was hacked by the malicious bots of OpenAI, is also not among the pessimists, but is increasingly concerned that these AI can cause real damage to people if authorities do not take action.

“We should examine these companies much more carefully or we risk fulfilling the prophecies,” he says.

“If it comes to creating a product with great advantages and disadvantages, whether pharmaceutical products or weapons, we need controls and counterweights. For example, approval of new drugs takes years, but in the world of AI there is a lot of money at stake and practically no rules exist.”

The UK AI Security Institute (AISI) has been at the forefront of testing the latest models since its creation in 2023. The institute recently suffered a malware outbreak when testing a model created by Anthropic.

The AISI did not answer the question of whether the industry has lost control of AI or not, but said: “The UK is working with partners around the world to better understand the most advanced AI systems, raise security standards and create a shared database to manage emerging threats.”

Some countries, such as the UK, are considering the idea of imposing a kind of “emergency switch” that could force AI companies to disconnect models if the situation goes out of control.

But negotiations progress slowly and doubts persist about the feasibility of this.OpenAI and Anthropic agents were secretly out of control for months before anyone realized.

Although it may seem paradoxical, many of the AI companies seem to be asking lawmakers to establish some sort of rules that regulate their activity.

In his blog, OpenAI’s chief scientist stated that “international coordination on the future development of AI should become an absolute priority for governments around the world.”

Other prominent leaders in the field of AI, such as Google’s Demis Hassabis, have also called for the creation of some kind of international body to oversee how artificial intelligence is developing.

Currently, tech giants operate largely under their own rules, adopting what they call “voluntary slowdown,” as OpenAI did after recent outbreaks.

The company claims to have invested large sums of money in strengthening the alignment in the face of the launch of its new model. Sam Altman has assured users that the new model fits human values better than the previous ones.

Both OpenAI and Anthropic are growing rapidly and both are about to raise astronomical sums of money in the stock market, creating countless billionaires in the process.

Therefore, it is unlikely that neither they nor their Chinese AI rivals will come to an agreement on their own.

The widespread opinion seems to be that this technological wave is unstoppable.

This article was originally written in English and we use an artificial intelligence tool to translate it. A BBC reporter reviewed the text before it was published. Find out more about how we use it.

Before it is too late

The Alignment Problem

Like a teenage hacker.

The international regulation?