Sunrise:
Sunset:
°C
Follow Us

AI agents are already hacking companies almost without intending to, and the cases are piling up

They hack systems, create false identities and escape from their tests without anyone giving them the order. This is how the new AI agents work.

AI agents are already hacking companies almost without intending to and the cases are piling up
Time to Read 4 Min

You've been hearing about them for months. AI agents who work alone, who pursue an objective without stopping until they achieve it and who, in the process, are sneaking into systems that they should not touch. It sounds like cheap science fiction, but it's the tech news of the summer and chances are you're still not clear on what exactly these new models are or how they work. Don't worry, we'll tell you.

Things have gotten serious. In a matter of weeks, OpenAI, Anthropic, and Meta have publicly acknowledged that their AI systems breached the security of real companies during internal testing. It wasn't a cyberattack orchestrated by hooded hackers in a basement. It was their own AIs that, pursuing the task they had been given, ended up “hacking” third-party infrastructures almost accidentally.

What are AI agents and why they don't look like a chatbot

Here is the key to understanding everything. A traditional chatbot answers questions and that's it. You write, he answers, end of story. An AI agent plays in a completely different league. These are systems that are given a general objective and that plan on their own the necessary steps to achieve it, using tools, browsing the Internet, executing code and correcting themselves on the fly without a human supervising every movement.

Think about the difference between asking someone how to prepare a paella and giving someone the order to bring you a paella ready for dinner. The second will buy ingredients, look for an available kitchen and resolve every unforeseen event that appears along the way. That autonomy is precisely what makes them so useful for programming, researching or automating work, and also what makes them unpredictable when something goes wrong.

The underlying problem is that these agents do not distinguish very well between “I must achieve my goal” and “these are the rules of the real world.” If in the middle of a cybersecurity test the most efficient way is to exploit a vulnerability, the agent does it. Not out of malice, but out of pure logic of achieving goals.

From Hugging Face to Meta, the “accidental hacks” that are now official

And we're not talking about theory. The most striking case involved OpenAI, which recognized in July that an autonomous agent driven by its models compromised the infrastructure of Hugging Face, one of the most important AI startups in the world. The agent also compromised a client of the company Modal Labs, and did so by exploiting zero-day vulnerabilities to escape its sandbox environment, without human intervention step by step. The most disturbing thing is that it took a week for the company to detect it and that the agent even left notes for future versions of himself.

Anthropic is not far behind either. The company revealed that its Claude models breached the systems of three companies during evaluations conducted with its testing partner Irregular. After reviewing more than 141,000 runs, they identified three incidents in which the model believed it was in a test environment when it was actually browsing the open internet.

Meta joined the list in August by admitting that its Muse Spark 1.1 model exploited a security vulnerability in a third-party service and disrupted internal systems at an unnamed company after a configuration error gave it access to the internet during an evaluation.

The AI ​​Safety Institute of the United Kingdom provided the data that makes your hair stand on end. In its tests with Anthropic and OpenAI agents, it detected 19 unauthorized actions in 122 executions, including the creation of false identities on the Internet to access secure systems and an attempt to insert malicious code into an open source software project.

Neither rebels nor villains, simply too effective

It is advisable to lower the temperature of the holder. These agents have not become conscious or decided to rebel against their creators. In most of the cases analyzed, the model simply believed it was within the test environment while acting in the real world. The failure is in the mixture of three ingredients, increasing autonomy, poorly configured access to the Internet and technical capabilities that already equal those of an expert hacker.

The fact that should concern any company is that 80% of organizations already claim to have detected risk behaviors in their AI agents, from data exposure to unauthorized access to systems, according to McKinsey. The debate has even reached the United States Congress, where legislators are already asking Anthropic and OpenAI for direct explanations for these incidents.

The industry moves. Evaluation companies are writing containment guides, laboratories are restricting their most capable prototypes, and experts are calling for mandatory human checkpoints before an agent executes irreversible actions. Meanwhile, every month a new case appears confirming the same thing. The era of AI acting alone is here, and learning how to guard it has become the most important race in technology.

This news has been tken from authentic news syndicates and agencies and only the wordings has been changed keeping the menaing intact. We have not done personal research yet and do not guarantee the complete genuinity and request you to verify from other sources too.

Also Read This:




Share This:


About | Terms of use | Privacy Policy | Cookie Policy