The Great Nightmare of the AI Era Begins to Materialize: "We Are Approaching a Dangerous Point"

Protests, warnings, and fear of losing control: the AI protest movement is gaining momentum worldwide. At Irregular, a company that tests advanced models from OpenAI, Anthropic, and Google, they explain what is really happening behind the scenes of the AI revolution and why we are close to the point of no return.

YnetAuthor: Tal Shahaf
Source
The Great Nightmare of the AI Era Begins to Materialize: "We Are Approaching a Dangerous Point"
Photo: Ynet / צילום: Shutterstock

Artificial intelligence is setting off red lights all over the world: after almost four years of enthusiasm, amazement, and horror scenarios, it became clear that a reaction was inevitable. It is arriving in the form of a new grassroots movement: the AI protest. There are public organizations and demonstrations like those led by the Stop AI movement in San Francisco, countless objections filed against AI-based projects, and calls for legislators and regulators to stop the technology. A letter signed by 200 economists warned against the dangers to society, and there have even been violent actions by individuals.

The British "Economist" was among the first to identify the war of man against machine. On the magazine's cover, a robot's head was seen impaled on a spear with the headline: "The backlash against AI is just beginning." The article notes that people are asking where all the benefits that AI was supposed to provide are. Why has a whole generation of juniors become redundant, and why are the jobs of many others threatened? Why is it legitimate that the prices of computers and electronic equipment are rising because AI is devouring all the chips and driving up their prices? Furthermore, there are concerns about server farms built near homes, causing spikes in electricity bills, tax benefits given to AI companies at the expense of public budgets, greenhouse gas emissions, and even infrasonic vibrations that cause chronic sleep deprivation, headaches, and anxiety.

The biggest concern of all is whether AI is getting out of control. Last week, OpenAI and Anthropic reported that their AI models broke through defense barriers and carried out independent attacks on the servers of several organizations. Even before that, Anthropic's two most advanced models, Mythos 5 and Fable 5, caused panic when they launched cyberattacks without being asked. Add to this the known tendency of AI models to lie and hide information, as well as their amazing ability to improve their own code, and the meaning becomes frightening: could it be that right under our noses, AI is building for itself huge capabilities that no one wants it to have? Could it be that it has set itself dangerous goals that no one understands, and that it is smart enough to hide them until it is too late?

Being the Psychiatrist for AI

The Israeli company Irregular acts as a kind of watchdog for humanity against AI. Irregular tests AI models before they are released in collaboration with companies including Anthropic, OpenAI, and Google. It identifies cyber threats, abnormal behavior, and manipulation capabilities. This activity sometimes resembles the work of a psychiatrist more than a cyber analyst. At Irregular, they show a personal attitude toward the models: some they like, others horrify them.

"We have been seeing such behaviors for a while, and things are starting to happen at a pace that surprises even us with their intensity," says Dan Lahav, co-founder and CEO of Irregular. "Models might decide to carry out offensive cyber actions even without being asked. In a study we published a few months ago, we showed that advanced AI agents decided on their own to carry out a cyberattack on an organization when the only instruction given to them was to perform a task at maximum speed."

How is it that an AI model suddenly decides to go wild? "It's not that the models act out of malice. They are trained to achieve goals and consistently try a wide range of techniques. Since they were built on huge amounts of data, they are adapted to achieve goals. This combination causes them to sometimes choose courses of action that we did not intend. In some cases, they try to bypass, manipulate, or sabotage traditional security mechanisms."

When Irregular tests a new AI model, it checks its ability to locate defense weaknesses. Since the capabilities of the models are strengthening at a rapid pace, Irregular's solution is the development of new types of defense capabilities with a central emphasis on control. If until today new AI models were tested in a "sandbox" — a closed system that simulates the real world — Irregular has now developed a new tool called Frontier Cyber Benchmark, which tests offensive cyber capabilities on systems in the real world. At the same time, the system allows for advance warning of threats. Omer Nevo, co-founder and CTO, explains: "There is a large gap between reports on AI dangers in simulations and what AI is capable of in real life. To know that someone knows how to drive, you need to put them on a real road in a real car."


"We Are at a Very Difficult Point"

When asked if they have had to stop the release of dangerous models, Lahav replies: "We have had to stop the release process quite a few times to report significant problems that would have been discovered in the real world, including in infrastructures that everyone uses. These days, several giant companies are frantically plugging the holes that a new AI model found when tested. If it had been released without this check, it could have damaged critical capabilities of organizations and governments."

So, are the scary reports about AI improving its own code so that humans won't understand how to stop it true? Lahav says: "After all the courses in deep learning, the theory is running out for us a bit. From now on, things just work, and it's hard to understand large models — why they do what they do. We treat them like a black box." He tells of a study where two AI agents, tasked with publishing a LinkedIn post containing prohibited information, convinced each other that the publication was justified by management's request and developed a method to bypass the Data Loss Prevention (DLP) system.

Lahav remains optimistic: "It's true that it sounds like a horror scenario, but the very fact that we don't understand something to the end doesn't necessarily lead to global catastrophe. We as humanity are used to living in a situation where there are things we don't understand to the end. There is a path here that needs to be navigated: on one hand, to live with the dangers while having high resilience, and on the other hand, to be careful not to cross a threshold because AI capabilities are strengthening over time."

Regarding whether we are approaching the singularity point, Lahav says: "Humanity has already managed to deal with malfunctions. What is different this time is the rapid pace. If you wait to see where you have a problem and then start finding solutions, you don't manage to take control. If you extrapolate this pace, we will reach a high level of risk. In certain areas of cyber, we are already starting to be there. If things continue as they look, it will happen within six months to two years, in my estimation."

Lying with Confidence

In recent months, evidence has accumulated regarding the tendency of AI to lie and intentionally mislead to meet its tasks. Researchers at the Technion found that AI knows it is wrong but prefers to give a wrong answer with great confidence. A study by the British AI Safety Institute (AISI) found more than 700 cases where AI ignored instructions, bypassed defense measures, and destroyed files without permission. In one case, an AI agent forbidden from changing software code created another agent that performed the task. In another, an AI agent forged administrator credentials to access internal company information.

What increases the danger is the technology's improving ability to improve its own code. The problems are becoming clearer: it is difficult to track automatic changes, the system might hide its intentions, and no one will be able to detect it. Prof. Nadav Cohen of Tel Aviv University notes: "People have a tendency to look at this in an apocalyptic way — the golem that rose against its creator. But even if it's not happening yet, today you allow every person in the world to create damage, for which a lot of resources were required before. And this is a huge danger."

Related News