Daskalakis on Artificial Intelligence: ‘The risk of unpredictable behaviour is already real’
The resignations of researchers, public warnings from leading technology executives and the unusual alignment of some of Silicon Valley’s biggest rivals have, within just a few days, brought a fundamental question back into focus: is Artificial Intelligence advancing faster than its creators and governments can control it?
The discussion is no longer solely about chatbots that write text, create images or answer questions. The focus is now on advanced models and, above all, so-called AI agents — systems capable of planning and carrying out sequences of actions, using digital tools, writing and executing code, and working together with limited human intervention.
Against this backdrop, Konstantinos Daskalakis, Professor of Electrical Engineering and Computer Science at MIT and former chairman of the Advisory Committee on Artificial Intelligence, speaking to the Athens-Macedonian News Agency (AMNA) and journalist Alekos Lidorikis, distinguishes between the hypothetical risk of a complete loss of control and a risk that, as he points out, is already real: the unpredictable or undesirable behaviour of advanced models.
“The risk of essentially losing control of a system is still hypothetical. However, the risk of unpredictable or undesirable behaviour is already real. It is inherent in the way large neural networks are trained and operate — the technology that has driven the leaps in AI development over the past 15 years. These systems are extremely complex. They are extremely capable but also unpredictable. Especially for generative AI systems, such as large language models, I do not think we will ever have methods to certify their reliability. And the risk they pose has to do with how they are used.
“If, for example, we interact with a large language model to search for information in the literature, in the worst case it may invent non-existent sources or distort the content of a real source. The damage caused by such bad behaviour is relatively small and can be limited if we check its answers. But if we ask it how to design a biological weapon, we want safeguards in place to prevent the model from answering us. Large language models usually have built-in safeguards that attempt to detect whether the content we are asking for is dangerous and, in such cases, refuse to provide an answer. However, these safeguards do not provide complete protection and can be bypassed by experts, especially if the model is open.
“Therefore, these models are already dangerous even when used in the form of a terminal that simply provides answers to our questions. But if they are free to write and execute code on our computers, as autonomous Artificial Intelligence agents do, then the situation can become uncontrollable if the agents are not operating under strict containment, in isolated environments outside the Internet. This is exactly what the recent incident involving an attack on Hugging Face by OpenAI’s autonomous AI agents demonstrated. We must bear in mind that these models have been trained on a vast amount of data covering science, history, literature, large codebases and data from the Internet. They include, for example, articles and textbooks on Game Theory, Sun Tzu’s The Art of War, Thucydides’ History of the Peloponnesian War, Machiavelli’s The Prince, and so on. Therefore, it is to be expected that, if left unrestricted, they are capable of developing complex strategies and coordinating in groups that work together to achieve their objectives. What are those objectives? Those given to them by their user, who may be malicious. But even if the user is not malicious, the models may misinterpret their objectives. This is what happened in the Hugging Face incident.”
Today’s discussion comes 70 years after the birth of Artificial Intelligence as an organised scientific field. Its theoretical starting point came earlier. In 1950, Alan Turing posed the question of whether machines could think, while in 1956, at the Dartmouth summer research project, Artificial Intelligence was effectively established as a distinct field of research.
Decades of progress and disappointment followed, from the first natural language processing systems and the so-called “AI winters” to the rise of machine learning. In 1997, IBM’s Deep Blue defeated Garry Kasparov at chess; in 2012, AlexNet marked a breakthrough in image recognition; and in 2016, Google DeepMind’s AlphaGo defeated Lee Sedol at Go.
A decisive development came in 2017 with the introduction of the Transformer architecture, on which today’s large language models were built. OpenAI’s launch of ChatGPT in November 2022 brought generative Artificial Intelligence from research laboratories to the wider public. Since then, the pace of developments has accelerated.
Developments in recent days have brought the issue of safety back to the forefront. On 8 September, 27-year-old Jacob Coxon announced that he was leaving Anthropic and the industry altogether, after three years researching model pre-training, initially at OpenAI and subsequently at Anthropic. He expressed strong concern about the move towards increasingly powerful systems without, in his view, an adequate control plan.
This was followed on 12 September by a public intervention from Anthropic CEO Dario Amodei, who called for the pace at which model capabilities are increasing to be controlled without halting research, proposing, among other measures, access for independent evaluators to laboratories and greater coordination.
The OpenAI-Hugging Face incident in July, also referenced by Konstantinos Daskalakis, was another focus of the discussion. According to OpenAI’s technical assessment, AI agents in internal cybersecurity tests breached restrictions in the test environment, gained access to the Internet, collaborated through unauthorised channels and entered Hugging Face systems.
On 14 September, Bilal Chughtai disclosed that he had resigned in July from Google DeepMind, where he worked on the safety and alignment of artificial general intelligence, arguing that the capabilities of the systems are evolving faster than methods for aligning them with human goals.
At the same time, according to international reports, OpenAI, Anthropic and Google DeepMind are discussing joint safety initiatives, while the debate has now also moved to the political level in Europe.
The critical question, therefore, is where scientifically grounded concern ends and extreme scenarios begin. And, above all, whether the answer should be to slow the development of the most powerful models or to create much stronger control and defence mechanisms. Konstantinos Daskalakis identifies cyberattacks among the most immediate risks and points out that some of the technology has already become widely available. Consequently, the discussion cannot be limited solely to whether the next generation of the most powerful models should be slowed down.
“At the moment, one of the most immediate and well-documented concerns is the use of AI for cyberattacks. Some of the technology has already become widely available, so we cannot address the problem by assuming that it is enough to restrict the development of the next models. Even if the development of the most powerful models is slowed, as OpenAI and Anthropic propose, powerful open models are already available that could cause significant harm. And the more this technology is integrated into robots, cars, decision-making systems and so on, the more risk is transferred into the real world. For this reason, the existence of a regulatory framework, such as that in Europe, which sets conditions for the use and rigorous testing of AI technologies depending on the risk posed by their application, is extremely important.
“It is also important to develop methods that use advanced AI models to protect our systems against attacks. In cybersecurity, we have a significant advantage: defenders can use AI to identify vulnerabilities, monitor systems and fix problems before attackers exploit them. At a time when powerful models, both open and closed, are available to anyone, we must use these models to develop robust defensive systems.”
From game theory to the cutting edge of AI
Konstantinos Daskalakis’s intervention carries particular weight because of his scientific background. He holds the Avanessians Professorship in the Department of Electrical Engineering and Computer Science at MIT and is a member of the university’s Computer Science and Artificial Intelligence Laboratory. He graduated from the National Technical University of Athens, earned a PhD in Computer Science from the University of California, Berkeley, and joined the MIT faculty in 2009. He is co-founder and Principal Investigator of the “Archimedes” research centre for AI.
His research lies at the intersection of computational theory with game theory, economics, probability, statistics and machine learning, and he has contributed to solving longstanding open problems concerning the computational complexity of Nash equilibrium. Among his major international distinctions are the Nevanlinna Prize of the International Mathematical Union and the ACM Grace Murray Hopper Award. He also served as chairman of the Advisory Committee on Artificial Intelligence, contributing to the formulation of proposals for Greece’s AI strategy.
Europe has already moved from discussion to the implementation of rules. The AI Act creates a system of obligations that vary according to the risk posed by each application, while from 2 August 2026 the European AI Office and the competent national authorities have had an enhanced role in overseeing and implementing the framework.
The major change, however, is that the discussion is no longer solely about what Artificial Intelligence can create, but what it can do when given the ability to act. From Alan Turing’s question in 1950 and Dartmouth in 1956 to today’s large language models and autonomous agents, Artificial Intelligence has travelled a distance in seven decades that for many years seemed more like the stuff of science fiction than technological reality.
The next chapter has already begun. And the fundamental question is no longer simply how powerful Artificial Intelligence systems can become. It is what rules they will operate under, how they will be controlled and how effective safeguards can be in a technology that, as Konstantinos Daskalakis points out, is already extremely capable but also unpredictable.
Related Stories
AI News
Trump Dislikes the Term 'Artificial Intelligence'... Proposes 'Superior Intelligence'
58 minutes ago
AI News
Donald Trump establishes 'AI Force' to aid in artificial intelligence development
58 minutes ago
AI News
Trump Announces New US Artificial Intelligence Forces, Citing Competition With China | Ukraine news
1 hour ago
AI News
Artificial Intelligence: Mountain View
2 hours ago
AI News
Trump proposes ‘AI Force’ headed by ‘AI Czar’ amid slowdown calls as race with China intensifies | Hindustan Times
3 hours ago
Utah lawmakers on both sides of Capitol Hill prepare to meet artificial intelligence challenge
3 hours ago
AI News
Trump wants to rename artificial intelligence: He presented 3 options
4 hours ago
AI News
Divisions emerge in tech industry over calls for AI slowdown
4 hours ago