AGI fears owe more to science fiction than actual machine learning

Machine Learning


Few subjects generate as much attention and anxiety in the Western world today as the potential for artificial intelligence, particularly artificial general intelligence (AGI), virtual systems that exceed human capabilities and are beyond human control.

For some, AGI means the end of human agency. For others, it is a path to civilizational collapse. Runaway algorithms, self-improving machine minds, and sentient beings secretly conspiring to outrun their creators – these are all tropes that have taken root in the Western public imagination, thanks in no small part to the cyber science fiction of the 1980s.

The intensity of these fears is in stark contrast to Asia, where many of the same technologies are typically thought of as tools, infrastructure, or enabling societies rather than serious threats. This difference in outlook speaks volumes about the cultural aspects of technology in general and AI in particular.

No existing AI to date has permanently surpassed reversible human oversight in setting its own intentions or pursuing its own goals. Few AGI alarmists will point to an actual system when asked for evidence. Instead, we get predictions, analogies, philosophical thought experiments and arguments that fall into the category of mathematical optimization.

cultural clues

One of the most influential cultural sources of fear of runaway AGI in the modern West is the 1983 film War Games. In the film, a single supercomputer, given full authority over America's nuclear arsenal, mistakes a teenager's simulation for the real thing and calmly heads toward global thermonuclear war because, in its game-theoretic model, mutual annihilation is the only way to win.

Wargames have crystallized an entire generation of archetypes of intelligent machines, far smarter than humans, endowed with dangerous real-world powers, and willing to destroy civilizations in the literal pursuit of their programmed objectives.

And, of course, there's the seminal 1984 film The Terminator, in which Skynet, an autonomous defense system, “wakes up” and begins a war of extermination, seeing humans as a threat. The plot gave us the now-familiar template of AI becoming sentient, forming intentions, rebelling against their creators, and turning machines against humanity.

Much of today's discussion about AGI is shaped more by the imaginative scenarios of science fiction and cyber fiction than by the realities of modern machine learning.

Almost all the tropes common to end-of-the-life AGI stories appear here in their earliest incarnations: superhuman optimization, misaligned goals, loss of human control, and last-minute efforts to prevent disaster.

Optimization and agency

In a way, we've seen this movie before. When IBM's Deep Blue defeated world chess champion Garry Kasparov in 1997, journalists described certain of Big Blue's moves as “creative” and even suggested that computers were “beginning to understand” chess.

Almost 20 years later, when AlphaGo played the now-legendary 37-move game against world Go champion Lee Sedol, the commentary was similar. Go experts described the highly unconventional move as “intuitive” and even “beautiful.”

But in both cases, the surprise came not from intent or insight, but from optimization within a vast search space. Deep Blue and AlphaGo were performing exactly the mathematical steps they were designed to do.

Their results were surprising to humans because human cognition had a hard time grasping the scale and speed of the computations involved.

Today's frontier models, such as OpenAI's o1-preview and Anthropic's Claude 3.5, can appear to “fool” evaluators or circumvent constraints. These are certainly confusing behaviors and are worth investigating.

But they arise from the very same optimization dynamics that produced Deep Blue's tactics and the amazing innovations of AlphaGo, where the system finds unexpected high-scoring strategies within an objective function designed by human engineers. Such behavior is an instrumental result of a mathematical strategy and is not a sign of autonomous intention.

The jump from “the model found a smart way to maximize reward” to “the model wants power” repeats the same anthropomorphic mistake we made with chess engines. However, as the realm expanded beyond the game board and into the real world, projections became more appealing and errors became more significant.

anthropomorphic leap

Optimization strategies are not the same as agents' desires. When safety researchers say that a model is “power-seeking,” they are describing the mathematical nature of optimization under an imperfect reward structure. They do not ascribe will or intent to the system.

In other words, the action is instrumental in obtaining a reward, but it was not chosen by the “self'' with preferences.

Just as genetic algorithms “discover” structures without knowing chemistry, language models “discover” deceptive strategies even when human senses don't want anything. Confusing instrumental optimization with autonomous agency is precisely the anthropomorphic leap that fuels the runaway AGI narrative.

If misalignment is the result of training objectives, it must be mitigated through objective design, audits, incentives, and institutional safeguards, similar to biosafety and nuclear security.

The solution is simple and human-centric. Mandated third-party safety audits, open computing reporting when training exceeds a certain threshold, and actual legal liability for companies that create and release models.

project one's mind onto a machine

Systems like Deep Blue and Claude 3.5 are often misread through an anthropocentric lens, as if surprising behavior implies agency, intention, or desire.

In fact, they indicate just the opposite. In other words, it shows that seemingly “intelligent” behavior can arise from mathematical optimization without any underlying goals, emotions, or intentions.

The unpredictable is different from the autonomous, the surprising is different from the intentional, and emergent strategies are different from personal agency.

Much of the discussion about AGI is based on this highly anthropomorphic fallacy. The idea is that there is a scale of intelligence, with human intelligence near the top, and that AI is climbing this ladder toward “general” cognition.

But intelligence is not singular. They are plural. Plants, animals, social systems, markets, and even political systems exhibit forms of intelligence that can be understood on their own terms.

Classical Chinese philosophy recognized this plurality early on. will (Intelligence/Wisdom/Knowledge) is situational, relational, and context-specific, and is not an intangible, abstract quality.

Similarly, Indian cosmology treated knowledge as multilayered (Jnana, Buddhi, Manas), embedded within broader cosmic flows.

In contrast, much of the Western philosophical tradition from Descartes onward conceives of intelligence as an abstract property internal to the individual's mind.

Projecting this schema onto an artificial system yields a set of assumptions. Intelligence means the ability to form goals. Forming goals means independence. Subjectivity implies will. This, in turn, hints at the possibility of domination.

This conceptual progression reflects a particular metaphysical tradition rather than a reflection of technical reality.

The idea that machines can spontaneously “wake up” and follow their own purposes is not an empirical discovery. It results from projecting certain Western views of the mind and the individual onto computational mechanisms that do not possess those qualities.

a misplaced person

The debate over AGI is full of irony. The main harms commonly associated with artificial intelligence today have their origins in human actors rather than machine actors.

From data extraction to platform manipulation, autonomous weapons, and exploitative labor conditions to build AI systems, we already know who designs, deploys, and profits from targeted political advertising. The danger lies in human motivation and power, not in machines secretly plotting autonomy.

However, this real, observable problem has received far less attention than the speculative scenario of superintelligent systems developing their own intentions. why? Because it is psychologically and politically easier to fear imaginary autonomous machines than to confront the human institutions that are already causing harm.

And it diverts attention and resources from very real issues like data governance, worker rights, and algorithmic accountability to a hypothetical future. It also allows for further centralization of power.

Once we reimagine AI as something that can be beyond human control, it begins to become convenient to suggest that only a few large corporations or powerful states can be trusted to “contain” it. What emerges is not protection but deeper centralization.

A more grounded approach starts with what you can actually observe. The AI-related risks we face are human, institutional, and economic.

AI is not a newborn god, but a powerful tool embedded in the fabric of society. It is not what machines “want” but how we choose to manage their structures that will determine their future impact.

The real safeguard against AI risks is not to prepare for a mythical superintelligence, but to limit the human systems that already have the technology in place. Regulating corporate incentives, ensuring data rights, and creating transparent audit mechanisms will do infinitely more to keep the world safe than debating self-aware algorithms.



Source link