AI cascading failures: How to protect your organization

AI For Business


It is well known that the flapping of a butterfly's wings in South America can cause a tornado in the Caribbean. The so-called butterfly effect (technically known as “sensitive dependence on initial conditions”) is of great importance for organizations considering implementing AI solutions. As systems become more interconnected with AI capabilities impacting more and more critical functions, the risk of cascading failures (local failures that ripple through to disruption throughout the organization) increases significantly.

It is natural to focus AI risk management efforts on individual systems where clear risks can be easily identified. Senior executives might ask how much the company could lose if the predictive model made an inaccurate prediction. How exposed are we if a chatbot provides information it shouldn't? What if we encounter edge cases that new automated systems can't handle? These are all important questions. However, focusing solely on these types of issues can provide a false sense of security. The most dangerous AI failures are not confined to one specific area. They are the ones who spread it.

How cascading failures work

Many AI systems currently operate as isolated nodes, but it is only when they are combined across an organization that artificial intelligence can fully fulfill its promise. A network of AI agents that communicate across departments. An automated ordering system that links customer service chatbots to logistics hubs and even factory floors. An executive decision support model that pulls information from every corner of the organization. These are the types of AI implementations that offer transformative value. But they are also the types of systems that create the greatest risk.

Consider how quickly the problem will multiply. Corrupted data at a single collection point can compromise the output of all downstream analysis tools. A security flaw in one model becomes a gateway to every system it touches. Additionally, when multiple AI applications compete for the same computing resources, the spike in demand can reduce overall performance, which is often a disaster.

When AI is siled, failures are suppressed. As AI becomes interconnected, failures can propagate in ways that are difficult to predict and become even more difficult to stop.

The 2010 US stock market “flash crash” showed that algorithms can interact in unexpected ways and cause problems of unimaginable scale. On the morning of May 6, more than $1 trillion was wiped from the value of the Dow Jones Industrial Average in minutes as automated systems triggered a downward spiral. Despite several years of investigation, the exact cause of the crash is still unknown.

The flash crash demonstrated that when autonomous systems interact, their combined behavior can deviate significantly from what a single system is programmed to do. None of the algorithms were designed to collapse the market, and if operated independently, they would not have. However, the interaction between them (each responding to signals generated by the others) produced unexpected results at the system level that diverged from the goals of any part of the system.



Source link