The data part is often the most difficult part of building an AI application: Srikant Goclnatha, Oracle

Applications of AI


By Abhijit Ahaskar & Uday Bhaskarwar

The proliferation of AI models and agents has left businesses with no choice, while also exposing critical gaps in their data infrastructure. Enterprises are increasingly realizing that the success of GenAI and agent AI is closely tied to the quality of the underlying data, making robust data governance and monitoring an essential prerequisite.

In a conversation with CXOtoday, Srikant Gokulnatha, Oracle’s senior vice president of Oracle AI Data Platform, Analytics, and Analytical Application Products, talks about the critical role data plays in the success of agent AI and what enterprises need to do to solve the problem of data fragmentation. Gokulnatha also pointed out that the AI ​​model and agent ecosystem is evolving so rapidly that companies should not be tied to any vendor or training method. Edited excerpt:

Q. What are the key data quality standards you recommend for companies to get the most value from GenAI and Agent AI?

The quality of data often determines the success of agent functionality in AI applications. The challenge is that unlike traditional analytics implementations, which have the luxury of bringing data into a warehouse for periodic loading, cleansing, and transformation, agent applications involve data that moves at high speeds.

Data may flow out of the warehouse, but at the same time unstructured data and real-time transactional data may flow in. While standard data quality techniques are still being applied, we are also seeing LLM used to measure data quality and enrich data. We are still in the learning phase. However, as AI adoption increases, we are discovering new techniques to ensure that the data is of sufficient quality to deliver robust results.

Q. As AI adoption increases, so does the amount of data it generates. Gartner recently warned that future LLMs will be trained more on output from previous AI models, increasing the risk of model collapse that further degrades the quality of AI output. Should this be something companies should worry about?

In general, this is a concern because many of the solutions we are building use LLM functionality. However, our solution applies it to the actual data that exists within the enterprise. So for the applications we promote, we don’t need to worry too much. Model collapse is a major concern when trying to build general-purpose apps based on synthetically generated public data using LLM. For business applications, this is less of an issue because the data is internal to the enterprise.

Q. Many companies suffer from fragmented data stored in silos. How does Oracle’s unified platform simplify this complexity and help CXOs make better decisions in real-time?

Our answer to that is to enable our customers to use an enterprise lakehouse approach that inherently recognizes that data resides across the enterprise and may come from a variety of systems. It could be oracle, it could be non-oracle, it could be relational, it could be unstructured, it could be historical, it could be real-time, it could be images or videos.

Therefore, data can be obtained in all kinds of formats. The lakehouse approach is useful for these solutions because data needs to be acquired from many locations, geometries, and modalities. The data part is often the most difficult part of building an AI application. You can pull data from all these different sources and manage it centrally.

Q. Is data fragmentation a bigger problem in India compared to other regions?

i don’t think so. This is a fairly generic problem and a characteristic of how businesses grow. If you’re starting a new business and need to define everything, you can adopt a more centralized data approach. Companies grow organically. Different departments often do different things or acquire new types of businesses. This complexity is inherent in how business models evolve. In some ways, India may have surpassed some traditional challenges. For example, if you look at banks in the United States, they often run mission-critical databases that were built 30 years ago. Because these systems still perform their functions, there is a reluctance to move away from them even though new systems are being introduced at the same time.

Q. Some hyperscalers offer companies the option to train frontier models directly on their own data. In your opinion, does this initial domain pre-training provide a significant upgrade in performance, or do you think fine-tuning is still the best approach to get the most benefit from Frontier models for enterprise customers?

Now, it looks like more tweaks are being made. A variety of new techniques are emerging. For example, there’s the concept of relational-based models, where models are pre-trained on large datasets and can predict patterns and make predictions without actually building a machine learning (ML) model. Therefore, various types of models are being actively researched and constructed.

Of course, we are considering all of them. However, I don’t think the pre-training part has become very common yet.

Another challenge is that these models are evolving so quickly that for a given agent or application, the model you chose six months ago is often no longer valid and is no longer the same vendor’s model. We currently do not limit our customers to any particular model, and some of the agents we are building are using two or three different models from different vendors. The functionality of these models changes dramatically over a period of 3-4 months, so you don’t want to be tied to a pattern.

Q. Whatever is happening in the AI ​​space is a continuum where models are constantly iterated and improved. But if this is a gradual transition, how do you explain the sudden volatility we saw a few weeks ago?

We don’t have much control over that, so we’re not too worried about it. What we know is that there are three things that are of great value to our customers.

First, much of the world’s data resides in Oracle systems. Make that data easily available without moving it. Moving data introduces high cost of ownership issues and leaves data pipelines vulnerable. Makes all your Oracle data highly effective.

Second, our system includes a lot of business context. When we build and deploy our systems, we put a lot of knowledge into configuring them in specific ways and building or configuring them for specific business processes.

You can automatically measure what it is and use that to inform your model. Therefore, the concept of context graph has become a very popular concept as it allows AI agents to operate more effectively. Third, once the agent has an insight or action that it wants to take, it needs to do so in the context of a business process. Most of the business processes reside within the Oracle system.

Q. As AI begins to make more autonomous decisions in Oracle Fusion applications, how do you ensure oversight and trust in AI-driven results?

It actually exists in multiple layers. Within OCI itself, there is a set of built-in guardrails for all of these models, regardless of vendor, as well as a set of customer-configurable guardrails. So, for example, if you are using Oracle AI Data Platform and do not want your end users or developers to use a set of models, you can define guardrails that apply to each as part of that configuration process.

At the agent level, techniques must be applied. There are now various techniques, such as using the LLM as a judge to check for hallucinations or bias. However, we see many customers still using human interaction for their most critical processes. Some actions are highly automated, but often there is a threshold, such as an amount or severity level, below which the risk is considered low and is automated without much oversight.



Source link