How Data Quality Shapes Machine Learning and AI Outcomes

AI For Business


“Garbage in, garbage out” is a well-known adage in programming. But it’s especially true in AI and machine learning, where model performance is often driven by the quality and relevance of training data.

AI and ML developers use the prepared datasets to train the models and algorithms they create and generate output. Their output is an in-depth analysis of the data that reveals trends and insights, or in the case of tools like ChatGPT, a platform that promises to answer user questions on seemingly endless topics.

Business leaders looking to deploy AI tools find a wealth of data sets available as inputs, covering areas such as healthcare, automotive and autonomous vehicle (AV) manufacturing, and finance and banking. need to recognize. The range of data sets used to train AI models is vast, and especially in the era of generative AI, the data must be relevant and of sufficient quality to meet end-user needs.

“Data quality is very important in the fields of data science and machine learning. [and] “AI has been around since the dawn of time. More people are discussing it,” he said.

Model performance depends on data quality and specificity

Techniques such as feature engineering and ensemble modeling can partially compensate for bad or inadequate training data, but the quality of the input data usually sets an upper limit on a model’s potential performance.

Therefore, ensuring data quality is critical to the success of business AI and ML initiatives.

“It’s clear that you can make terrible models out of high-quality data,” says Carlson. “But the quality of the data limits what the model can do.”

Six dimensions of data quality: accuracy, completeness, consistency, relevance, completeness and timeliness.
Various factors are considered to determine the quality of the dataset.

Companies use AI models for specific reasons. In other words, enterprise models require training with customized and relevant datasets. Therefore, when evaluating what data to retrieve and use, it’s important to consider the end systems that use that data. “How do you know what quality you’re trying to achieve until you’ve decided what you’re going to use the data for?” Carlson said.

Because of the importance of data relevance and specificity, popular but highly versatile models such as GPT-4 are not always optimal for enterprise use cases. A model trained on a large but unspecified dataset is unlikely to contain a good representative sample of the types of conversations, tasks, and data relevant to a particular industry or organizational workflow.

Instead of viewing data as objectively good or bad, think of data quality as a relative characteristic that is closely tied to the real-world purpose of your model. Even if the dataset itself is comprehensive, unique, and well-structured, if the team cannot use it to make the necessary predictions for the planned use case, the organization will not be able to achieve the desired outcomes. It may turn out to be possible.

As an example, Carlson shared his experience with a previous project for an electronic medical records platform. His team found that, despite having extensive data on how doctors used the platform, they could not predict when customers would leave the service. The decision to switch services was made by the clinic manager who did not use the platform directly. That is, their actions were not tracked.

“What you get is incredibly high quality data that is completely useless,” says Carlson. “It was poor quality data for what we wanted to use it for.”

Specialized datasets exist for a wide range of industries

While it takes resources and time for organizations to effectively train AI models, they now have easy access to the industry-specific datasets themselves.

In the financial sector, sites such as Data.gov and the American Economic Association feature data sets that provide macroeconomic data on US employment, economic output, trade, and many other related topics. On the other hand, the official websites of the International Monetary Fund and the World Bank provide datasets covering global financial markets and financial institutions.

Certain data sets in Data.gov’s vast catalog include titles such as “Car Sales” and “Food Price Outlook.” These types of datasets are provided by the US Department of Transportation and USDA respectively and are useful for specific business use cases within the financial sector.

Many of these data sets are available free of charge to enterprises. Just as ChatGPT was trained on text culled from various websites, articles, and online forums, companies are expected to look online and in data marketplaces for information to bring their models up to speed. increase.

Ethics and privacy considerations are part of data quality decisions

However, as organizations seek to incorporate external datasets and models, the underlying data collection practices are coming under increasing scrutiny.

“The question is, when we talk about AI, are models generated as a result of non-consensual information?” said Daniel Barber, CEO of privacy management platform DataGrail.

OpenAI, creator of ChatGPT, has already started facing lawsuits over its use of personal data. When evaluating whether to use or collect data from outside the organization, it is imperative that organizations consider data ethics and privacy considerations in a structured manner from the outset.

“The first step in making sure a business is adopting the right approach is to establish an ethics policy on how the business operates with AI,” said Barber. This internal ethics policy was formulated by the AI ​​Ethics Council and should be reviewed regularly to ensure it is functioning as intended. Similarly, organizations should appoint a data protection officer to be involved in decisions about the use or acquisition of data from outside the organization.

Incorporate diverse perspectives in developing ethics policies and planning AI initiatives. A group of individuals from a variety of backgrounds and functions offers potential outcomes not possible with a technical team alone.

“Data quality efforts are often siloed, simply doing their own thing and working against historical goals, largely disconnected from what the results look like. If you’re doing it, you’re much less likely to succeed at something,” Carlson said.

Ensuring data quality is necessary to avoid real-world impact

The need becomes even more apparent in certain areas, where the lack of reliable data can hurt consumers.

For example, high-quality data is essential when developing AV algorithms in the automotive industry. Companies are consistently working to improve the capabilities of their AVs to prevent serious real-world impacts. The datasets available for AV algorithms typically feature data captured from his LIDAR and camera systems on real self-driving cars to improve object detection and motion prediction.

In the healthcare industry, AI and ML are eagerly embraced as a way to not only address tedious administrative tasks, but also aid in diagnosis. High-quality datasets are therefore particularly important when training AI to understand health problems well enough to avoid misdiagnosis.

The website HealthData.gov features datasets on the impact of the COVID-19 pandemic in the United States. Text is not the only type of data relevant to the medical field. For example, thousands of chest X-ray images. can also be analyzed.

When evaluating whether and how to use medical datasets, keep in mind that user privacy and data ethics are often the most important areas in these areas. Barber pointed out that health-related information and biometric data are among the most sensitive types of data that can be collected about an individual.

“I think most people understand why that information is particularly sensitive to individuals,” he says. “So how was that information collected and was consent included in that process? That is very important for businesses to understand.”

Failure to ensure data privacy and security impacts business

Using data that is later found to violate privacy laws and industry standards can have serious repercussions for your company.

Nor should companies ignore this issue simply because they risk being fined in the future. Violations of security and privacy regulations not only have a negative economic and reputational impact, but can also force companies to remove algorithms and software that rely on illegally and unethically obtained data.

In May, the US Federal Trade Commission (FTC) accused Ring, an Amazon-owned company that sells internet-connected home security cameras, for not restricting internal access to its customers’ videos. settled the lawsuit for violating the privacy of For example, one employee watched thousands of videos of areas such as bathrooms and bedrooms from a female user’s device, according to the complaint.

“Ring’s disregard for privacy and security has exposed consumers to espionage and harassment,” FTC Consumer Protection Director Samuel Levin said in a press release.

And, as the complaint states, Ring used these videos to train its algorithms without user consent, which could have far-reaching repercussions for the company. Under the proposed settlement order, which is currently awaiting court approval, Ring will be required to remove data, models and algorithms derived from illegally reviewed videos.

If such outcomes become the norm, “the real business risk here is greater than just the compliance factor,” Barber said. “Rather, a wrong implementation can erode the business value of an entire model that may have taken hundreds of hours to build.”



Source link

Leave a Reply

Your email address will not be published. Required fields are marked *