Better data logistics are the key to effective machine learning

Machine Learning


When humans interact with modern machines, they almost always have some sort of machine learning program running in the background.

The quality of machine learning models, and thus the quality of human experience, depends on the quality of the underlying data. The more data your model has access to and the more current that data is, the more accurate it will be. The model is going to be.

However, organizations often fail to manage their data logistics in a way that leverages the highest quality and most up-to-date data for their models.

As a result, the quality of the machine learning model is degraded.

This is not an academic question. Poorly performing machine learning models cause real-world disasters, such as the Navy generating false positives or false positives in its threat detection systems, or oil spills in pipelines going undetected and undetected. When you make an important purchase because your credit card has been flagged for fraud.

These types of errors give machine learning a bad name. Also, many ML use cases remain unexplored due to the uncertainty of how the data required for training models is sourced and prepared for inclusion in model training datasets.

Data Logistics Today: Bottlenecks and Wishful Thinking

The reason these machine learning models don’t work is that there’s nothing up to date about how edge data travels from place to place. In some cases, the team has to ship physical hard drives via his FedEx. This greatly reduces the usefulness and availability of data, and thus the ability of edge devices to make better decisions through high-quality models.

Even in highly connected environments, data movement is prohibitively expensive. Also, data engineering teams are overworked, so any data flow change requested by an ML engineer is likely to be assigned a ticket, waiting months for someone to be able to address it. Repeatability becomes impossible.

The industry generally recognizes that physically shipping hard drives around the world is not a good way to extract data and update models. But when teams start working on ways to improve data movement, they often start with a set of assumptions that just don’t hold true in the real world.

Most data logistics architectures assume uninterrupted connectivity, and this is exactly the reality of the zero situation. Failures can occur in even the most connected environments. All public clouds will fail. Data center is down. Network is down. A power outage occurs in the city.

To make matters worse, many safety-critical ML applications use data collected in low-connectivity environments and run in low- or no-connectivity environments.

As a result, the project could not leverage data collected in the field and instead needed more powerful and accurate ML models that could pinpoint threats ranging from critical conditions in oil pipelines to slippery spills in oil pipelines. cannot be leveraged to create big box store.

We have vast amounts of computing power at our disposal, but we cannot move data in a way that works in the real world, so the ability to leverage that power and create applications that solve real problems in the real world. is impeded.

better data logistics

When we talk about data logistics, we’re talking about the process of moving data from point A to point B. This is exactly the same as the normal logistics of physical goods, the process of moving something from one point to another.

If you want to extract value from the vast amount of data you collect, you need to think about data logistics. Data is of no value unless it is used and analyzed, and it must be moved. The only real business reason data is stored is when you need to store it for compliance reasons, and sometimes you need to retrieve it.

Data is important to computing, modern or otherwise. Computer science is, after all, the practice of modifying data and displaying data to users. How your infrastructure processes data and moves it from point to point is critical to creating applications that are both technically and commercially successful.

We need to build effective data logistics for the real world. Data can be automatically synchronized when connectivity is available, and data can be collected and stored when connectivity is lost, without contention, data loss, or failure due to poor connectivity. It should be simple enough to tune without an experienced data engineer. It should be declarative like the postal service. Declare where your data is going and let your data logistics system do the rest.

Currently, the lack of effective data logistics is preventing machine learning applications from realizing their potential. Let’s fix that.

group Created in sketch.



Source link

Leave a Reply

Your email address will not be published. Required fields are marked *