(Tee11/Shutterstock)
A great way to streamline access to data is to create data products. Rather than simply exposing raw data to data scientists, data products offer a more sophisticated and governed approach. But without some level of built-in quality, data users cannot trust data products.
The concept of data products dates back to 2012 when DJ Patil, then Chief Data Scientist in the United States, published his book Data Jujitsu: The Art of Turning Data into Product. According to Patil, a data product is one that “facilitates the achievement of an end goal through the use of data.”
As we move into 2019, our data products will Data Nami People to watch in 2022. According to Deghani’s original article in his defining data mesh, “Domain data teams should apply product thinking to the datasets they serve. Think of your data assets as your product, and think of other data within your organization. We see scientists, machine learning and data engineers as our customers.”
Thomas H. Davenport, Randy Bean, and Shail Jain further the concept of data products in a 2022 Harvard Business Review article, in which they describe data products as “creating reusable datasets that can be analyzed in many ways. attempt to do so.” Move different users over time to solve a specific business problem. “
The authors of HBR further refined the term by distinguishing between data products that are good for reuse and analytics products with built-in analytics or AI capabilities. In either case, companies implementing data products should consider creating a new position to oversee the creation and use of data products: the data product manager.
One of the data product manager’s jobs is to ensure the high quality of the data. Anthony Deighton, recently promoted to the role of data product general manager at Tamr, sees similarities between a typical software development product manager and a data product manager.
(Peschkova/Shutterstock)
“Product managers think about customer goals and map them to product features,” he says. “Data He is a product manager a lot like this. He thinks about a business goal and maps it to the characteristics of the data. Is there? It’s very similar.”
Tamr was co-founded by legendary computer scientist and database creator Mike Stonebreaker and recently took over the baton of data products. The company was originally founded to commercialize the “data tamer system” that Stonebraker helped develop at his MIT and wrote in 2012. This helps curb the data chaos that exists in many large enterprises with many data silos. (“The more databases you create, the more silos you create. So perhaps Tamr has a microphone to help solve the problem he caused by creating more databases around the world. It might be the way we do things,” says Dayton.
According to Tamr’s recent research, businesses are turning to data products to address the customer data problem. He found that 69% of his respondents ranked “business value” as the most important metric for measuring the success of their data products, second only to user experience.
“This is very interesting,” says Dayton. “Because they’re not thinking about technical metrics like uptime, performance, or data volume, they’re talking about tying it to the business value they’re delivering to their customers and partners.”
Nearly three-quarters of respondents to the survey said companies wanting to solve the customer data problem should develop a data product strategy that focuses on business value rather than other tactics. was also found. That’s no surprise. He joined Tamr a few years ago after working for his Qlik for many years.
“Data is a real asset. [But] It’s crazy,” he says. “That’s impacting the customer relationship. You’re a business owner and you say, ‘If I could do a better job with customer data, I could offer a better service, product, etc.’ to my customers.” It’s a serious business challenge. “
Data silos (TFoxFoto/Shutterstock)
Tamr uses machine learning and AI to help companies find errors in their data. The company’s technology has been developed to extend entity resolution capabilities across many silos, giving enterprises a clearer, more distinct view of their assets, and better support for downstream AI and analytics use cases. It will give you a good starting point.
“Our view from a data product perspective is, can we use machine learning as a mechanism to automate that process and let machines do the work instead of relying on humans,” says Dayton. “The idea that computers can be tasked with tasks that were previously considered necessary by humans is becoming much more mainstream.”
Quality and trust issues tend to get more complex as the amount of customer data grows. Instead of relying on data to automate more customer interactions, businesses must rely on manual methods of interaction until the data can be cleansed (which itself can be a time-consuming manual task). often).
I’m looking for a better approach.
“Customer data is where the primary interface pointing to the customer faces this post-apocalyptic trash can fire. That is your data. And we are all going through this,” Dayton said. says.
We repeat that when you call a horizontally integrated telecom company to address a concern about an account, it gets handed off to other departments where the concern is reiterated and handed off again. We experience it when we call the hospital for test results and the nurse says, “Oh, that’s a whole different system, isn’t it?”
“We experience this instinctively all the time,” Dayton says. “And what this research shows is that organizations are going through that as well. Literally not having access to your data doesn’t feel good, it’s just as frustrating as not being able to answer questions, fill prescriptions, or whatever the problem is on the other end of the conversation. is.”
There is no escaping Conway’s Law that the design of a system reflects the organization that built it and how it communicates. Conway’s Law leads to specialization and compartmentalization of the system. The corollary that Dayton derives from this law is what he calls Dayton’s Law, which states that the data in an organization also reflects the way the organization is structured.
“Organizing the company by product means that customer data is organized by product silo,” he says. “Or, if you organize your business geographically, if you have Europe and the United States, and you have the northeastern and southeastern United States, then you’re organizing your customer data geographically. If you organize your business by go-to-market segments, There is a large enterprise versus a midsize market, and that’s how the data is organized.”
However, companies can begin to work out Dayton’s Law piece by piece with their data product strategy. “The idea behind data products is can you articulate the view of the data at the top of your organizational structure and the underlying data structure,” he says.
Moving data to the cloud (huge data warehouses, data lakes, data lakehouses) has removed some of the technical barriers that kept data in isolated silos. In some cases, the physical separation was eliminated, which was very helpful. It’s the organizational structure of the company that currently holds back progress, and that’s where a data product strategy can help.
“What our customers are doing is pulling data from these multiple silos and creating two tables in the data lake, but they are not related and we have no idea how to integrate them. You just don’t understand,” says Deighton. Say. “It is the data product that connects them.
“Warehouses and data lakes are necessary but not sufficient conditions for success,” he continues. “This is a key enabler, but a data product strategy is about being able to see across silos, whether those silos are in the data lake or in the system of record.”
Related products:
How to build great data products
How ML-Based Data Mastering Saves Millions of Dollars in Clinical Trials Business
Tamr Helps Air Force Data Management
