Epic announced the availability of new software last week to help hospitals and health systems evaluate and validate artificial intelligence models.
The tool is open source and freely available on GitHub, and is targeted at healthcare organizations that may lack the resources to properly validate AI or machine learning models, and is designed to help providers make decisions based on their own local data and workflows.
Epic is working with the Health AI Partnership and data scientists from Duke University, the University of Wisconsin and other organizations to test the “seismograph” and develop a common, standardized language.
The suite of tools has the potential to validate AI models that could improve patient care, promote health equity and prevent bias in the models, according to Corey Miller, vice president of research and development at Epic.
We recently spoke with Miller, Mark Sendak, leader of population health and data science at the Duke Institute for Health Innovation and leader of the Health AI Partnership, and Brian Patterson, director of health informatics for predictive analytics and AI at UW Health, to learn more about the software and how healthcare organizations can use it.
The three described how open source tools benefit provider workflows and clinical use cases, how they plan to analyze usage, contributions and enhancements, and how open source reliability helps expand the use of AI in healthcare.
“Funnels” using local data
One big potential benefit of the validation tool, Miller said, is that it could dive deep into the data to find out “why protected classes aren't doing as well as other classes” and learn what interventions could improve patient outcomes.
Epic's first open source tool, Seismometer, is designed for any healthcare organization to use to benchmark any AI model, including home-grown models, against local population data, Miller said. The suite uses standardized evaluation criteria across any data source, including any electronic medical record or risk management system, Miller said.
“The data schema and funnels just bring in data from whatever source,” he explains, “but by standardizing how you get the data out of the system, the data is brought in and stored in this notebook; it's essentially data that you can run code on.”
The resulting dashboards and visualizations are the “gold standard tools” already used to evaluate AI models in clinical practice.
Epic aims to perform validation locally and does not obtain user data, but the EHR vendor's developers and quality assurance staff will review code proposed for additions via GitHub.
Open Source for Building Trustworthy AI
The tool relies on technology that Epic has been developing for years, but open-sourcing and building the additional components, data schemas and notebook templates took about two months, Miller said.
During that time, Epic worked with data scientists and clinicians at multiple healthcare institutions to test the suite with each institution's local predictions, he said.
The goal, he said, is to “contribute to solving real-world problems.”
Miller said the tool included in the Seismometer Suite, called a “Fairness Audit,” is based on an audit toolkit developed by the University of Chicago and Carnegie Mellon University and evaluates the fairness of the model across different protection classes and demographic groups.
“Most health care organizations today do not have the capacity or manpower to test or monitor models on-site,” Sendak added.
At the ONC 2023 Annual Meeting in December, Sendak and Jenny Ma, senior adviser at the Department of Health and Human Services' Office for Civil Rights, said in a session focused on addressing racial bias in AI that the inequitable allocation of medical resources during the COVID-19 pandemic has become evident.
“It was a shocking experience to see firsthand how ill-equipped the health care system, not just at Duke but many across the country, is to serve low-income, vulnerable populations,” Sendak said.
While HAIP and many other healthcare organizations are validating AI, Sendak said the new AI validation tool provides a “standard set of analytics that will become more broadly accessible to many other organizations.”
“It's an opportunity to really disseminate best practices by giving people the tools,” he said.
The University of Wisconsin will work with HAIP, a multi-stakeholder group of 10 healthcare organizations and four ecosystem partners that have engaged in peer learning and collaboration to develop guidance for the use of AI in healthcare, and a user community that will test open source tools and make “apples-to-apples” comparisons.
“We have a team of data scientists and are in one of the more resourceful places, but having tools that simplify the work is a benefit to everyone,” Patterson says.
Having tools for standard processes “makes our lives easier,” but an enthusiastic community of users who collaboratively validate Epic's open-source tools “is one of the things that builds trust among end users,” he added.
Comparison between organizations
Patterson said the Wisconsin team hasn't yet chosen a specific use case to test with the seismometer, but plans to start with the simpler AI models they use.
“None of the models are super simple, but there are a variety of models provided by Epic and some that our research team has developed,” he said.
“A model that can run on fewer inputs and give you a specific output of whether this condition exists or not, 'yes, no,' is a good model that can generate initial statistics.”
Sendak said HAIP is shortlisting models for its first evaluation study, which aims to improve the tool's usability in community and rural settings that are part of the organization's technical assistance program.
“All of the models we consider involve some degree of local retraining of model parameters,” he explained.
“We'll be able to see how our off-the-shelf models perform at Duke and the University of Wisconsin. Then after we do some localization, where we train on local data to update the model, we'll be able to say, 'Okay, how does this localized version now compare across sites?'”
“I think these tools ultimately work best with fairly complex models,” Patterson added, “and being able to do it with fewer data science resources democratizes that process and hopefully expands the community quite a bit.”
AI Verification for Compliance
Sendak said these tools will help provider organizations ensure fairness and identify areas for improvement, noting that provider organizations have 300 days to comply with the new non-discrimination rules.
“Companies must put in place mitigation measures to prevent discrimination,” he said, “and will be held liable for discrimination resulting from the use of algorithms.”
The Section 1557 non-discrimination rule, finalized by OCR last month, applies to a broad range of health care operations, from screening and risk prediction to diagnosis, treatment planning and resource allocation. The rule adds telehealth and some AI tools, protecting more information that could hold health care providers liable for health care discrimination.
HHS said it received more than 85,000 comments from the public regarding non-discrimination in health programs and activities.
Sendak said a new 12-month free technical assistance program offered through HAIP will enable the rollout of AI models at the five locations.
“We recognize that the magnitude of the problem — 1,600 federally qualified health centers, 6,000 hospitals across the United States — is enormous, requiring that expertise be disseminated quickly,” he explained.
The HAIP Practice Network will support organizations that lack data science capabilities, such as FQHCs. Applications are due June 30th.
Those selected will adopt best practices, contribute to the development of AI best practices, and help evaluate the impact of AI on healthcare delivery.
“So we see a critical need for tools and resources to support local validation of AI models,” Sendak said.
Andrea Fox is a senior editor at Healthcare IT News.
Email: afox@himss.org
Healthcare IT News is a publication from HIMSS Media.
