A team of researchers from the University of Waterloo has developed a new machine learning method that can detect hate speech on social media platforms with 88% accuracy, potentially saving employees hundreds of hours of emotionally damaging work.
The method, called multimodal discussion transformer (mDT), differs from traditional hate speech detection methods in that it understands the relationship between text and images and can place comments in a larger context, which is particularly useful for reducing false positives, where comments are often incorrectly flagged as hate speech due to culturally sensitive language.
“We sincerely hope that this technology can help reduce the emotional burden of humans manually sifting through hate speech,” said Liam Hebert, a computer science doctoral student at the University of Waterloo and first author of the study. “We believe that by taking a community-centered approach to applying AI, we can create safer online spaces for everyone.”
Researchers have been building models to analyze the meaning of human speech for years, but until now, these models have struggled to understand nuanced speech and contextual statements. Previous models could only identify hate speech with 74 percent accuracy, below what the Waterloo study was able to achieve.
“Context is extremely important in understanding hate speech,” Hebert said. “For example, the comment 'Gross!' may be harmless in itself, but its meaning changes dramatically depending on whether it's in response to a picture of a pizza with pineapple on it or to a member of a marginalized group.”
“Understanding that difference is easy for humans, but training a model to understand the contextual connections in a discussion, including taking into account images and other multimedia elements in the discussion, is actually a much harder problem.”
Unlike previous efforts, the Waterloo team built and trained their model on a dataset that included not only individual hate comments but also the context of those comments: the model was trained on 8,266 Reddit discussions containing 18,359 labeled comments from 850 communities.
“More than 3 billion people use social media every day,” Hebert said. “The influence of these social media platforms has reached unprecedented levels, making the detection of hate speech at scale urgently necessary to create a respectful and safe space for everyone.”
research, Multimodal Discussion Transformer: Integrating Text, Image, and Graph Transformers to Detect Hate Speech on Social Mediawas recently published in the Proceedings of the 38th AAAI Conference on Artificial Intelligence.
