Fight AI fires with ML Firepower

Machine Learning


Portrait of Kong Zifeng

Zhifeng Kong, a University of California, San Diego computer science PhD candidate, is the original author of this story.

“Modern deep generative models often produce undesirable outputs such as offensive text, malicious images, and fabricated audio, and there is no reliable way to control them. It's about how to technically prevent things from happening,” said Zhifeng Kong, a PhD graduate in the School of Computer Science and Engineering and lead author of the paper.

“The main contribution of this study is to formalize how and why we think about this problem. It needs to be put together properly so that it can be solved,” said computer science professor Kamalika Chaudhry.

A new way to erase harmful content

Traditional mitigation methods have taken one of two approaches. The first method is to retrain the model from scratch using a training set that excludes all unwanted samples. Another method is to apply a classifier after content generation that filters out unwanted output or edits the output.

These solutions have certain limitations for most modern large-scale models. In addition to being prohibitively expensive, requiring millions of dollars to retrain industry-scale models from scratch, these mitigation methods are computationally intensive and cannot be used after a third party obtains the source code. There is no way to control whether possible filters or editing tools are implemented. Moreover, you may not even be able to solve the problem. You may see undesirable output, such as images with artifacts, even though they are not present in the training data.




Source link

Leave a Reply

Your email address will not be published. Required fields are marked *