High-quality labeled data is required for many NLP applications, especially for training classifiers and evaluating the effectiveness of unsupervised models. For example, scholars frequently attempt to classify texts into different thematic or conceptual categories, filter noisy social media data for relevance, and assess moods and positions. Labeled data are necessary to provide a training set or benchmark for comparing results, regardless of whether supervised, semi-supervised, or unsupervised techniques are used for these tasks. Such data may be provided for semantic analysis, high-level tasks such as hate speech, and in some cases more specialized purposes such as party ideology.
Researchers usually need to make original annotations to ensure that labels correspond to conceptual categories. Until recently, there were only two basic approaches. For example, research assistants can be hired by researchers and trained as programmers. Second, she may rely on freelancers working on her website, such as Amazon Mechanical Turk (MTurk). These two approaches are frequently combined, with cloud her workers augmenting the labeled data and trained annotators generating small gold standard datasets. Each tactic has its own strengths and weaknesses. Training annotators often produce high quality data, even though the service is expensive.
However, there are concerns about the quality of MTurk data deteriorating. Other platforms such as CrowdFflower and FigureEight are no longer available for academic research after being acquired by business-focused organization Appen. Cloud employees are much more affordable and adaptable, but quality can be better, especially for difficult activities and languages other than English. A researcher from the University of Zurich, with particular emphasis on his ChatGPT, published in November 2022, explores the potential of large-scale language models (LLMs) for text annotation tasks. This study shows that zero-shot ChatGPT classification outperforms them at a fraction of the cost of MTurk annotation (i.e. without additional training).
🚀 Build high-quality training datasets, solve NLP machine learning challenges, and develop powerful ML applications with Kili Technology
LLM has worked very well on a variety of tasks such as classifying legislative ideas, scaling ideologies, solving cognitive psychology problems, and emulating human samples for research studies. Some research has found that ChatGPT can perform the kind of text annotation tasks they specified, but to the best of their knowledge it has not been fully evaluated yet. A sample of his 2,382 tweets collected for preliminary research was used for analysis. In this project, tweets were annotated by trained annotators (research assistants) against her five separate tasks: relevance, pose, subject, and her two types of frame identification.
They distributed jobs to cloud workers in MTurk and zero-shot classification in ChatGPT using the same codebook they created to train their research assistants. They then evaluated ChatGPT’s performance against his two benchmarks. (i) accuracy relative to crowd workers; (ii) inter-coder agreement when compared to both crowd workers and their trained annotators; They found that his ChatGPT’s zero-shot accuracy was higher than his MTurk’s on four tasks. ChatGPT outperforms MTurk and trained annotators on all inter-coder consensus features.
Also, ChatGPT is much more affordable than MTurk. His five classification jobs in ChatGPT cost about $68 (25,264 annotations), while the same task in MTurk costs $657 (12,632 annotations). So ChatGPT costs only $0.003, or a third of a penny, making him about 20 times more affordable than MTurk while offering excellent quality. It is possible to annotate the entire sample at this cost or build a large training set for supervised learning.
They tested 100,000 annotations and found that it cost about $300. These findings show how ChatGPT and others of his LLMs could change the way researchers conduct data annotation, upending some aspects of the business model of platforms such as MTurk. However, further research is needed to fully understand how ChatGPT and other LLMs work in a broader context.
Please check paper. All credit for this research goes to the researchers of this project.Also, don’t forget to participate 17,000+ ML SubReddit, Discord channeland email newsletterShare the latest AI research news, cool AI projects, and more.
Aneesh Tickoo is a consulting intern at MarktechPost. He is currently pursuing his Bachelor of Science in Data Science and Artificial Intelligence from the Indian Institute of Technology (IIT), Bhilai. He spends most of his time working on projects aimed at harnessing the power of machine learning. His research interest is image processing and he is passionate about building solutions around it. He loves connecting with people and collaborating on interesting projects.
🔥 Gain a competitive edge with data: Actionable market intelligence for global brands, retailers, analysts and investors. (with sponsorship)
