To increase the relevance of responses generated by Dropbox Dash, Dropbox engineers began using LLM to augment human labeling. This plays an important role in identifying the documents that should be used to generate the response. Their approach provides useful insights for any system built on search augmentation generation (RAG).
As Dmitriy Meyerzon, principal engineer at Dropbox, explains, document retrieval quality is a bottleneck for RAG systems that select relevant content from large document repositories before passing them to LLM.
Because there are millions (billions for very large companies) of documents in the enterprise search index, Dash can only pass a small portion of the documents it retrieves to LLM. Therefore, the quality of the search rankings and the labeled relevance data used to train them is critical to the quality of the final answers.
This means that the quality of your search ranking model directly impacts the answers it generates. Dash uses a ranking model that is trained using supervised learning techniques. In this model, query-document pairs are labeled according to how well each document satisfies a particular query. The main challenge of this approach lies in producing a large number of high-quality relevance labels.
To address the limitations of purely human judgment-based labeling, which is expensive, time-consuming, and inconsistent, Dropbox introduced a complementary approach in which LLM generates relevance judgments at scale. This method is cheap, consistent, and easily scales to large document sets. However, LLMs are not perfect raters and their judgments must be evaluated before being used for training.
In practice, using LLM for relevance assessment requires a structured process that combines automation and human oversight.
This approach, called “Human-moderated LLM labeling,” is simple. A small, high-quality dataset is human-labeled and later used to tune the LLM evaluator. LLM then generates hundreds of thousands or even millions of labels, amplifying human effort approximately 100 times. Importantly, the LLM is not a replacement for a ranking system. Using LLM directly for query-time ranking would be too time-consuming and limited by context.
The evaluation step involves comparing the relevance ratings generated by the LLM to human judgments on a test subset of query-document pairs that are not included in the training set. The evaluation also highlights the most significant mistakes where LLM decisions do not match user behavior, such as users clicking on documents with low ratings in LLM or skipping documents with high ratings in LLM. This produces the strongest learning signal.
One important consideration is that context is often important in determining relevance. For example, Dropbox’s “diet sprite” refers to an internal performance tool, not a beverage. To address this, LLM is now able to perform additional searches, examine context, and understand internal terminology, greatly improving labeling accuracy.
Based on Dropbox Dash’s experience, Meyerzon says this approach allows LLM to consistently amplify human judgment at scale, proving an effective way to improve RAG systems.
