
Large-scale language models (LLMs) have achieved remarkable success in various domains, but their centralized training requires huge data collection and annotation efforts, which are costly for individual parties. Federated learning (FL) has emerged as a promising solution that can jointly train LLMs on distributed data while preserving privacy (FedLLM). Frameworks such as OpenFedLLM, FederatedScope-LLM, and FedML-LLM have been developed with methods to address the data quality, intellectual property, privacy, and resource constraints of FedLLM, but the lack of realistic benchmarks remains a major challenge. Current research builds artificial FL datasets by partitioning centralized datasets, but fails to capture the characteristics of real-world cross-user data.
Various methods have been proposed to address data heterogeneity in federated learning, which is a major challenge when client datasets come from different distributions. These include regularization, gradient correction, feature tuning, aggregation weight tuning, introducing momentum, and leveraging pre-trained models. FedLLM has recently attracted attention with frameworks such as OpenFedLLM, FederatedScope-LLM, and FedML-LLM, as well as methods such as FedbiOT for model property protection and FFA-LoRA for differential privacy, but it still suffers from significant limitations. Previous studies have evaluated artificially created federated datasets by partitioning centralized datasets, but they fail to capture the complexity of real-world cross-user data.
Researchers from Shanghai Jiao Tong University, Tsinghua University, and the Shanghai AI Research Institute FedLLM Benchis the first realistic benchmark for FedLLM. It provides a comprehensive testbed with four datasets: Fed-Aya (reconciliation of multi-language instructions), Fed-WildChat (reconciliation of multi-turn chat instructions), Fed-ChatbotIT (reconciliation of single-turn chat instructions), and Fed-ChatbotPA (reconciliation of preferences). These datasets are naturally split by real user ID across 38-747 clients and capture realistic federation properties such as data split across devices. The datasets exhibit diversity in language, data quality, quantity, sequence length, and user preferences, reflecting real-world complexity. FedLLM-Bench integrates these datasets with eight baseline methods and six evaluation metrics to facilitate comparison of methods and exploration of new research directions.
FedLLM-Bench is presented from four perspectives: training method, dataset, dataset analysis, and evaluation metrics. For the training method, it covers eight baseline FL methods, including FedAvg, FedProx, SCAFFOLD, FedAvgM, FedAdagrad, FedYogi, and FedAdam, in addition to federated instruction tuning and configuration tuning tasks using parameter-efficient LoRA fine-tuning. The benchmark includes four diverse datasets, including Fed-Aya (multilingual instruction tuning), Fed-ChatbotIT, Fed-WildChat, and Fed-ChatbotPA, capturing realistic characteristics such as different languages, quality, quantity, length, and user preferences. Extensive dataset analysis reveals inter- and intra-dataset diversity in aspects such as length, instructions, quality, embeddings, and quantity. Six metrics are used in the evaluation: four open-ended (MT-Bench, Vicuna bench, AdvBench, Ref-GPT4) and two closed-ended (MMLU, HumanEval).
In the benchmark, we evaluate the implemented methods across a range of datasets. On multilingual Fed-Aya, most federated methods outperform local training on average, but no single method dominates across all languages, highlighting the opportunity for language personalization. For Fed-ChatbotIT, all federated approaches enhance the ability to follow instructions over local training without compromising general functionality, with FedAdagrad performing best overall. On Fed-WildChat, federated methods consistently outperform local training in single-turn and multi-turn conversations, with FedAvg proving to be the most effective in multi-turn. For preference adjustment on Fed-ChatbotPA, federated training improves the ability to follow instructions and safety compared to local, with FedAvgM, FedProx, SCAFFOLD, and FedAvg performing best. Across datasets, federated learning shows clear advantages over individual training by leveraging joint data.
In this work, the researchers present FedLLM-Bench, the first realistic benchmark for FedLLM. At its core is a suite of four diverse datasets spanning instruction alignment and configuration tuning tasks, representing real-world characteristics such as a wide variety of languages, data quality, quantity, instruction style, sequence length, embeddings, and user preferences across 38-747 clients. Integrating eight training methods, four training datasets, and six evaluation metrics, FedLLM-Bench's extensive experiments benchmark traditional federation approaches and explore research directions such as cross-linguistic collaboration and differential privacy. By providing a comprehensive and practical testbed that reflects real-world complexities, FedLLM-Bench aims to reduce effort, enable fair comparisons, and drive progress in the emerging field of FedLLM. This timely benchmark will greatly benefit the research community working on collaborative and privacy-preserving training of large-scale language models.
Please check paper. All credit for this research goes to the researchers of this project. Also, don't forget to follow us. twitter. participate Telegram Channel, Discord Channeland LinkedIn GroupsUp.
If you like our work, you will love our Newsletter..
Please join us 44k+ ML Subreddit

Asjad is an Intern Consultant at Marktechpost. He is pursuing a B.Tech in Mechanical Engineering from Indian Institute of Technology Kharagpur. Asjad is an avid advocate of Machine Learning and Deep Learning and is constantly exploring the application of Machine Learning in Healthcare.
