Jackwang111/M2RL-RL_Coding
Jackwang111/M2RL-RL_Coding is a 4 billion parameter language model developed by Haoqing Wang et al. from Samsung Research and Peking University, focusing on multi-domain reinforcement learning for large language models. This model is part of the M2RL collection, exploring methods for mixing or merging models to enhance performance across various domains. It is specifically designed for post-training research, including weight merging and multi-teacher on-policy distillation, offering a foundation for advanced RL applications in LLMs. The model's architecture and training methodology are detailed in a paper accepted to COLM 2026.
Loading preview...
Overview
Jackwang111/M2RL-RL_Coding is a 4 billion parameter language model developed by Haoqing Wang, Xiang Long, Ziheng Li, Yilong Xu, Tingguang Li, and Yehui Tang from Samsung Research and Peking University. This model is a key component of their research into "To Mix or To Merge? Toward Multi-Domain Reinforcement Learning for Large Language Models," with the paper accepted to COLM 2026.
Key Capabilities
- Multi-Domain Reinforcement Learning: The model is designed to facilitate research in applying reinforcement learning across multiple domains for large language models.
- Post-Training Research: It is specifically released to support further research in areas such as weight merging and multi-teacher on-policy distillation.
- Research Foundation: Provides model checkpoints for the M2RL project, enabling other researchers to build upon this work.
Good For
- Academic Research: Ideal for researchers exploring novel techniques in multi-domain reinforcement learning for LLMs.
- Model Merging Experiments: Suitable for experiments involving the merging of model weights from different domains.
- On-Policy Distillation: Useful for developing and testing multi-teacher on-policy distillation strategies.
- Understanding M2RL: Provides a practical resource for those interested in the M2RL framework and its applications.