Jackwang111/M2RL-MT_OPD
Jackwang111/M2RL-MT_OPD is a 4 billion parameter language model developed by Samsung Research and Peking University, focusing on multi-domain reinforcement learning for large language models. This model is specifically designed for post-training research, including weight merging and multi-teacher on-policy distillation. It offers a 32768 token context length, making it suitable for advanced research in RLHF and model adaptation.
Loading preview...
Overview
Jackwang111/M2RL-MT_OPD is a 4 billion parameter language model developed by a collaboration between Samsung Research and Peking University. It is a product of research presented at COLM 2026, focusing on "Multi-Domain Reinforcement Learning for Large Language Models." The model is released with a substantial context length of 32768 tokens, providing ample capacity for complex tasks.
Key Capabilities
- Multi-Domain Reinforcement Learning: Designed to explore techniques for combining or merging knowledge from various domains within reinforcement learning frameworks for LLMs.
- Post-Training Research: Specifically intended as a base for further research in areas like weight merging and multi-teacher on-policy distillation.
- High Context Length: Features a 32768 token context window, enabling the processing of extensive inputs and facilitating more complex reasoning or generation tasks.
Good For
- Academic Research: Ideal for researchers investigating advanced reinforcement learning techniques for large language models.
- Model Adaptation Studies: Useful for experiments involving the adaptation of LLMs to new domains or tasks through methods like weight merging.
- On-Policy Distillation: Provides a foundation for exploring multi-teacher on-policy distillation strategies to improve model performance or efficiency.